Google's cheapest 2026 Gemini tier โ sub-cent pricing at scale
Gemini 3 Flash-Lite is Google's 2026 budget tier, sitting below Gemini 3 Flash. It targets the price-sensitive high-volume use cases โ classification, extraction, summarization, chat boilerplate โ where GPT-4o-mini, Claude Haiku 4.5, and Gemini 2.5 Flash-Lite used to fight. It keeps the 1M-token context window, adds stronger multilingual support, and ships with the lowest list pricing of any major-provider 2026 model.
Who it's for: Teams running massive classification/extraction pipelines, builders of real-time chat features, and anyone who needs "good enough" LLM calls at a price that rounds to zero.
P50 latency under 400ms on typical prompts, fast enough to sit inline in a UI without a spinner.
Retains the Gemini long-context window, useful for batch document work and long transcript parsing.
Roughly 1/3 the price of Gemini 3 Flash and an order of magnitude below flagships โ viable for massive parallelizable jobs.
Strong on low-resource languages that smaller open-weight models mangle.
If your workload is high-volume and "good enough" really is good enough, Gemini 3 Flash-Lite is the obvious 2026 default โ nothing else at the frontier labs comes close on cost-per-token. For anything nuanced, step up to Flash or Pro.