Google's cheapest, fastest model โ built for high-volume, latency-sensitive tasks.
Gemini 2.5 Flash-Lite is Google's answer to 'I need a million classifications a day.' It trades a little quality for the lowest price and latency in the Gemini lineup, making it ideal for summarization, tagging, routing, and other high-throughput jobs where cost-per-token dominates.
Who it's for: Engineering and product teams running massive volumes of simple AI tasks on a budget.
Among the cheapest capable models per token.
Low latency for real-time pipelines.
Inherited million-token window for big inputs.
Text, image, and more in one model.
Flash-Lite is the right call when you're processing huge volumes and every fraction of a cent matters. For quality-sensitive work, step up to Gemini Flash or Pro.