buildfastwithaibuildfastwithai
AI WorkshopsAll blogsAgentic AI Launchpad
Agentic AI Launchpad
Download Unrot App
Free AI Workshop
Mentorship

Agentic AI Launchpad

Go from user to builder in 6 weeks.

Explore Program
Claude Mastery Course
Share
Back to blogs
Analysis
Benchmarks
Reviews

Gemini 3.5 Flash-Lite Review: Price, Speed & Benchmarks

July 22, 2026
15 min read
Share:
Gemini 3.5 Flash-Lite Review: Price, Speed & Benchmarks
Share:

Gemini 3.5 Flash-Lite Review: The Cheap Tier Grew Up

Google's smallest new model beats a bigger one. Gemini 3.5 Flash-Lite scores 54.2 percent on SWE-Bench Pro against 49.6 percent for Gemini 3 Flash, and 74.0 percent on OSWorld-Verified against 65.1 percent. A model that costs 0.30 dollars per million input tokens is outperforming a model from the tier above it.

That result is the story of this release, and almost nobody covered it because Gemini 3.6 Flash launched the same day and took all the attention. Flash-Lite shipped on July 21, 2026 as the budget option in a three-model drop, and it quietly did something the budget option is not supposed to do.

I want to be precise about what this does and does not mean, because cheap-tier hype gets ahead of reality fast. Flash-Lite is not beating the current mid-tier model, it is beating a previous generation one. It also has a real latency problem that makes its positioning slightly strange. Here is the honest breakdown, including where I think Google's own framing is misleading.

What Is Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite is Google's cheapest and fastest current model, released on July 21, 2026 and built for high-throughput, low-latency workloads like agentic search and document processing. It is the entry tier of the Flash family, sitting below Gemini 3.6 Flash, and it targets jobs where you make an enormous number of inexpensive calls rather than a few expensive ones.

The specs are more serious than the Lite name suggests. It carries a full one million token context window, roughly 1,500 A4 pages, and accepts text, image, speech, and video as input while producing text output. It is classified as a reasoning model, not a stripped-down completion model, which is the first sign that Google's tiering is about economics rather than a hard capability ceiling.

It launched alongside Gemini 3.6 Flash and the restricted-access Gemini 3.5 Flash Cyber, and our Google Gemini and Google AI collection tracks how each tier in the family compares as new models ship.

The intended jobs are specific: agentic search where a system fires hundreds of queries to gather information, document processing at volume, classification and extraction pipelines, and any workload where per-call cost multiplied by call count is the number that decides whether your product is viable.

Flash-Lite is not built to give you the best answer. It is built to give you a good enough answer ten thousand times before lunch.

Pricing: The Cheapest Serious Model Google Sells

Gemini 3.5 Flash-Lite costs 0.30 dollars per million input tokens and 2.50 dollars per million output tokens. Against Gemini 3.6 Flash at 1.50 and 7.50 dollars, that is five times cheaper on input and three times cheaper on output.

Run those multiples through a real workload and the gap stops being abstract. Process a million documents averaging 2,000 input tokens and 200 output tokens each, and Flash-Lite costs roughly 1,100 dollars while 3.6 Flash costs about 4,500 dollars. Same job, same day, one number a business approves without a meeting and one that needs a forecast.

There is a nuance worth flagging, because Artificial Analysis ranks Flash-Lite around 87th of 152 models on price, describing its input as somewhat expensive and its output as expensive. That reads as contradictory until you notice what the ranking includes: a long tail of tiny open-weight models that cost fractions of a cent and cannot do this work. Against models with comparable capability, Flash-Lite is inexpensive. Against every model in existence, it is mid-priced. Both statements are true and only one is useful.

For a sense of how these price tiers have moved over the last few generations, our Gemini 3.5 Flash review with pricing and API details documented the previous baseline, and the direction has been consistently downward.

The model is also economical with tokens, which compounds the saving. In Artificial Analysis testing it produced 43 million output tokens against a median of 58 million for comparable models, meaning it says less to reach the same place. Concision is an underrated cost lever, and it is the same trait Google engineered into 3.6 Flash.

Benchmarks: A Generational Jump, Not a Trim

Gemini 3.5 Flash-Lite improves on its predecessor, Gemini 3.1 Flash-Lite, by margins that look like a new model rather than a refresh. Coding, computer use, and terminal work all jumped by 16 to 23 percentage points.

Benchmarks: A Generational Jump, Not a Trim

Terminal-Bench nearly doubling, from 31.0 to 54.0 percent, is the standout. That eval measures whether a model can operate a command line competently, which is a prerequisite for any agent that touches infrastructure. A budget model crossing 50 percent there changes what you can reasonably delegate to the cheap tier.

On the wider field, Flash-Lite scores 36 on the Artificial Analysis Intelligence Index, ranking 13th of 152 models, against an average of 16 for comparable models. A cheap model landing in the top 15 on an intelligence index is not what the Lite label historically meant.

If you want to see where that places it against every current frontier release rather than just its own family, our ranked analysis of the best AI models in 2026 scores the whole field on one scale.

The Result That Should Worry the Tier Above

The most consequential number in this release is that Gemini 3.5 Flash-Lite beats the larger Gemini 3 Flash on two benchmarks: SWE-Bench Pro at 54.2 against 49.6 percent, and OSWorld-Verified at 74.0 against 65.1 percent. Google published this comparison itself, which is a confident thing to do.

Be careful with what it proves. Gemini 3 Flash is an earlier generation, so this is a new small model beating an old medium model, which is the normal rhythm of progress rather than a violation of it. It is not beating Gemini 3.6 Flash, and it does not come close on the intelligence index, 36 against 50.

What it does prove is that the capability floor is rising faster than the tier structure. If your product was built on Gemini 3 Flash and you have been avoiding the Lite tier on the assumption that cheap means weak, that assumption is now wrong by a measurable margin, and you are paying more for less on at least two dimensions.

My take: the interesting pressure here is internal. Google is selling a model at one fifth the input price that outperforms its own previous mid-tier on coding and computer use. Every generation, the cheap tier absorbs more of what the mid tier was for, and the mid tier has to justify itself on harder ground. That is excellent for anyone building on these APIs and uncomfortable for anyone whose pricing depends on capability tiers staying neatly separated.

The cheap tier is not catching up to the mid tier. It is catching up to what the mid tier was one generation ago, on schedule, every time.

πŸš€ Cohort Waitlist Open
Go From AI User to AI Builder

Don't just use ChatGPT. Learn to build custom LLM agents, RAG pipelines, and full-stack Agentic AI apps in our intensive 6-week program.

6 Weeks Live Mentorship
Deploy 5+ Real-world Apps
Weekly App Templates & Code
No Coding Experience Required
Explore Program
Join 1,000+ graduatesβ€’Free Registration

Speed and the Latency Contradiction

Gemini 3.5 Flash-Lite generates 489.9 output tokens per second, second fastest of 152 models Artificial Analysis tracks, but its time to first token is 11.73 seconds. For a model marketed at latency-sensitive workloads, that second number is a genuine contradiction.

Throughput and latency are different things and the difference matters here. Once Flash-Lite starts producing output it is close to the fastest thing available, which is excellent for bulk generation. But the wait before the first token appears is long, and in a live interface, that is the part users experience as slowness.

Reported figures vary, with some sources citing a lower time to first token around 7 seconds and output speed near 350 tokens per second, likely reflecting different providers, regions, and reasoning settings. Treat any single latency number as a starting point and measure on your own infrastructure, because this is the spec most sensitive to how you deploy.

Where this profile genuinely shines is asynchronous, high-volume work. Firing ten thousand classification calls in parallel, processing a document backlog overnight, running agentic search where the user is waiting for a final synthesized answer rather than a token stream. In all of those, front-loaded latency is invisible and raw throughput is everything.

The same throughput-versus-latency split showed up in the previous Lite generation, and we measured it directly in our Gemini 3.1 Flash Lite versus 2.5 Flash speed and cost comparison, which is a useful baseline for how Google's Lite tiers behave under load.

Flash-Lite vs 3.6 Flash: Which One Do You Actually Need?

Choose Gemini 3.5 Flash-Lite when call volume drives your cost and good enough is genuinely good enough. Choose Gemini 3.6 Flash when a single task's quality determines whether the product works at all. The gap in capability is real but smaller than the gap in price.

The honest comparison: 3.6 Flash scores 50 on the intelligence index against 36 for Flash-Lite, and leads on every published benchmark including SWE-Bench Pro at 58.7 against 54.2 percent and OSWorld-Verified at 83.0 against 74.0. It also has far stronger long-context retrieval, 91.8 against 72.2 percent, which is the widest gap between the two models and the clearest reason to pay up.

But you are paying five times more on input for those gains. A 4.5 point SWE-Bench difference is not worth a 5x price multiplier on a classification pipeline. It absolutely is worth it on an autonomous coding agent where a wrong fix costs an engineer an hour.

A pattern I would actually recommend: route by task, not by product. Use Flash-Lite for the high-volume, low-stakes calls in your system, retrieval, classification, extraction, first-pass filtering, and escalate to 3.6 Flash only for the steps where quality decides the outcome. Most teams pick one model for everything and overpay on 90 percent of their calls to protect the other 10 percent.

For a structured way to run that evaluation across vendors rather than within one family, our comparison of Gemini 3.5 Flash against GPT-5.5, Claude, and DeepSeek lays out the testing method, and it transfers directly to the Lite tier.

Where Flash-Lite Is the Wrong Choice

Flash-Lite is the wrong model for interactive chat, long-document reasoning, and anything where a single wrong answer is expensive. Its weaknesses are specific and predictable, which at least makes them easy to design around.

  • Interactive chat interfaces, because a time to first token near 11.73 seconds reads as broken to a user watching a blank screen, no matter how fast the streaming is afterwards.
  • Long-context reasoning, where 72.2 percent on GDM-MRCR v2 is a large drop from 91.8 percent on 3.6 Flash. It has the million token window but it is measurably less reliable at using the far end of it.
  • High-stakes single decisions, like medical, legal, or financial reasoning, where the correct move is the best model you can afford rather than the cheapest one that usually works.
  • Frontier coding work, since even 3.6 Flash trails GPT-5.6 Luna and Grok 4.5 on coding benchmarks, and Flash-Lite sits below 3.6 Flash.

There is also a subtler risk with cheap models that nobody puts in a spec sheet. When calls are almost free, teams stop rationing them, and volume quietly grows until the cheap model is producing a large fraction of your output with nobody checking quality at the tail. Cheap tokens change behavior, not just budgets. Set up sampling and evaluation before you scale volume, not after the complaint arrives.

How to Access Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite is available now through the Gemini API in Google AI Studio and Android Studio, in the Gemini Enterprise app, and rolling out across the Gemini app and Google Search. AI Studio is the fastest way to test it, and you can compare it against 3.6 Flash in the same session by swapping the model string.

Three things to set up before you scale. Turn on context caching if you re-send the same instructions or documents across calls, since that discount applies across the Gemini family and most teams forget it on the cheap tier where it matters most in aggregate. Batch your requests to exploit the 489.9 tokens per second throughput rather than firing them serially. And benchmark time to first token from your own region, because reported figures vary widely and it is the number most likely to surprise you in production.

If you want runnable patterns for high-volume pipelines rather than single API calls, the gen-ai-experiments cookbook collection has batching, extraction, and retrieval implementations you can point at Flash-Lite and measure in an afternoon.

One timing consideration. Google teased Gemini 4 in the same week it shipped these three models, and the Lite tier historically updates shortly after a flagship generation lands. If you are architecting around Flash-Lite specifically, keep your model selection configurable rather than hardcoded, because this tier moves fast and the next jump will likely be as large as the one from 3.1 Flash-Lite.

πŸš€ Cohort Program Open
Claude Mastery: Cowork & Code

The only comprehensive program designed to take you from basic prompting to building interactive Artifacts, custom integrations, and deploying production-ready code with Claude Code.

No coding experience needed
Build interactive Artifacts & Agents
Deploy apps with Claude Code
Cohort-based learning & mentorship
Explore Program
Cohort-based trainingβ€’Register Now

Frequently Asked Questions

What is Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite is Google's cheapest and fastest current model, released on July 21, 2026 for high-throughput, low-latency workloads like agentic search and document processing. It has a one million token context window, accepts text, image, speech, and video input, and is classified as a reasoning model despite sitting in the budget tier.

How much does Gemini 3.5 Flash-Lite cost?

It costs 0.30 dollars per million input tokens and 2.50 dollars per million output tokens. That makes it five times cheaper on input and three times cheaper on output than Gemini 3.6 Flash, which is priced at 1.50 and 7.50 dollars respectively. It is also token-efficient, producing 43 million output tokens in testing against a 58 million median for comparable models.

Is Gemini 3.5 Flash-Lite better than Gemini 3.6 Flash?

No, 3.6 Flash is stronger on every published benchmark, scoring 50 on the Artificial Analysis Intelligence Index against 36 for Flash-Lite, with the widest gap in long-context retrieval at 91.8 versus 72.2 percent. Flash-Lite wins only on price and raw output speed. Pick it when call volume drives your costs, not when single-task quality decides your outcome.

How fast is Gemini 3.5 Flash-Lite?

It generates 489.9 output tokens per second, ranking second out of 152 models tracked by Artificial Analysis, but its time to first token is 11.73 seconds. That combination suits batch and asynchronous work well and interactive chat poorly. Reported latency varies by provider and region, so measure it on your own setup.

What is the context window of Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite has a one million token context window, roughly 1,500 A4 pages of text. However, its long-context retrieval accuracy is 72.2 percent on GDM-MRCR v2, well below the 91.8 percent that Gemini 3.6 Flash achieves, so the window is large but less reliable at depth.

Is Gemini 3.5 Flash-Lite good for coding?

It is surprisingly capable for a budget model, scoring 54.2 percent on SWE-Bench Pro and 54.0 percent on Terminal-Bench 2.1, both large jumps over Gemini 3.1 Flash-Lite at 38.3 and 31.0 percent. It even beats the older, larger Gemini 3 Flash on SWE-Bench Pro. For frontier coding work, though, GPT-5.6 Luna and Grok 4.5 remain ahead.

What is Gemini 3.5 Flash-Lite best used for?

Agentic search, document processing at volume, classification and extraction pipelines, and any workload where you make a very large number of calls and per-call cost is the deciding factor. It also handles computer-use tasks well, scoring 74.0 percent on OSWorld-Verified, which suits automation agents that operate desktop interfaces.

Where can I access Gemini 3.5 Flash-Lite?

It is available through the Gemini API in Google AI Studio and Android Studio, in the Gemini Enterprise app, and across the Gemini app and Google Search. Google AI Studio is the quickest place to test it and lets you compare it directly against Gemini 3.6 Flash by changing the model string.

Recommended Blogs

  • Gemini 3.5 Flash Review
  • Gemini 3.1 Flash Lite Speed Test
  • Best AI Models 2026 Ranked
  • Gemini 3.5 Flash vs GPT-5.5
  • GPT-5.4 Review and Benchmarks
  • GPT-5.4 vs Gemini 3.1 Pro
  • Gemini Omni Video Model

Resources & Community

Join our community of 70,000+ AI enthusiasts and learn to build powerful AI applications! Whether you're a beginner or an experienced developer, Build Fast with AI helps you understand and implement AI in your projects.

  • Website: buildfastwithai.com
  • LinkedIn: Build Fast with AI
  • Instagram: @buildfastwithai
  • Founder Twitter: @satvikps
  • Twitter: @BuildFastWithAI

Agentic AI Launchpad 2026

A structured 6-week cohort program that takes you from AI basics to building and deploying real-world agentic AI systems. Includes live sessions, expert mentorship, project reviews, and a builder community network.

Ready to go from learning to building? Join the next cohort: Agentic AI Launchpad 2026

Free AI Resources

Access free tools, workshops, and micro-learning to keep building:

  • AI Workshops: Free resources, upcoming events & past recordings
  • Unrot: Learn AI in 5 minutes a day (free micro-learning app)

The cheap tier gets better every quarter and most teams never re-test it. Subscribe to Build Fast with AI for the benchmark breakdown the week each model ships.

References

  • Flash-Lite launch announcement (Google)
  • Intelligence, speed and price analysis (Artificial Analysis)
  • Provider performance benchmarking (Artificial Analysis)
  • Benchmarks and model tests (DataCamp)
  • Cheaper token-efficient Flash tier (MarkTechPost)
  • Launch coverage and Gemini 4 tease (9to5Google)
  • Flash-Lite versus 3.6 Flash comparison (LLM Stats)

Benchmark roundup across models (OfficeChai)

Enjoyed this article? Share it β†’
Share:
    You Might Also Like
    Best AI Models April 2026: Ranked by Benchmarks
    Interactive
    Best AI Models April 2026: Ranked by Benchmarks

    GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, GLM-5 - every major AI model ranked by SWE-bench, ARC-AGI-2, and real-world scores. April 2026 breakdown.

    Latest AI Models April 2026: Rankings & Features
    Benchmarks
    Latest AI Models April 2026: Rankings & Features

    Meta Description GPT-5.4, Gemini 3.1 Ultra, Gemma 4, Muse Spark, GLM-5.1: every major AI model released March-April 2026, compared by benchmark, price, and use case.