Gemini 3.5 Flash
Gemini 3.5 Flash is a fast frontier large language model developed by Google DeepMind, announced at Google I/O 2026 on May 19, 2026 and made generally available the same day [1][2]. It is the first model in the Gemini 3.5 family and the successor to Gemini 3 Flash, which shipped in December 2025 [3][4]. Google positioned the model around agentic execution and coding rather than conversational chat, describing it as the company's "most intelligent Flash model" and claiming it surpasses the larger Gemini 3.1 Pro on most coding and agentic evaluations while running about four times faster on output [1][5].
Overview
Gemini 3.5 Flash sits in the "Flash" tier of the Gemini line, the segment Google tunes for low latency and high throughput rather than maximum raw capability. With this release Google argued that the speed-optimized tier had caught up to, and on several axes overtaken, the previous flagship Pro model [1][6]. The pitch centered on long-horizon agentic work: tasks where a model plans, calls tools, writes and runs code, and iterates over many steps with limited human input [5].
The model is multimodal on input. It accepts text, images, video, audio, and PDF documents, and returns text [7]. It carries a one-million-token context window and supports tool use including Google Search grounding, code execution, file search, URL context, and function calling [8]. Notably, the computer-use capability available in some earlier Gemini releases is not supported in Gemini 3.5 Flash at launch [8].
Announcement and release
Google unveiled Gemini 3.5 Flash during the Google I/O 2026 developer keynote on May 19, 2026, and shipped it to production the same day rather than as a staged preview [1][2]. Koray Kavukcuoglu, who leads Google DeepMind's model work, framed the launch around the shift from "AI as a conversational tool to AI as an agentic tool," and said the model offers "an incredible combination of quality and low latency" [2]. Tulsee Doshi noted the model could "run autonomously for multiple hours," pausing to check in with a person at decision points [2].
At launch the model carried the API identifier gemini-3.5-flash and replaced the earlier preview identifier used for Gemini 3 Flash [8]. Google described it as generally available, stable, and ready for scaled production use [9]. The same release powered consumer-facing features announced at I/O, including the model behind the Gemini app and AI Mode in Google Search [2].
Place in the Gemini line
Gemini 3.5 Flash is the opening model of the Gemini 3.5 generation. Its direct predecessor in the Flash tier is Gemini 3 Flash (released December 17, 2025), and the wider 3.x family includes Gemini 3 Pro and the intermediate Gemini 3.1 Pro that Google used as its comparison baseline at launch [3][4][6]. Google reported that 3.5 Flash outperforms Gemini 3.1 Pro on challenging coding and agentic benchmarks, which several outlets described as the first time a Flash-tier model had surpassed a Pro-tier model on those workloads [1][5].
A larger sibling, Gemini 3.5 Pro, was previewed as "rolling out next month," with reporting putting that timeline in June 2026 [1][9][10]. Some coverage noted audience disappointment that the Pro variant was not shipping at the event itself [10]. For background on the broader generation, see Gemini 3.
Architecture and training
Google disclosed little about the underlying architecture or training data. The published model card and developer documentation describe the model's interface and behavior rather than its internals, and Google did not release parameter counts or training-corpus details [7][8]. The stated knowledge cutoff is January 2025 [7][9].
What Google did detail is the reasoning interface. Gemini 3.5 Flash uses a thinking_level parameter with values minimal, low, medium, and high, where medium is the default. This changed the default from the high setting used by Gemini 3 Flash, and Google said the low setting was retuned to be markedly stronger on code and agentic tasks [8]. The model also preserves intermediate reasoning across turns of a multi-turn conversation automatically, a behavior Google labels "thought preservation" [8].
Capabilities and modalities
The model accepts text, image, video, audio, and PDF input and produces text output [7]. Google lists its primary strengths as everyday tasks, agentic coding, advanced reasoning, multimodal understanding, and long-context understanding [7]. Tool capabilities include function calling, structured output, Google Search and Google Maps grounding, file search, URL context, and code execution [7][8].
The agentic framing was the centerpiece. Google described the model as able to execute long, multi-step workflows, and reporting cited early uptake among banks and fintechs automating multi-week processes [2][5]. On raw output speed, Kavukcuoglu said the model is "4x faster than other frontier models," and that Google had developed an optimized version that is "12x faster with the same quality" [2].
Benchmark performance
Google reported that Gemini 3.5 Flash leads Gemini 3.1 Pro across most coding, agentic, and multimodal evaluations, though it trails on a few reasoning and long-context tests. The table below lists scores published by Google and compiled in independent launch analyses; comparison figures are versus Gemini 3.1 Pro [1][7][6].
| Benchmark | Category | Gemini 3.5 Flash | Gemini 3.1 Pro |
|---|---|---|---|
| Terminal-Bench 2.1 | Agentic coding | 76.2% | 70.3% |
| SWE-Bench Pro (public) | Coding | 55.1% | 54.2% |
| MCP Atlas | Multi-step tool use | 83.6% | 78.2% |
| Toolathlon | Real-world tool use | 56.5% | 49.4% |
| OSWorld-Verified | Computer tasks | 78.4% | 76.2% |
| Finance Agent v2 | Agentic finance | 57.9% | 43.0% |
| GDPval-AA | Agentic (Elo) | 1656 | 1314 |
| CharXiv Reasoning | Chart reasoning | 84.2% | 83.3% |
| MMMU-Pro | Multimodal reasoning | 83.6% | 80.5% |
| Blueprint-Bench 2 | Multimodal | 33.6% | 26.5% |
| MRCR v2 (128k) | Long context | 77.3% | 84.9% |
| MRCR v2 (1M) | Long context | 26.6% | 26.3% |
| Humanity's Last Exam | Reasoning | 40.2% | 44.4% |
| ARC-AGI-2 | Reasoning | 72.1% | 77.1% |
On the Artificial Analysis Intelligence Index, the high-effort configuration of Gemini 3.5 Flash scored 55 and measured roughly 173 output tokens per second [11]. Google highlighted four headline numbers from the model card: 76.2% on Terminal-Bench 2.1, 83.6% on MCP Atlas, 84.2% on CharXiv Reasoning, and an Elo of 1656 on GDPval-AA [1][7].
Availability, context window, and pricing
Gemini 3.5 Flash launched across Google's developer and enterprise surfaces as well as its consumer apps. It is available through the Google AI Studio interface and the Gemini API, in Android Studio, on Google Vertex AI and the Gemini Enterprise Agent Platform, in Google Antigravity, and to everyone in the Gemini app and AI Mode in Search [1][2][7].
The context window is 1,048,576 input tokens with a maximum of 65,536 output tokens [8][12]. Pricing is volume-based per million tokens; rates differ slightly between Google's global and non-global serving regions, and batch processing is offered at a discount [12][13].
| Item | Rate (per 1M tokens) |
|---|---|
| Input (global) | $1.50 |
| Output (global) | $9.00 |
| Cached input | $0.15 |
| Cache storage | $1.00 per 1M token-hours |
| Input / output (non-global) | $1.65 / $9.90 |
| Batch input / output | $0.75 / $4.50 |
These rates are higher than the prior generation. Simon Willison noted the model is about three times the price of Gemini 3 Flash Preview and six times that of Gemini 3.1 Flash-Lite, and read the increase as the major labs probing how much their API customers will pay [12]. Despite the higher per-token cost, Google deployed the model broadly across free consumer products, which Willison read as an aggressive bet on reach [12].
Reception
Coverage of the launch split between the benchmark story and the pricing story. Outlets including TechCrunch and MarkTechPost emphasized Google's claim that a Flash-tier model now beat its own Pro-tier model on coding and agentic work, and the explicit reorientation toward agents over chatbots [2][5]. The model landed in the top-right quadrant of Artificial Analysis's intelligence-versus-speed chart, the region Google wanted to occupy [1][11].
Other commentary was cooler. Trending Topics characterized the release as "more of a solid incremental improvement than a milestone" and pointed out that the version number is 3.5, not 4, while flagging the "considerable" price increase over the predecessor [10]. The same report noted a muted early standing on community arena rankings and described audience disappointment that Gemini 3.5 Pro was delayed rather than shipped at the event [10].
July 2026 variants
Google expanded the Gemini 3.5 family on July 21, 2026 with Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. Despite the shared version number, the two releases have different relationships to Gemini 3.5 Flash. Flash-Lite is a generally available efficiency model based on Gemini 3.1 Flash-Lite, while Flash Cyber is a specialized fine-tune of Gemini 3.5 Flash that Google announced for a forthcoming restricted pilot [14][17][20]. Independent launch coverage likewise treated Flash-Lite as a public high-volume model and Flash Cyber as a security model limited to selected partners [23].
| Model | Relationship | Public identifier | Status on July 24, 2026 | Primary role |
|---|---|---|---|---|
| Gemini 3.5 Flash | Base model | gemini-3.5-flash | Generally available | Agentic execution, coding, and multimodal work |
| Gemini 3.5 Flash-Lite | Based on Gemini 3.1 Flash-Lite | gemini-3.5-flash-lite | Generally available | High-volume, latency-sensitive, and low-cost work |
| Gemini 3.5 Flash Cyber | Fine-tuned from Gemini 3.5 Flash | No public standalone identifier | Forthcoming limited CodeMender pilot | Vulnerability discovery, validation, and patching |
Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite was released as a stable production model on July 21. Google Cloud lists it as generally available and gives a retirement date of July 21, 2027 or later [18]. It accepts text, images, video, audio, and PDFs, produces text, and has the same published token limits as the base model: 1,048,576 input tokens and 65,536 output tokens [15]. Its intended workloads include translation, classification, document processing, agentic search, and subagent execution [14][17].
The model supports thinking, caching, structured output, code execution, file search, function calling, URL context, and grounding with Google Search and Google Maps. It does not support audio generation, image generation, or the Live API [15]. Computer Use differs by serving surface: the Gemini Developer API page lists it as supported in preview, while the Gemini Enterprise Agent Platform page lists the preview capability as unsupported [15][18]. Google set the default thinking level to minimal for latency and cost. Its Cloud guidance recommends medium or high for autonomous subagents because minimal thinking can end multi-step tool work prematurely [18].
Standard paid Gemini API pricing is $0.30 per million input tokens and $2.50 per million output tokens, including thinking tokens. A standard cache hit costs $0.03 per million tokens, with storage billed at $1.00 per million token-hours. Batch and Flex processing each cost $0.15 per million input tokens and $1.25 per million output tokens; Priority processing costs $0.54 and $4.50 respectively [16]. These rates place Flash-Lite below the base model's $1.50 input and $9.00 output rates, although the default reasoning settings and target workloads also differ.
Flash-Lite evaluation and safety
Google evaluated Flash-Lite across coding, agentic, multimodal, and long-context tests. The published results below compare it with Gemini 3.1 Flash-Lite, the model on which it is based [17]. They are vendor-reported scores, and results can depend on the named harness, tool access, and context length.
| Benchmark and configuration | Gemini 3.5 Flash-Lite | Gemini 3.1 Flash-Lite |
|---|---|---|
| SWE-Bench Pro (Public) | 54.2% | 38.3% |
| Terminal-bench 2.1, Terminus-2 harness | 54.0% | 31.0% |
| MLE-Bench | 39.2% | 22.0% |
| GDPVal-AA v2 | 1140 Elo | 642 Elo |
| OSWorld-Verified | 74.0% | 54.3% |
| CharXiv Reasoning, no tools | 74.5% | 73.2% |
| CharXiv Reasoning, with tools | 76.5% | 75.6% |
| GDM-MRCR v2, 128k average | 72.2% | 60.1% |
| GDM-MRCR v2, 1M pointwise | 21.3% | 12.3% |
Google's launch post cited an Artificial Analysis speed measurement of 350 output tokens per second [14]. The current Artificial Analysis page reports 459.5 output tokens per second, a 10.65-second time to first answer token, and an Intelligence Index score of 36 under its own test configuration [19]. The measurements were recorded under different configurations or at different times, so neither figure should be read as a fixed throughput guarantee.
The Flash-Lite model card lists hallucinations and occasional slowness or timeouts among the known limitations, with a March 2026 knowledge cutoff. Google's automated internal evaluation found better text-to-text and multilingual safety than Gemini 3.1 Flash-Lite, no change in image-to-text safety, improved refusal tone, and a 5.32-percentage-point regression in unjustified refusals, where a lower score is better. Google also reported specialist red teaming outside the model development team, satisfaction of child-safety launch thresholds, and no egregious concerns [17]. Its frontier-safety conclusion was inferred from Gemini 3.1 Pro's assessment rather than derived from a separately published full threshold table for Flash-Lite.
Gemini 3.5 Flash Cyber
Gemini 3.5 Flash Cyber is a cybersecurity-focused variant built on top of Gemini 3.5 Flash and fine-tuned to find, validate, and patch software vulnerabilities [14][20]. It runs inside CodeMender, where multiple subagents can inspect code paths and combine their findings into one report. Google reports using the system internally across Chrome, Android, Cloud, Ads, and YouTube codebases [20].
The announcement was not a public API release. Google said Flash Cyber would become available "soon" in a limited-access CodeMender pilot restricted to governments and trusted partners, with access expanding over time. The company cited the technology's dual-use risk as the reason for the controlled deployment [14][20]. General CodeMender capabilities offered through the Gemini Enterprise Agent Platform with generally available Gemini models are separate from the Flash Cyber pilot [20][22].
Google has not published a standalone Flash Cyber API identifier, numerical token price, context or output limits, modality specification, general-release date, or dedicated model card [20][21][22]. Although the launch post says it has a lower token price than larger cybersecurity models, that statement does not establish a public rate. Specifications and safety findings from the base Gemini 3.5 Flash therefore cannot be assumed to apply unchanged to the specialized model.
Cyber evaluation and restrictions
Google's headline CyberGym result measures the combined CodeMender system, not one raw model call. CodeMender could invoke Flash Cyber up to five times before producing a final report, and the system reached an 83.2% pass@1 success rate on a benchmark covering hundreds of real-world software vulnerabilities. Google noted that competitor figures in the comparison chart were self-reported by their providers [20].
| Evaluation | Configuration | Reported result |
|---|---|---|
| CyberGym | CodeMender with up to five Flash Cyber calls per final report | 83.2% pass@1 |
| V8 JavaScript engine | Fixed number of invocations | 55 unique confirmed issues, versus 47 for Gemini 3.5 Flash and 36 for Claude Opus 4.6 |
| V8 unique findings | Same fixed-invocation test | 10 issues not found by either comparison model |
| Big Sleep evaluation | Complex Chrome and Safari code without safety guardrails | Google reported a substantial lead over mainline Gemini 3.5 Flash and Gemini 3.6 Flash, without publishing a numerical score in the article text |
The Big Sleep test was built by another Google team and intentionally removed safety guardrails, so it was an internal stress test rather than an independent evaluation of deployed pilot behavior [20]. Google also used private, undisclosed vulnerabilities in Chrome commit scanning to reduce contamination, but it did not publish enough dataset detail for outside reproduction. TechRepublic noted further evidence gaps, including unpublished false-positive, patch-acceptance, and regression rates and the absence of independent production testing [22]. The limited access policy reduces exposure to misuse, but no dedicated public frontier-safety assessment for the specialized fine-tune accompanied the announcement.
Limitations
Several limits were disclosed or evident at launch. Computer Use is not supported, unlike some earlier Gemini models [8]. The model output is text only, so it cannot generate images or audio directly [7]. On a handful of evaluations it trails Gemini 3.1 Pro, including Humanity's Last Exam (40.2% versus 44.4%), ARC-AGI-2 (72.1% versus 77.1%), and the 128k-token slice of MRCR v2 (77.3% versus 84.9%), which suggests the gains concentrated in agentic and coding workloads rather than across the board [6]. As with prior Gemini releases, Google did not publish architecture or training-data details, so independent verification of those aspects is not possible [7][8].
References
- ^Google DeepMind. "Gemini 3.5: frontier intelligence with action." blog.google, May 19, 2026. blog.google/...gemini-3-5
- ^Maxwell Zeff. "With Gemini 3.5 Flash, Google bets its next AI wave on agents, not chatbots." TechCrunch, May 19, 2026. techcrunch.com/...t-ai-wave-on-agents-not-chatbots
- ^DataNorth AI. "Google Releases Gemini 3.5 Flash: Frontier-Level Coding and Agentic Performance at 4x Speed." datanorth.ai/...google-releases-gemini-3-5-flash
- ^Digital Applied. "Gemini 3.5 Flash: Benchmarks, Thinking & API Guide 2026." digitalapplied.com/...5-flash-benchmarks-api-guide
- ^MarkTechPost. "Google Introduces Gemini 3.5 Flash at I/O 2026: A Faster and Cheaper Model for AI Agents and Coding." May 20, 2026. marktechpost.com/...model-for-ai-agents-and-coding
- ^LLM Stats. "Gemini 3.5 Flash: Benchmarks, Pricing, and Complete Specs." llm-stats.com/...gemini-3.5-flash-launch
- ^Google DeepMind. "Gemini 3.5 Flash" (model card). deepmind.google/...flash
- ^Google. "What's new in Gemini 3.5 Flash." Gemini API documentation. ai.google.dev/...whats-new-gemini-3.5
- ^Google. "Gemini 3 Developer Guide." Gemini API documentation. ai.google.dev/...gemini-3
- ^Trending Topics. "Google Launches Gemini 3.5 Flash With Higher Prices but No Generational Leap." trendingtopics.eu/...ices-but-no-generational-leap
- ^Artificial Analysis. "Gemini 3.5 Flash: Intelligence, Performance & Price Analysis." artificialanalysis.ai/...gemini-3-5-flash
- ^Simon Willison. "Gemini 3.5 Flash: more expensive, but Google plan to use it for everything." May 19, 2026. simonwillison.net/...gemini-35-flash
- ^OpenRouter. "Gemini 3.5 Flash - API Pricing & Benchmarks." openrouter.ai/...gemini-3.5-flash
- ^Google. "Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber." July 21, 2026. blog.google/...lash-3-5-flash-lite-3-5-flash-cyber
- ^Google. "Gemini 3.5 Flash-Lite." Gemini API model documentation. ai.google.dev/...gemini-3.5-flash-lite
- ^Google. "Gemini Developer API pricing." ai.google.dev/...pricing
- ^Google DeepMind. "Gemini 3.5 Flash-Lite" (model card). Published July 21, 2026. deepmind.google/...gemini-3-5-flash-lite
- ^Google Cloud. "Gemini 3.5 Flash-Lite." Gemini Enterprise Agent Platform documentation. docs.cloud.google.com/...3-5-flash-lite
- ^Artificial Analysis. "Gemini 3.5 Flash-Lite Intelligence, Performance & Price Analysis." artificialanalysis.ai/...gemini-3-5-flash-lite
- ^Google DeepMind. "Introducing Gemini 3.5 Flash Cyber." July 21, 2026. deepmind.google/...roducing-gemini-3-5-flash-cyber
- ^Google DeepMind. "Gemini 3.5 Flash Cyber." deepmind.google/...cyber
- ^TechRepublic. "Google Holds Back Gemini 3.5 Flash Cyber as CodeMender Enters Preview." July 23, 2026. techrepublic.com/...-google-flash-cyber-codemender
- ^Axios. "Google releases series of new cheaper Gemini models." July 21, 2026. axios.com/...google-gemini-ai-models
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
2 revisions · v3 · 3,084 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Cite this page: AI Wiki. "Gemini 3.5 Flash." aiwiki.ai, updated 24 Jul 2026. CC BY 4.0. https://aiwiki.ai/wiki/gemini_3_5_flash