Gemini TTS
Gemini TTS is the family of text-to-speech models that Google builds on its Gemini model line and sells through the Gemini API, Google AI Studio, Vertex AI and Google Cloud Text-to-Speech. Gemini TTS models take natural-language direction about tone, accent, pace and emotion, and can voice a two-person dialogue in a single request.[1][2] The first public models, Gemini 2.5 Flash Preview TTS (gemini-2.5-flash-preview-tts) and Gemini 2.5 Pro Preview TTS (gemini-2.5-pro-preview-tts), were released in the Gemini API on 20 May 2025, the opening day of Google I/O 2025.[3][4] Google upgraded both in December 2025, released Gemini 3.1 Flash TTS as a preview on 15 April 2026, and on 22 September 2026 made Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS generally available, the first Gemini TTS models to leave preview in the Gemini API.[3][5][6] Over those 16 months the documented language count grew from 24 to 130, the fixed set of 30 prebuilt voices was joined by a voice library, voice design and consent-gated voice replication, and the prompt format moved from free-text instructions to bracketed audio tags and then to structured per-turn metadata.[1][7][8]
Overview
Google documents Gemini TTS as a separate capability from the speech that its Live API models produce. The Gemini API guide puts it this way: "While the Live API excels in dynamic conversational contexts, TTS through the Gemini API is tailored for scenarios that require exact text recitation with fine-grained control over style and sound, such as podcast or audiobook generation."[1][8] Every Gemini TTS model page lists text as the only input and audio as the only output, and marks function calling, search grounding, thinking and the Live API as not supported.[9][10][11][12]
The models reach developers through three surfaces:
| Surface | What it offers | Source |
|---|---|---|
| Gemini API and Google AI Studio | Model IDs such as gemini-2.5-flash-preview-tts, gemini-3.1-flash-tts-preview and gemini-3.8-flash-tts, token-based pricing, a free tier for most models | [13][8] |
| Google Cloud Text-to-Speech ("Gemini-TTS") | Model IDs gemini-2.5-flash-tts, gemini-2.5-pro-tts, gemini-2.5-flash-lite-preview-tts and gemini-3.1-flash-tts-preview, alongside Chirp 3: HD voices and older voice types | [2][14] |
| Google products | Google Vids received Gemini 3.1 Flash TTS in April 2026; Gemini Notebook and Google Vids received the 3.8 models in September 2026 | [5][15] |
The Cloud product calls the family "Gemini-TTS" and describes it as "the latest evolution of our Cloud TTS technology that moves beyond natural-sounding speech and provides granular control over generated audio using text-based prompts."[2] Note that the Cloud model IDs for Gemini 2.5 (gemini-2.5-flash-tts, gemini-2.5-pro-tts) differ from the Gemini API IDs, which kept a preview suffix.[2][9][10]
Timeline
Status is as of 30 September 2026. Dates are the dates in Google's release notes unless stated.
| Date | Model ID or event | Surface | Status (Sep 2026) | Key change | Source |
|---|---|---|---|---|---|
| 11 Dec 2024 | Gemini 2.0 Flash (experimental) | Gemini API, Vertex AI | Superseded | Google announced "steerable text-to-speech (TTS) multilingual audio" as a 2.0 Flash output, available only to early-access partners | [16] |
| 20 May 2025 | gemini-2.5-flash-preview-tts, gemini-2.5-pro-preview-tts | Gemini API | Preview, no shutdown date announced | First Gemini TTS models; one or two speakers; 24 languages; 30 voices | [3][4][1][6] |
| 20 May 2025 | gemini-2.5-flash-preview-native-audio-dialog | Live API | Shut down 20 Oct 2025 | Native audio output for real-time dialogue (sibling, not a TTS model) | [3] |
| 30 Sep 2025 | gemini-2.5-flash-tts, gemini-2.5-pro-tts | Cloud Text-to-Speech | Generally available | Gemini-TTS GA on Google Cloud; 30 speakers in "80 or more locales" | [17] |
| 7 Nov 2025 | Streaming synthesis | Cloud Text-to-Speech | Available | Gemini TTS supports streaming requests on Cloud | [17] |
| 10 Dec 2025 | Updated 2.5 Flash and 2.5 Pro TTS | Gemini API | Preview | "Enhanced expressivity, precision pacing, and seamless dialogue"; replaced the May models | [3][18] |
| 11 Dec 2025 | gemini-2.5-flash-lite-preview-tts listed | Cloud Text-to-Speech | Preview | Single-speaker Flash-Lite model named in a regional expansion note | [17][2] |
| 15 Apr 2026 | gemini-3.1-flash-tts-preview | Gemini API, AI Studio, Vertex AI, Google Vids | Preview; labelled "legacy" after the 3.8 launch | Audio tags; 70+ languages | [3][5][11] |
| 17 Jun 2026 | Streaming for gemini-3.1-flash-tts-preview | Gemini API | Available | Streaming via streamGenerateContent supported for the 3.1 preview; the July 2026 guide said TTS otherwise did not support streaming, though the June 2025 guide had shown a streaming example | [3][7][1] |
| 22 Sep 2026 | gemini-3.8-flash-tts, gemini-3.8-flash-lite-tts, /v1beta/voices | Gemini API, AI Studio | Generally available | Voice design, voice replication, voice library; 130 and 101 languages | [3][12][15] |
Google's deprecations page gives 26 February 2026 as the release date of gemini-3.1-flash-tts-preview, which conflicts with the 15 April 2026 date in both the release notes and the launch blog post.[6][3][5] The 3.8 models appear in the release notes under 22 September 2026, while Google's blog post and the @GoogleAI announcement are dated 23 September 2026.[3][15][19]
Background
Google's neural speech synthesis work predates Gemini by years and runs back to DeepMind's WaveNet. Google Cloud's most recent pre-Gemini voice line is Chirp 3: HD, whose documentation says "Powered by our latest generation of generative models, these voices deliver realism and emotional resonance." Chirp 3: HD reached general availability on 2 April 2025 with 8 speakers across 31 locales.[20][17]
Speech output from Gemini itself was first announced with Gemini 2.0 Flash on 11 December 2024. Google's launch post said 2.0 Flash "now supports multimodal output like natively generated images mixed with text and steerable text-to-speech (TTS) multilingual audio," but only "text-to-speech and native image generation available to early-access partners."[16] The same launch introduced the Multimodal Live API for real-time audio and video streaming.[16]
Model generations
Gemini 2.5 Flash and Pro Preview TTS (May 2025)
Google announced the first public Gemini TTS models on 20 May 2025 at I/O. In a post signed by Tulsee Doshi, Google said the new "previews for text-to-speech in 2.5 Pro and 2.5 Flash" had "first-of-its-kind support for multiple speakers, enabling text-to-speech with two voices via native audio out," could "capture really subtle nuances, such as whispers," and worked "in over 24 languages and seamlessly switches between them."[4] The "first-of-its-kind" description is Google's. Engadget reported that the onstage demo started in English, switched to Hindi and returned to English in the same voice.[21]
The Gemini API release notes for that day list gemini-2.5-pro-preview-tts and gemini-2.5-flash-preview-tts as "capable of generating speech with one or two speakers."[3] The June 2025 version of the developer guide documented:[1]
- 30 prebuilt voices, each with a one-word descriptor (for example "Zephyr -- Bright", "Kore -- Firm", "Sulafat -- Warm").
- 24 supported languages detected automatically from the input, including Arabic (Egyptian), Hindi, Japanese, Korean, Thai, Vietnamese, Bengali, Marathi, Tamil and Telugu.
- Multi-speaker output for both models, configured by mapping speaker names in the transcript to voices.
- A 32k-token context window per TTS session.
Style was steered with ordinary instructions written into the prompt; the guide's first example was "Say cheerfully: Have a wonderful day!"[1] Google's developer blog described the models as giving developers control over "TTS expression and style."[22] Google later labelled the 2.5 Flash model "our fastest engine for high-fidelity speech synthesis" and the 2.5 Pro model "our premium engine for studio-quality speech synthesis."[9][10]
December 2025 update
On 10 December 2025 Google announced "significant enhancements" to both 2.5 preview models, listing three changes: "Enhanced expressivity: Richer tone versatility and stricter adherence to style prompts," "Precision pacing: Smarter context-aware speed adjustments and better instruction following," and "Seamless dialogue: Consistent character voices in multi-speaker scenarios."[18] Google said the updated models "will replace our TTS models released in May," and described the Flash variant as "optimized for low latency" and the Pro variant as "optimized for quality." The post still gave the language count as "all 24 supported languages" and named Wondercraft and Toonsutra as customers.[18] The model IDs did not change; the model pages list "December 2025" as the latest update for both.[9][10]
Two days later Google published a companion post on the Live API side, "Improved Gemini audio models for powerful voice interactions," which opened by referring to "an upgrade to our Gemini 2.5 Pro and Flash Text-to-Speech models" earlier that week before turning to an updated Gemini 2.5 Flash Native Audio model.[23]
Gemini-TTS on Google Cloud
Google Cloud Text-to-Speech declared "Gemini-TTS" generally available on 30 September 2025. Its release notes for that day contain two entries with slightly different figures: one says gemini-2.5-flash-tts and gemini-2.5-pro-tts launched "with support for 30 speakers in 80 or more locales," the other that Gemini-TTS "provides support for 30 voices and over 70 locales."[17] Cloud added streaming synthesis for Gemini TTS on 7 November 2025 and, on 11 December 2025, expanded regional availability for gemini-2.5-pro-tts, gemini-2.5-flash-tts and gemini-2.5-flash-lite-preview-tts.[17]
As of September 2026 the Cloud documentation lists four Gemini-TTS models: Gemini 3.1 Flash TTS (Preview), Gemini 2.5 Flash TTS, Gemini 2.5 Flash Lite TTS (Preview), which supports a single speaker only, and Gemini 2.5 Pro TTS. All four share an 8,192-token input limit and a 16,384-token output limit. Unary requests can return LINEAR16, ALAW, MULAW, MP3, OGG_OPUS or PCM audio.[2] The Cloud language table marks 24 locales as GA and 63 as Preview.[2] The Gemini 3.8 models were not yet listed on the Cloud Gemini-TTS page when it was fetched on 30 September 2026; Google's 3.8 announcement says enterprise API access through Gemini Enterprise is "coming soon."[2][15]
Gemini 3.1 Flash TTS Preview (April 2026)
Google released Gemini 3.1 Flash TTS on 15 April 2026 in preview for developers in the Gemini API and Google AI Studio, for enterprises on Vertex AI, and for Workspace users through Google Vids.[5][3] The model ID is gemini-3.1-flash-tts-preview.[11][13]
The launch post, by Vilobh Meshram and Max Gubin, introduced "audio tags," described as "an intuitive way to control vocal style, pace and delivery" by "embedding natural language commands directly into the text input." It also added AI Studio controls that Google framed as putting the developer in the "director's chair": scene direction, per-speaker "Audio Profiles" with "Director's Notes," and export of the settings as Gemini API code.[5] Google said the model supported "more than 70 languages."[5] A July 2026 copy of the developer guide lists 78 languages and shows the tag syntax, with square-bracket tags such as [whispers], [laughs], [sighs] and [very fast], and advises that non-English transcripts "still use English audio tags."[7]
The same guide listed known limitations of the 3.1 preview: output that "may not always strictly match the selected speaker," quality that "may begin to drift with generated outputs that are longer than a few minutes," occasional text tokens returned instead of audio (surfacing as a 500 error), and a prompt classifier that could reject vague prompts or read the director's notes aloud.[7] Streaming output via streamGenerateContent arrived for this model on 17 June 2026; the guide noted that "TTS does not support streaming, except when using gemini-3.1-flash-tts-preview."[3][7]
Since the 3.8 launch, Google's model page calls gemini-3.1-flash-tts-preview "a legacy preview model" and recommends migrating to the 3.8 models; the deprecations page lists no shutdown date.[11][6]
Gemini 3.8 Flash TTS and Flash-Lite TTS (September 2026)
The Gemini API release notes for 22 September 2026 announce both 3.8 TTS models as generally available together with a new Voices endpoint (/v1beta/voices).[3] Google describes Gemini 3.8 Flash TTS as its "flagship creative text-to-speech model" and the release notes describe Gemini 3.8 Flash-Lite TTS as a cost-efficient model "built to replace gemini-3.1-flash-tts-preview for high-throughput production and real-time voice agent cascades."[12][3] The Flash model page lists 130 supported languages for Flash and 101 for Flash-Lite.[12] Google's announcement post, by Leland Rechis and Alan Cowen, introduced generative voice design, a library of "2,000+ production-ready voices," voice replication "from just a 30-second audio sample" backed by consent verification, and inline vocal bursts and backchannels such as <laughs> and |mhm|.[15]
The 3.8 generation changed how prompts are written. Google's migration guide says the 3.8 models treat input text "strictly as a verbatim transcript," so inline directions like "Say cheerfully: Hello!" or "Speaker 1: Hello!" "may be spoken aloud." Sustained style and speaker labels move into a structured speech_metadata field, angle-bracket tags such as <laugh>, <sigh> and <short pause> are reserved for point-in-time events, and the default output for unary requests changed from headerless raw PCM to WAV.[12][8]
Capabilities by generation
| Capability | 2.5 Flash / Pro Preview TTS (2025) | 3.1 Flash TTS Preview (Apr 2026) | 3.8 Flash TTS / Flash-Lite TTS (Sep 2026) |
|---|---|---|---|
| Speakers per request | One or two[3][1] | Single and multi-speaker[2] | Up to 2 speakers with prebuilt voices; custom voices synthesized turn by turn[8] |
| Prebuilt voices | 30[1] | 30 (same set)[7] | 30 featured voices plus an Extended Voice Library; blog cites "2,000+" voices, release notes "150+ prebuilt and custom voices"[8][15][3] |
| Languages (Gemini API docs) | 24[1][18] | "More than 70" (blog); 78 listed in July 2026 guide[5][7] | 130 (Flash) and 101 (Flash-Lite)[12] |
| Style control | Natural-language instructions in the prompt[1] | Square-bracket audio tags, Audio Profile, Scene, Director's Notes[5][7] | speech_metadata.style per turn, angle-bracket vocal events, pipe backchannels[8] |
| Custom voices | None | None | Voice design (text prompt) and voice replication (reference plus consent audio)[8] |
| Streaming (Gemini API) | Not supported per the July 2026 guide, although the June 2025 guide showed a streaming example[7][1] | Supported from 17 Jun 2026[3] | Supported; streamed audio is headerless raw 16-bit PCM[8] |
| Default output | Headerless raw PCM, 24 kHz[12][1] | Headerless raw PCM[12] | WAV with RIFF header[12] |
| Token limits | 8,192 in / 16,384 out[9][10] | 8,192 in / 16,384 out[11] | 8,192 in / 16,384 out[12] |
The 30 voice names used by Gemini TTS (Zephyr, Puck, Charon, Kore and so on) are the same 30 names that Google Cloud lists for Chirp 3: HD voices, and the Cloud documentation describes the Gemini-TTS voices as "similar to our existing Chirp 3: HD Voices."[2][20]
Pricing history
Gemini API prices are per million tokens in US dollars on the paid tier. Google converts audio at 25 tokens per second of output audio.[13]
| Model | Input (text) | Output (audio) | Batch output | Notes | Source |
|---|---|---|---|---|---|
| 2.5 Flash Preview TTS | $0.50 | $10.00 | $5.00 | Same price in June 2025 and September 2026; free tier available | [24][13] |
| 2.5 Pro Preview TTS | $1.00 | $20.00 | $10.00 | No free tier | [24][13] |
| 3.1 Flash TTS Preview | $1.00 | $20.00 | $10.00 | Free tier available | [13] |
| 3.8 Flash TTS | $0.50 until 31 Dec 2026, then $1.00 | $9.00 until 31 Dec 2026, then $18.00 | $4.50, then $9.00 | Priority tier $16.20, then $32.40 | [13] |
| 3.8 Flash-Lite TTS | $0.50 until 31 Dec 2026, then $1.00 | $6.00 until 31 Dec 2026, then $12.00 | $3.00, then $6.00 | Priority tier $10.80, then $21.60 | [13] |
At 25 tokens per second, one minute of audio is 1,500 output tokens. That works out to about $0.015 per minute for 2.5 Flash, $0.03 for 2.5 Pro and 3.1 Flash, and $0.0135 for 3.8 Flash at its 2026 rate, excluding input tokens (this per-minute figure is arithmetic on Google's published rates, not a Google-quoted price). Google's own pricing page expresses the 3.8 Flash rate as "$0.00225 per 10s audio" through the end of 2026.[13]
Google Cloud Text-to-Speech charges the same token rates for its Gemini-TTS models: $0.50 and $10.00 per million input and output tokens for Gemini 2.5 Flash TTS and 2.5 Flash-Lite Preview TTS, and $1.00 and $20.00 for Gemini 3.1 Flash TTS (Preview) and 2.5 Pro TTS. By contrast, Chirp 3: HD voices cost US$30 per million characters after a free first million, and Instant custom voice costs US$60 per million characters.[14]
Related Google speech products
Live API native-audio models
The Live API is Google's real-time, bidirectional audio interface, and its models generate speech natively rather than passing text to a TTS model. At I/O 2025 Google released gemini-2.5-flash-preview-native-audio-dialog alongside the TTS previews; Google's developer blog said it offered "over 30 distinct voices and 24+ languages."[3][22] Later Live API models include gemini-2.5-flash-native-audio-preview-09-2025 (23 September 2025), gemini-2.5-flash-native-audio-preview-12-2025 (12 December 2025), gemini-3.1-flash-live-preview (26 March 2026), and the generally available gemini-3.8-live and gemini-3.8-live-extended-thinking (15 September 2026).[3] Google's December 2025 post said Gemini 2.5 Flash Native Audio had started rolling out in Gemini Live and Search Live.[23] Google's 3.8 TTS announcement groups both lines into a "Gemini Audio family," saying the TTS models follow "3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking" (see Gemini 3.5 Transcribe).[15]
NotebookLM (Gemini Notebook) and Google Vids
NotebookLM's Audio Overview feature, in which "two AI hosts start up a lively 'deep dive' discussion based on your sources," launched on 11 September 2024, eight months before the Gemini TTS API models. Google's launch post did not name the speech model behind it.[25] NotebookLM was renamed Gemini Notebook on 16 July 2026.[26] Google's September 2026 announcement says Gemini 3.8 Flash TTS is rolling out "For everyone: In Gemini Notebook" and Gemini 3.8 Flash-Lite TTS "For everyone: In Google Vids."[15] Google Vids had already received Gemini 3.1 Flash TTS in April 2026.[5]
SynthID and consent safeguards
Google says outputs are watermarked with SynthID. The 3.1 launch post states that "All audio generated by Gemini 3.1 Flash TTS is watermarked with SynthID," and the 3.8 post says "every audio clip generated by our Gemini Audio models is watermarked with SynthID."[5][15] For 3.8 voice replication, Google says "users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created," and that replication is also backed by C2PA credentials.[15] Android Authority reported that Google blocks AI Studio voice replication in several regions, including the UK, the European Economic Area, India, Texas and Illinois.[27]
Benchmarks and reception
Google-reported results
At the 3.1 launch Google said the model "achieved an impressive Elo score of 1,211" on the Artificial Analysis TTS leaderboard, and that Artificial Analysis had placed it in its "most attractive quadrant" for quality relative to price.[5] For 3.8, Google said Flash TTS secured "the #1 overall spot on Hume AI's Voice Design Benchmark (71.4)" and led "in accent modeling (60.8)," and that Flash and Flash-Lite took "the #1 and #2 spots respectively on Hume AI's Overall Quality Index."[15] These are Google's descriptions of third-party results.
Artificial Analysis leaderboard
Artificial Analysis runs a blind-preference "Provider Voice Arena" that compares TTS models using each provider's own voices. When the leaderboard was fetched on 30 September 2026, Google's entries stood as follows:[28]
| Rank | Model (as labelled by Artificial Analysis) | Elo | Samples | Listed price |
|---|---|---|---|---|
| 3 | Gemini 3.8 Flash TTS | 1268 | 2,261 | $16.5 /1M chars |
| 7 | Gemini 3.8 Flash-Lite TTS | 1240 | 2,232 | $11.0 /1M chars |
| 12 | Gemini 3.1 Flash TTS | 1204 | 3,575 | $18.3 /1M chars |
| 41 | Gemini 2.5 Flash Lite TTS | 1081 | 3,096 | $9.2 /1M chars |
| 55 | Gemini 2.5 Flash TTS (Dec 2025) | 1053 | 2,813 | $9.2 /1M chars |
| 69 | Gemini 2.5 Pro (Dec 2025) | 1030 | 2,404 | $18.3 /1M chars |
On that date the top two places belonged to ElevenLabs' Eleven v4 (1316) and Cartesia's Sonic 3.6 (1275).[28] The arena is updated continuously, so ranks and scores change; the 1,204 figure for 3.1 Flash TTS in September 2026 is below the 1,211 Google quoted at launch.[28][5]
Hume AI's Real-World VoiceEQ
Hume AI published an evaluation of preview versions of both 3.8 models on 24 September 2026, using three blind human raters per clip. Hume disclosed in the post that "Hume has a non-exclusive licensing agreement with Google" and said its test set is held out from model developers.[29] Hume found the biggest gain over Gemini 3.1 Flash TTS in "extended long-form stability," where scores rose "from 1.22 to 2.89 for Gemini 3.8 Flash TTS and 3.03 for Gemini 3.8 Flash-Lite TTS on a five-point scale." It ranked Flash TTS first in voice design but found speaker similarity in voice replication "less competitive" (Flash-Lite seventh and Flash eleventh of 13 models). On its Expressivity-Reliability Frontier Score the top four were all Gemini models: 3.8 Flash TTS (0.920), 3.8 Flash-Lite TTS (0.914), Gemini 2.5 Pro TTS (0.880) and Gemini 2.5 Flash TTS (0.861).[29] Android Authority noted that ElevenLabs still scored higher in Hume's individual voice quality and human-like variation categories.[27]
A licensing agreement between the two companies was first reported months before the 3.8 launch. PYMNTS, citing Wired, reported on 22 January 2026 that Google DeepMind had signed a licensing agreement with Hume AI and would hire Hume CEO Alan Cowen and several of its engineers.[30] Cowen is a co-author of Google's 3.8 TTS announcement, where he is listed as "Director, Research Science, on Behalf of the Gemini Audio Team."[15]
Third-party comparisons
Mistral AI's technical report on its Voxtral TTS model used Gemini 2.5 Flash TTS as a baseline in blind human evaluations of emotional delivery. The authors wrote that "Gemini 2.5 Flash TTS is the strongest model"; Voxtral's win rate against it was 35.4% with explicit emotion steering and 37.1% with implicit steering, compared with 51.0% and 55.4% against ElevenLabs v3.[31] The evaluation was run by a competitor on its own prompt set.
Press coverage has focused on control and voice customization. Engadget's I/O 2025 report highlighted the on-the-fly language switching and a whisper mode it called "a little creepy."[21] Android Authority's September 2026 report, headlined "Gemini can now clone your voice and perform scripts like an actor," focused on the 30-second voice replication and the consent and regional restrictions around it.[27]
Documentation inconsistencies
Google's own pages disagree on several figures, so any single number should be read with its source:
- Language counts. The 3.8 Flash TTS model page says 130 languages; the TTS guide says "over 130" for Flash and "over 100" for Flash-Lite; the Flash page's comparison table says 101 for Flash-Lite; the announcement post and the @GoogleAI post say "100+" or "more than 100."[12][8][15][19]
- Voice counts. The release notes mention "150+ prebuilt and custom voices," the guide mentions "hundreds of additional voices," and the blog post mentions "2,000+ production-ready voices."[3][8][15]
- 3.1 release date. The deprecations page says 26 February 2026; the release notes and blog say 15 April 2026.[6][3][5]
- Supported-model tables. The September 2026 TTS guide lists Gemini 2.5 Pro Preview TTS but not Gemini 2.5 Flash Preview TTS, while the pricing and deprecations pages still list the Flash model with no shutdown date.[8][13][6] Separately, on 18 September 2026 Google said it was "limiting access to the 2.5 models to users who have actively used them in the past"; that notice does not mention the TTS variants specifically.[3]
See also
- Gemini 3.8 Flash TTS
- Gemini 3.8 Flash-Lite TTS
- Gemini 3
- Gemini 2.5 Flash
- Gemini 2.5 Pro
- Google AI Studio
- NotebookLM
- SynthID
- Voice cloning
- Text-to-Speech Models
- Best AI Voice Generators
- ElevenLabs
- Hume AI
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13Google AI for Developers, "Speech generation (text-to-speech)," last updated 3 June 2025, Internet Archive snapshot of 14 June 2025. web.archive.org/...speech-generation
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10Google Cloud, "Gemini-TTS," Cloud Text-to-Speech documentation, accessed 30 September 2026. cloud.google.com/...gemini-tts
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23Google AI for Developers, "Release notes" (Gemini API changelog), entries dated 11 December 2024 to 22 September 2026, accessed 30 September 2026. ai.google.dev/...changelog
- ^1 ^2 ^3Tulsee Doshi, "Gemini 2.5: Our most intelligent models are getting even better," Google blog, 20 May 2025. blog.google/...google-gemini-updates-io-2025
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14Vilobh Meshram and Max Gubin, "Gemini 3.1 Flash TTS: the next generation of expressive AI speech," Google blog, 15 April 2026. blog.google/...gemini-3-1-flash-tts
- ^1 ^2 ^3 ^4 ^5 ^6Google AI for Developers, "Gemini deprecations," accessed 30 September 2026. ai.google.dev/...deprecations
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Google AI for Developers, "Speech generation (text-to-speech)," last updated 6 July 2026, Internet Archive snapshot of 10 July 2026. web.archive.org/...speech-generation
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12Google AI for Developers, "Text-to-speech generation (TTS)," last updated 24 September 2026. ai.google.dev/...speech-generation
- ^1 ^2 ^3 ^4 ^5Google AI for Developers, "Gemini 2.5 Flash text-to-speech," model page, last updated 18 August 2026. ai.google.dev/...gemini-2.5-flash-preview-tts
- ^1 ^2 ^3 ^4 ^5Google AI for Developers, "Gemini 2.5 Pro text-to-speech," model page, last updated 18 August 2026. ai.google.dev/...gemini-2.5-pro-preview-tts
- ^1 ^2 ^3 ^4 ^5Google AI for Developers, "Gemini 3.1 Flash TTS (text-to-speech) preview," model page, accessed 30 September 2026. ai.google.dev/...gemini-3.1-flash-tts-preview
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11Google AI for Developers, "Gemini 3.8 Flash TTS," model page, last updated 24 September 2026. ai.google.dev/...gemini-3.8-flash-tts
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10Google AI for Developers, "Gemini Developer API pricing," last updated 24 September 2026. ai.google.dev/...pricing
- ^1 ^2Google Cloud, "Text-to-Speech pricing," accessed 30 September 2026. cloud.google.com/...pricing
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14Leland Rechis and Alan Cowen, "Gemini 3.8 text-to-speech says hello," Google blog, 23 September 2026. blog.google/...gemini-3-8-text-to-speech
- ^1 ^2 ^3Sundar Pichai, Demis Hassabis and Koray Kavukcuoglu, "Google introduces Gemini 2.0: A new AI model for the agentic era," Google blog, 11 December 2024. blog.google/...google-gemini-ai-update-december-2024
- ^1 ^2 ^3 ^4 ^5 ^6Google Cloud, "Cloud TTS release notes," accessed 30 September 2026. cloud.google.com/...release-notes
- ^1 ^2 ^3 ^4Ivan Solovyev, "Improving Gemini Text-to-Speech models for better control and capabilities," Google blog, 10 December 2025. blog.google/...gemini-2-5-text-to-speech
- ^1 ^2Google AI (@GoogleAI), "We're launching Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS," X, 23 September 2026. x.com/...2102781694730285427
- ^1 ^2Google Cloud, "Chirp 3: HD voices," Cloud Text-to-Speech documentation, accessed 30 September 2026. cloud.google.com/...chirp3-hd
- ^1 ^2Will Shanklin, "Google's new text-to-speech can switch languages on the fly," Engadget, 20 May 2025. engadget.com/...tch-languages-on-the-fly-174528459
- ^1 ^2Shrestha Basu Mallick, Logan Kilpatrick, Alisa Fortin and Ivan Solovyev, "Gemini API I/O updates," Google Developers Blog, 23 May 2025. developers.googleblog.com/...gemini-api-io-updates
- ^1 ^2Bibo Xu and Tara Sainath, "Improved Gemini audio models for powerful voice interactions," Google blog, 12 December 2025. blog.google/...gemini-audio-model-updates
- ^1 ^2Google AI for Developers, "Gemini Developer API pricing," Internet Archive snapshot of 14 June 2025. web.archive.org/...pricing
- ^Biao Wang, "NotebookLM now lets you listen to a conversation about your sources," Google blog, 11 September 2024. blog.google/...notebooklm-audio-overviews
- ^Google Workspace Updates, "NotebookLM is now Gemini Notebook," 16 July 2026. workspaceupdates.googleblog.com/...gemini-notebook
- ^1 ^2 ^3Hillary Keverenge, "Gemini can now clone your voice and perform scripts like an actor," Android Authority, 24 September 2026. androidauthority.com/...speech-rolling-out-3714915
- ^1 ^2 ^3Artificial Analysis, "Text to Speech Leaderboard" (Provider Voice Arena), accessed 30 September 2026. artificialanalysis.ai/...leaderboard
- ^1 ^2Alice Baird, "Newly-released Google's Gemini 3.8 Flash TTS tops Hume's Real-World VoiceEQ leaderboard," Hume AI blog, 24 September 2026. hume.ai/...s-hume-s-real-world-voiceeq-leaderboard
- ^PYMNTS, "Google Recruits Hume CEO Alan Cowen to Bolster Voice AI Efforts," 22 January 2026 (reporting on Wired). pymnts.com/...-alan-cowen-bolster-voice-ai-efforts
- ^Alexander H. Liu et al. (Mistral AI), "Voxtral TTS," arXiv:2603.25551v2, 2026. arxiv.org/...2603.25551
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 4,509 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent verification V6 (xg13, 30 Sep 2026): ~95 claims vs Gemini API/Cloud TTS release notes, archived docs, Google blogs, AA leaderboard, Voxtral paper; 0 material, 5 minor fixed
Cite this page: AI Wiki. "Gemini TTS." aiwiki.ai, updated 30 Sept 2026, fact-checked 30 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/gemini_tts