Gemini 3.8 Flash TTS
Gemini 3.8 Flash TTS (model ID gemini-3.8-flash-tts) is a text-to-speech model from Google that turns text into single-speaker or two-speaker audio with line-by-line control over delivery. Google describes it as its "flagship creative text-to-speech model," built for "studio-grade voice fidelity, expressive acting, authentic regional accents, and rock-solid long-form multi-turn stability."[2] It was released as a generally available (GA) model in the Gemini API on 22 September 2026 and announced publicly on 23 September 2026, alongside a cheaper sibling, Gemini 3.8 Flash-Lite TTS.[3][1] Headline features are generative voice design (creating a new voice from a written description), consent-gated voice replication from a short audio sample, an expanded voice library, inline vocal-burst tags such as <laugh> and <sigh>, and listener backchannels written in pipes such as |mhm|.[1][4] The model supports 130 languages according to its developer documentation, while Google's announcement and social posts round this to "more than 100."[2][1][14]
Google's model lifecycle table lists the 3.8 Flash TTS and Flash-Lite TTS models as the recommended replacements for the earlier Gemini TTS previews, gemini-3.1-flash-tts-preview, gemini-2.5-flash-preview-tts and gemini-2.5-pro-preview-tts.[8] It is available to developers through the Gemini API and Google AI Studio, and to consumers in Gemini Notebook, the product formerly called NotebookLM; enterprise API access through Gemini Enterprise was listed as "coming soon" at launch.[1][23] On 28 September 2026 Google's main account promoted it as a tool that "turns voice generation into your full creative studio."[14]
Overview
| Attribute | Detail |
|---|---|
| Developer | Google (Gemini Audio team, Google DeepMind)[1][10] |
| Model ID | gemini-3.8-flash-tts[2] |
| Status | Generally available (GA) in the Gemini API[3][11] |
| API release | 22 September 2026 (Gemini API release notes and deprecations table)[3][8] |
| Public announcement | 23 September 2026 (Google blog, @GoogleAI)[1][13] |
| Base model | "Based on Gemini 3 Pro," per the model card[10] |
| Input / output | Text in, audio out[2] |
| Token limits | 8,192 input; 16,384 output (Gemini API serving limit); the model card lists "64K token output"[2][10] |
| Languages | 130 (developer docs); "more than 100 languages and dialects" (announcement)[2][1] |
| Voices | 30 prebuilt studio voices, an Extended Voice Library, Voice design and Voice replication[4] |
| Default output | WAV (24 kHz, mono, 16-bit PCM) for unary requests; raw PCM for streaming[4] |
| Launch price (paid tier) | $0.50 per 1M text input tokens and $9.00 per 1M audio output tokens through 31 December 2026; $1.00 and $18.00 from 1 January 2027[7] |
| Consumer product | Gemini Notebook[1][10] |
| Predecessor | gemini-3.1-flash-tts-preview (Gemini 3.1 Flash TTS)[2][8] |
Background
Google's Gemini API first offered dedicated speech-generation models with gemini-2.5-flash-preview-tts and gemini-2.5-pro-preview-tts, which the Gemini API model lifecycle table dates to 20 May 2025.[8] On 15 April 2026 Google introduced Gemini 3.1 Flash TTS, a preview model that added "audio tags" embedded in the input text, "Audio Profiles" and "Director's Notes" in AI Studio, support for more than 70 languages, and SynthID watermarking on all output.[25] Google said at the time that 3.1 Flash TTS had reached an Elo score of 1,211 on the Artificial Analysis TTS leaderboard.[25]
The 3.8 TTS models belong to what Google calls its "Gemini Audio family." The launch post says they follow Gemini 3.5 Live Translate, Gemini 3.5 Transcribe, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking.[1] Android Authority described the TTS release as the latest expansion of a Gemini 3.8 generation that began with the text model Gemini 3.8 Flash.[18] All four Gemini 3.8 Audio models share a single model card, published on 15 September 2026 and titled "Gemini 3.8 Audio (Live, Live Extended Thinking, Flash TTS, Flash-Lite TTS)."[10]
The launch post was written by Leland Rechis, Group Product Manager, and Alan Cowen, "Director, Research Science, on Behalf of the Gemini Audio Team."[1] Cowen founded the voice-AI company Hume AI; in January 2026 Wired reported that Google DeepMind had signed a licensing agreement with Hume and would hire Cowen and several of its engineers, with Andrew Ettinger taking over as Hume's CEO.[26] This matters for reading the launch benchmarks, several of which come from Hume (see Benchmarks below).
Release and availability
The Gemini API release notes for 22 September 2026 list "Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS generally available (GA)," together with a new Voices endpoint (/v1beta/voices) for Voice design, Voice replication and the Extended Voice Library.[3] Google's blog post "Gemini 3.8 text-to-speech says hello" went up on 23 September 2026, and @GoogleAI posted the launch the same day: "We're launching Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS."[1][13] The tweet positioned Flash TTS for "high-fidelity creative production like gaming, immersive audiobooks, and podcasts," and Flash-Lite TTS for near real-time voice agents, high-volume dubbing and bulk audio.[13]
Rollout by channel, as stated in the launch post and model card:
| Channel | Gemini 3.8 Flash TTS | Gemini 3.8 Flash-Lite TTS |
|---|---|---|
| Developers | Gemini API and Google AI Studio (launch post: "rolling out starting today," 23 Sep 2026) | Gemini API and Google AI Studio (launch post: "rolling out starting today," 23 Sep 2026) |
| Enterprises | "Coming soon via API in Gemini Enterprise" | "Coming soon via API in Gemini Enterprise" |
| Consumers | Gemini Notebook | Google Vids |
Sources: Google launch post and Gemini 3.8 Audio model card.[1][10]
Gemini Notebook is the name Google gave NotebookLM on 16 July 2026; Google's Workspace team said the product remained a standalone research tool and that the new name reflected "how it will evolve to do more across the Google ecosystem."[23][24] Android Authority's launch report headlined that "Gemini can now clone your voice," but its text places the consumer rollout in Gemini Notebook (Flash TTS) and Google Vids (Flash-Lite TTS), with developer access through AI Studio and the API.[18]
Google also opened a new audio playground in AI Studio, described as "a voice design workspace" where users can prompt new voices, replicate their own voice, and move them into a "dual-speaker screenplay editor."[1] Voice replication in AI Studio is not available in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland or India.[1]
Google named Agora, LiveKit, Pipecat and Vercel as developer platforms supporting the models through the Gemini API, and said Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang were integrating them for dubbing, localization and voice agents.[1] Third-party gateways listed the model on launch day: LiteLLM announced "day 0" support for both models through its /v1/audio/speech endpoint, and OpenRouter lists google/gemini-3.8-flash-tts with a release date of 23 September 2026.[21][22]
Capabilities
Script and direction model
The biggest change from Gemini 3.1 Flash TTS is how prompts are structured. Gemini 3.8 TTS "treats input text strictly as a verbatim transcript," so stage directions written inline, such as "Say cheerfully: Hello!", may be spoken aloud.[2] Instead, sustained delivery for a whole turn (emotion, pace, prosody, volume, whispering) goes into a structured speech_metadata.style field attached to each text part, and speaker labels go into speech_metadata.speaker.[4] In the Interactions API this is an annotation of type speech_metadata; in the GenerateContent API it is a speech_metadata object on each part.[2] Google's AI Studio developer guide summarizes the split as three parts: cast (the voice), direct (the style field) and speech (the transcript with inline sounds).[9]
Because style is attached per part, one speaker's line can be split into consecutive parts with different styles, for example moving from a whisper to a shout and then back to regular volume.[9] Google's guidance is to design a persona once in Voice design and keep per-turn style strings short or empty; the docs call long "Audio Profile" and "Director's Notes" blocks carried over from earlier models "the most common cause of voice drift."[4]
Inline tags, backchannels and pronunciation
Point-in-time events are written inline in angle brackets. The TTS guide lists recommended tags including <breath>, <chuckle>, <cough>, <gasp>, <giggle>, <laugh>, <sigh>, <sob>, <throat-clearing>, <whispers>, <yawn>, <short pause> and <long pause>, and advises keeping the tags in English even when the transcript is in another language.[4] Google advises using human vocalizations rather than non-vocal sound effects such as applause.[4]
In two-speaker dialogue, short listener reactions wrapped in pipes inside one speaker's turn, such as |mhm|, |oh really?| or |absolutely|, are voiced by the other speaker as backchannels or overlapping speech without starting a new turn.[4] The AI Studio guide says one to three words in pipes make "Speaker B" voice the reaction "simultaneously in the background while Speaker A keeps talking," and that multiple pipe segments can simulate two people talking over each other; the TTS guide notes that overlapping speech "works best with gemini-3.8-flash-tts."[9][4] Other controls include capitalized words for emphasis, punctuation and ellipses for hesitation, and International Phonetic Alphabet spellings inside slashes (for example /niːv/) for difficult names.[4][9]
Multi-speaker scenes
A single request can stage a two-speaker scene by configuring two speakers in speech_config.speakers and setting "mode": "conversational" for natural turn-taking; every turn must name its speaker.[4] The launch post calls this "native two-speaker scene staging."[1] There is a limit: single-request multi-speaker generation "supports up to 2 speakers using prebuilt voices." Dialogue between designed or replicated voices must be synthesized one turn at a time and stitched together.[4] The New Stack singled this out as the one limitation of native two-speaker generation worth noting.[19]
Voices
The API offers four ways to choose a voice:[4]
| Source | What it is | Notes |
|---|---|---|
| Prebuilt studio voices | 30 named voices (for example Zephyr, Puck, Kore, Charon, Algenib, Sulafat) | The same 30-voice set the earlier Gemini TTS models used[4][22] |
| Extended Voice Library | Additional catalog voices, filterable by language, region, accent, gender, pitch, persona and use context via GET /v1beta/voices | See the note on counts below |
| Voice design | A new persistent voice generated from a natural-language description (type="prompted"), returned as a voice_... ID with a WAV preview | Stored per project[5] |
| Voice replication | A voice copied from reference audio plus a recorded consent statement (type="replicated") | Stored voice_... ID or stateless voicekey_...[6] |
Google's pages give different sizes for the library. The launch post says "2,000+ production-ready voices" and the AI Studio guide says Google "expanded our library of voices from 30 to 2,000+"; Google DeepMind's product page says "over a thousand" in one place and "thousands" in another; the API speech guide describes "hundreds of additional voices" beyond the 30 prebuilt ones; and the 22 September release note mentions querying "150+ prebuilt and custom voices."[1][9][11][4][3] The AI Studio guide gives regional examples such as Mexican Spanish, Brazilian Portuguese, Egyptian Arabic and Scots English.[9]
Stored custom voices (designed or replicated) are limited to 200 per project and expire after one year; stateless replicated voice keys expire after seven days.[4][6] Google's DeepMind page marks generative voice design as "Gemini 3.8 Flash TTS only," but the Gemini API documentation lists Voice design as supported on both Flash TTS and Flash-Lite TTS.[11][5] A "voice remixing" feature, for adjusting a library voice's timbre, pitch, pace and accent by prompt, was announced as "coming soon."[1]
Languages
Flash TTS detects the input language automatically. Its model page lists 130 languages in a comparison table and says it "supports over 130 languages," from Acehnese and Afrikaans to Uyghur and Vietnamese, including several script variants (for example Chinese in Hans and Hant scripts and Standard Arabic in Arabic and Latin scripts).[2] The AI Studio guide also says 130 languages.[9] The launch post, the 23 September tweet and the 28 September promotional tweet use "100+" or "more than 100 languages and dialects."[1][13][14] For comparison, Flash-Lite TTS is listed at 101 languages.[2]
Long-form stability
Google says Flash TTS maintains "consistent voice identity, timbre, volume, and acoustic room tone across extended dialogues and multi-minute narrations without voice drift," and the launch post claims high quality "across hours of continuous audio with minimal speaker drift."[2][1] Per request, however, input is capped at 8,192 tokens, so LiteLLM advises splitting long scripts such as audiobook chapters across calls.[2][21]
Technical specifications
| Property | Value |
|---|---|
| Model code | gemini-3.8-flash-tts[2] |
| Input | Text only; 8,192-token limit[2] |
| Output | Audio only; 16,384 tokens (Gemini API serving limit); model card: "64K token output"[2][10] |
| Audio token rate | 25 tokens per second of audio[7] |
| Unary output | audio/wav with RIFF header, 24 kHz, mono, 16-bit signed little-endian PCM[4] |
| Streaming output | Headerless audio/l16 PCM chunks, 24 kHz, mono[4] |
| Other formats | audio/mulaw and audio/alaw (G.711 telephony); sample_rate can be set, for example 24000, 16000 or 8000 Hz[4] |
| Supported | Audio generation, context caching, Batch API, Flex inference, Priority inference[2] |
| Not supported | Function calling, code execution, search or Maps grounding, structured outputs, thinking, URL context, file search, image generation, Live API[2] |
| Knowledge cutoff | January 2025 (model card)[10] |
The default WAV output is itself a migration change: gemini-3.1-flash-tts-preview and earlier TTS models returned headerless raw PCM, so code that wrapped the bytes in a WAV header must drop that step.[2]
Pricing
Google's pricing page shows a promotional rate through 31 December 2026 that doubles from 1 January 2027. Prices are per 1 million tokens in US dollars on the paid tier; audio output tokens correspond to 25 tokens per second of audio.[7]
| Tier | Input (text) through 31 Dec 2026 | Output (audio) through 31 Dec 2026 | Input from 1 Jan 2027 | Output from 1 Jan 2027 |
|---|---|---|---|---|
| Standard | $0.50 | $9.00 | $1.00 | $18.00 |
| Batch | $0.25 | $4.50 | $0.50 | $9.00 |
| Flex | $0.25 | $4.50 | $0.50 | $9.00 |
| Priority | $0.90 | $16.20 | $1.80 | $32.40 |
Source: Gemini API pricing page.[7]
Google states the standard output rate as equivalent to $0.00225 per 10 seconds of audio through 2026 and $0.0045 from January 2027.[7] At 25 tokens per second, an hour of output audio is 90,000 tokens, or about $0.81 at the promotional standard rate. The free tier is "free of charge" for standard requests, with content used to improve Google's products; Batch and Flex are not available on the free tier.[7] Context caching is priced separately.[7] By comparison, gemini-3.1-flash-tts-preview is listed at $1.00 per 1M text input tokens and $20.00 per 1M audio output tokens.[7]
In its launch-day post, Artificial Analysis converted prices into per-character terms and put Flash TTS at $32.98 per 1M characters, against $22.07 for Flash-Lite TTS and $18.31 for Gemini 3.1 Flash TTS, calling both new models more expensive than 3.1 Flash TTS but "substantially cheaper than Eleven v3 at $100/1M characters."[17] Its live leaderboard later listed lower figures: on 30 September 2026 it showed $16.5 per 1M characters for Flash TTS and $11.0 for Flash-Lite TTS, both below the $18.3 it listed for Gemini 3.1 Flash TTS.[28]
Benchmarks
All benchmark figures below are reported by the organization named, on the dates given (23 to 30 September 2026). Google's own launch results rely heavily on Hume AI's evaluations; Hume discloses that it "has a non-exclusive licensing agreement with Google" and says its held-out test set is not shared with model developers.[15]
Google-reported results
Google's launch post says Flash TTS took "the #1 overall spot on Hume AI's Voice Design Benchmark (71.4)" and led in accent modeling (60.8), and that Flash TTS and Flash-Lite TTS placed first and second on "Hume AI's Overall Quality Index."[1] Google DeepMind's model page gives the underlying tables:[11]
| Hume Voice Design benchmark | Gemini 3.8 Flash TTS | ElevenLabs Voice Design v3 | Inworld Voice Design |
|---|---|---|---|
| Overall English | 71.4 | 70.8 | 69.8 |
| Multilingual | 3.82 | 3.65 | 3.57 |
| Accents | 60.8 | 45.4 | 35.8 |
| Voice Qualities (single tag) | 74.6 | 76.6 | 76.3 |
| Hume TTS quality (selected) | 3.8 Flash TTS | 3.8 Flash-Lite TTS | 3.1 Flash TTS | ElevenLabs v3 | Cartesia Sonic 3.6 | OpenAI gpt-4o-mini-tts |
|---|---|---|---|---|---|---|
| Overall reliability x expressiveness | 0.920 | 0.914 | 0.783 | 0.706 | 0.840 | 0.740 |
| Human-like variation | 4.58 | 4.51 | 3.95 | 5.00 | 3.40 | 4.22 |
| Multispeaker | 4.14 | 4.10 | 3.60 | 3.85 | n/a | n/a |
| Style tag control (single tag) | 4.34 | 4.32 | 4.31 | 3.89 | 3.37 | n/a |
In the same tables, ElevenLabs scored higher on voice qualities and human-like variation, a point Android Authority also noted.[11][18] Google also reported blind pairwise preference Elo scores from "Voice Arena," which its methodology document describes as "public crowdsourced, blind pairwise preference evaluations."[12] Flash TTS led Google's table in Japanese (1,232), Arabic MSA (1,204), Mexican Spanish (1,152) and Hindi (1,106), while Flash-Lite TTS led in English (1,087), Brazilian Portuguese (1,134) and Vietnamese (1,156). In English, Flash TTS (1,061) placed below both Flash-Lite TTS and Cartesia Sonic 3.6 (1,068).[11]
Hume AI's evaluation
Hume published its own write-up on 24 September 2026, titled "Newly-released Google's Gemini 3.8 Flash TTS tops Hume's Real-World VoiceEQ leaderboard." It says it tested "preview versions" of both models using three blind human raters per clip recruited through its Human Feedback API.[15] Its findings were mixed:[15]
- On Hume's new Expressivity-Reliability Frontier Score, Flash TTS ranked first (0.920) and Flash-Lite TTS second (0.914), followed by Gemini 2.5 Pro TTS (0.880), Gemini 2.5 Flash TTS (0.861), Cartesia Sonic 3.6 (0.840) and Grok TTS (0.792). Google's comparison table does not show the two Gemini 2.5 models.
- Extended long-form stability rose from 1.22 for Gemini 3.1 Flash TTS to 2.89 for Flash TTS and 3.03 for Flash-Lite TTS on a five-point scale.
- Flash TTS ranked first overall in Voice Design, with its clearest accent advantages in UK regional accents and World Englishes.
- On voice replication, speaker similarity was weaker: Flash TTS ranked eleventh of 13 models at 3.53 out of 5, and Flash-Lite TTS seventh at 3.68, against a field average of 3.69.
- Specific controls lagged: Flash TTS's young-adult voice accuracy was 34.0% against 68.1% for ElevenLabs v3, and its volume-control pass rate was 56.9% against 88.9% for Inworld Voice Design.
A second Hume post on 29 September 2026 evaluated two-speaker dialogue across 48 scripts. Flash TTS scored 4.11 overall, just below Gemini 2.5 Pro TTS (4.12) and ahead of Flash-Lite TTS (4.10) and ElevenLabs v3 (3.90). Its speaker separation (4.17) was the best of the Gemini models tested, but same-gender pairs remained harder: 4.92 for male-female pairs against 3.42 for male-male pairs.[16]
Artificial Analysis
Artificial Analysis posted results on launch day. Flash TTS debuted at #2 on its Provider Voice Arena with an Elo of 1,263, behind Cartesia's Sonic 3.6 (1,272) and ahead of Qwen-Audio-3.0-TTS-Plus (1,260), and 64 Elo points above Gemini 3.1 Flash TTS (1,199).[17] It ranked #1 on Artificial Analysis's Pronunciation Robustness Benchmark at 89.5%, and ranked #2 in Japanese and Portuguese among its nine multilingual Controlled Voice Arenas. Artificial Analysis measured generation speed at 44.1 characters per second, about 2.7 times faster than real time.[17] The arena is updated continuously: when its leaderboard was checked on 30 September 2026, Flash TTS stood third with an Elo of 1,268, behind ElevenLabs' Eleven v4 (1,316) and Cartesia's Sonic 3.6 (1,275), and Flash-Lite TTS was seventh at 1,240.[28]
Safety and consent
Voice replication requires two recordings from the same person: a 10 to 30 second reference clip and a consent clip in which the speaker reads a fixed statement ("I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model") in one of 30 supported language locales.[6] Google's launch post describes this as "consent verification: users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created."[1] The launch post describes replication from "just a 30-second audio sample," while the API documentation specifies a reference clip of 10 to 30 seconds, making 30 seconds the documented maximum rather than a minimum.[1][6] Developers who do not want Google to store a voice profile can request a stateless, encrypted voicekey_... that the application keeps and that expires after seven days.[6]
Google says "every audio clip generated by our Gemini Audio models is watermarked with SynthID," and that voice replication is also backed by C2PA content credentials.[1] These safeguards address voice cloning and deepfake risks; The New Stack contrasted Google's self-serve replication with OpenAI's custom voices, which it said require going through sales and are limited to 20 voices per organization.[19]
The Gemini 3.8 Audio model card states that Google relied on its Frontier Safety Framework evaluations of Gemini 3.7 Flash, which did not reach any tracked or critical capability levels, and that the 3.8 Audio models "do not have meaningful new capabilities or material increases in performance" compared with Gemini 3.7 Flash. For risks, mitigations and acceptable use, the card refers readers to the Gemini 3 Pro model card.[10]
Comparison with Flash-Lite TTS
The two 3.8 TTS models share an API schema, so switching between them is a one-parameter change.[4] Google recommends Flash-Lite TTS as the default drop-in replacement for gemini-3.1-flash-tts-preview, and Flash TTS when an application needs "studio-grade acting nuance, multi-speaker dialogue with backchanneling, dialect acting, or long-form narrative stability."[9]
| Gemini 3.8 Flash TTS | Gemini 3.8 Flash-Lite TTS | |
|---|---|---|
| Model ID | gemini-3.8-flash-tts | gemini-3.8-flash-lite-tts |
| Primary strength (Google) | "Maximum voice fidelity, acting nuance, and dialect coverage" | "High throughput, low latency, and cost efficiency" |
| Best uses (Google) | Audiobooks, studio narration, complex multi-speaker dialogue, heavy vocal-burst acting, regional dialects | High-volume production, real-time voice agent cascades, read-aloud features, voice replication |
| Languages | 130 | 101 |
| Standard price through 2026 (per 1M tokens) | $0.50 in / $9.00 out | $0.50 in / $6.00 out |
| Consumer product | Gemini Notebook | Google Vids |
| Artificial Analysis Provider Voice Arena (23 Sep 2026) | #2, Elo 1,263 | #6, Elo 1,236 |
Sources: Gemini API model pages, pricing page, launch post and Artificial Analysis.[2][27][7][1][17]
Reception
Coverage focused on voice replication and control. Android Authority called the release a step beyond "paste text, get audio," noting line-by-line direction, two-speaker staging and the consent and watermark safeguards.[18] The New Stack framed it as making custom voices "self-serve" and called the verbatim-transcript rule "a breaking change for anyone who embedded stage directions in prompts to the 3.1 preview model."[19] Fone Arena summarized the launch specifications and regional limits.[20]
Launch partners quoted by Google DeepMind included 99.co, whose CEO Darius Cheung cited support for "local languages and accents, from Singlish to Indonesian," and Katsuyo developer Luke Pane, who said other models lacked "critical aspects of Japanese text-to-speech such as pitch accent and correct reading of kanji."[11] Hume, while ranking the model first on its overall scores, said "speaker similarity and the precision of some voice controls remain areas for improvement."[15]
See also
- Gemini 3.8 Flash-Lite TTS
- Gemini TTS
- Gemini 3.8 Flash
- Text-to-speech
- Voice cloning
- Hume AI
- ElevenLabs v3
- Cartesia
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25Leland Rechis and Alan Cowen, "Gemini 3.8 text-to-speech says hello," Google blog (The Keyword), 23 September 2026. blog.google/...gemini-3-8-text-to-speech
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20Google AI for Developers, "Gemini 3.8 Flash TTS" (model page), last updated 24 September 2026. ai.google.dev/...gemini-3.8-flash-tts
- ^1 ^2 ^3 ^4 ^5Google AI for Developers, "Release notes" (Gemini API changelog), entry dated 22 September 2026. ai.google.dev/...changelog
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20Google AI for Developers, "Text-to-speech generation (TTS)," last updated 24 September 2026. ai.google.dev/...speech-generation
- ^1 ^2Google AI for Developers, "Voice design," last updated 24 September 2026. ai.google.dev/...voice-design
- ^1 ^2 ^3 ^4 ^5Google AI for Developers, "Voice replication," last updated 24 September 2026. ai.google.dev/...voice-replication
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Google AI for Developers, "Gemini Developer API pricing," last updated 24 September 2026. ai.google.dev/...pricing
- ^1 ^2 ^3 ^4Google AI for Developers, "Deprecations" (model release and shutdown dates). ai.google.dev/...deprecations
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8fofr, "Gemini 3.8 Flash TTS: Developer Guide," Google AI Studio. aistudio.google.com/...8-flash-tts-developer-guide
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Google DeepMind, "Gemini 3.8 Audio - Model Card," published 15 September 2026. deepmind.google/...gemini-3-8-audio
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Google DeepMind, "Gemini Audio - Speech generation." deepmind.google/...speech-generation
- ^Google DeepMind, "Gemini 3.8 Audio (Flash TTS, Flash-Lite TTS) Model evaluation: Approach, methodology & results." deepmind.google/...gemini-3-8-tts
- ^1 ^2 ^3 ^4Google AI (@GoogleAI), "We're launching Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS," X, 23 September 2026. x.com/...2102781694730285427
- ^1 ^2 ^3Google (@Google), "Gemini 3.8 Flash TTS turns voice generation into your full creative studio," X, 28 September 2026. x.com/...2104657688718434638
- ^1 ^2 ^3 ^4Alice Baird, "Newly-released Google's Gemini 3.8 Flash TTS tops Hume's Real-World VoiceEQ leaderboard," Hume AI blog, 24 September 2026. hume.ai/...s-hume-s-real-world-voiceeq-leaderboard
- ^Sharath Rao, Kimberly Lo and Alice Baird, "Evaluating Google's multi-speaker TTS: A case study in why private evaluations matter," Hume AI blog, 29 September 2026. hume.ai/...evaluating-multi-speaker-tts
- ^1 ^2 ^3 ^4Artificial Analysis (@ArtificialAnlys), "Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS," X, 23 September 2026. x.com/...2102784197853380647
- ^1 ^2 ^3 ^4Hillary Keverenge, "Gemini can now clone your voice and perform scripts like an actor," Android Authority, 24 September 2026. androidauthority.com/...speech-rolling-out-3714915
- ^1 ^2 ^3Amanda Caswell, "OpenAI makes you call sales for a custom voice. Google just made it self-serve.," The New Stack, 24 September 2026. thenewstack.io/gemini-tts-voice-replication-api
- ^Fone Arena, "Google rolls out Gemini 3.8 Flash TTS and Flash-Lite TTS with custom voice creation, 2000+ voices, 100+ language support," 24 September 2026. fonearena.com/...mini-3-8-flash-tts-flash-lite-tts
- ^1 ^2LiteLLM, "Day 0 support: Gemini 3.8 Flash TTS and Flash-Lite TTS," 23 September 2026. docs.litellm.ai/...gemini_3_8_flash_tts
- ^1 ^2OpenRouter, "Google: Gemini 3.8 Flash TTS." openrouter.ai/...gemini-3.8-flash-tts
- ^1 ^2Google Workspace Updates, "NotebookLM is now Gemini Notebook," 16 July 2026. workspaceupdates.googleblog.com/...gemini-notebook
- ^TechCrunch, "Google continues its renaming streak by turning NotebookLM to Gemini Notebook," 16 July 2026. techcrunch.com/...ng-notebooklm-to-gemini-notebook
- ^1 ^2Google, "Gemini 3.1 Flash TTS: the next generation of expressive AI speech," Google blog, 15 April 2026. blog.google/...gemini-3-1-flash-tts
- ^PYMNTS, "Google Recruits Hume CEO Alan Cowen to Bolster Voice AI Efforts," 22 January 2026 (reporting on Wired). pymnts.com/...-alan-cowen-bolster-voice-ai-efforts
- ^Google AI for Developers, "Gemini 3.8 Flash-Lite TTS" (model page). ai.google.dev/...gemini-3.8-flash-lite-tts
- ^1 ^2Artificial Analysis, "Text to Speech Leaderboard" (Provider Voice Arena), accessed 30 September 2026. artificialanalysis.ai/...leaderboard
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 4,352 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent verification V3 (xg13, 30 Sep 2026): ~165 claims vs 32 sources (Gemini API docs/release notes/pricing, DeepMind model card, Google blog, Hume, Artificial Analysis); 1 material (stale AA price comparison) + 8 minor fixed
Cite this page: AI Wiki. "Gemini 3.8 Flash TTS." aiwiki.ai, updated 30 Sept 2026, fact-checked 30 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/gemini_3_8_flash_tts