Gemini 3.8 Flash-Lite TTS
Gemini 3.8 Flash-Lite TTS (model ID gemini-3.8-flash-lite-tts) is a text-to-speech model from Google that converts text into single-speaker or two-speaker audio. Google's developer documentation calls it a "fast, cost-efficient workhorse text-to-speech model built to replace gemini-3.1-flash-tts-preview for high-throughput production workloads."[1] It was released as a generally available (GA) model in the Gemini API on 22 September 2026 and announced publicly on 23 September 2026, together with its higher-fidelity sibling Gemini 3.8 Flash TTS.[3][2][6] Google positions Flash-Lite TTS for near real-time voice agents, high-volume dubbing and bulk audio generation, while reserving Flash TTS for studio-grade creative work such as audiobooks, studio narration and complex multi-speaker dialogue.[6][1] It uses the same API schema as Flash TTS, supports 101 languages according to its model page, and costs $0.50 per million text input tokens and $6.00 per million audio output tokens under launch pricing that runs through 31 December 2026.[1][5]
Background
Google has offered text-to-speech through the Gemini API since its Gemini 2.5 preview TTS models (gemini-2.5-flash-preview-tts and gemini-2.5-pro-preview-tts, released 20 May 2025), followed by Gemini 3.1 Flash TTS Preview, launched on 15 April 2026.[3] The Gemini API deprecation table now lists the 3.8 TTS models as the recommended replacements for all three of those previews.[13] See Gemini TTS for the wider model line.
The 3.8 TTS pair belongs to what Google calls the Gemini Audio family. Google's launch post says the two models follow Gemini 3.5 Live Translate, Gemini 3.5 Transcribe, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking.[2] All four Gemini 3.8 Audio models share one model card, which states that "Gemini 3.8 Audio is based on Gemini 3 Pro" and refers readers to the Gemini 3 Pro card for architecture, training data and hardware details.[7]
Release and availability
The Gemini API release notes entry for 22 September 2026 lists "Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS generally available (GA)" and describes Flash-Lite TTS as built "for high-throughput production and real-time voice agent cascades."[3] The same entry introduced the Voices endpoint (/v1beta/voices) used for voice design, voice replication and the Extended Voice Library.[3] Google's blog post "Gemini 3.8 text-to-speech says hello," written by group product manager Leland Rechis and research science director Alan Cowen "on Behalf of the Gemini Audio Team," is dated 23 September 2026, and the @GoogleAI account announced both models that day at 15:26 UTC.[2][6] The Gemini API models overview labels the model "New" and "Stable," and the deprecation table gives it a release date of 22 September 2026 with no shutdown date announced.[14][13]
| Channel | Status at launch (23 September 2026) |
|---|---|
| Gemini API | GA in the release notes from 22 September 2026; launch post: "rolling out starting today"[3][2] |
| Google AI Studio | Rolling out from launch day[2] |
| Google Vids (video creation app) | Rolling out from launch day[2][7] |
| Gemini Enterprise (API) | "Coming soon"[2] |
The consumer channel differs between the siblings: Google put Flash-Lite TTS into Google Vids and Flash TTS into Gemini Notebook.[2][7] Android Authority's launch report made the same split.[18]
Third-party model gateways listed Flash-Lite TTS on launch day. LiteLLM announced "day 0" support for both 3.8 TTS models through its /v1/audio/speech endpoint on 23 September 2026.[20] OpenRouter lists google/gemini-3.8-flash-lite-tts with a release date of 23 September 2026 and a single upstream provider, Google AI Studio.[21] The Vercel AI Gateway lists it as google/gemini-3.8-flash-lite-tts with a release date of 23 September 2026, and fal.ai lists a google/gemini-3.8-flash-lite-tts endpoint.[22][23] These listings were checked on 30 September 2026.
Positioning and intended uses
Google's launch tweet framed the choice between the two models as a question: "Want the AI to automatically adjust its tone and pacing on the fly for near real-time voice agents? This is your engine! Built for cost-efficient scale, high-volume dubbing, and bulk audio creation."[6] The blog post describes Flash-Lite TTS as "Built for high-volume, cost-efficient scale. Optimized for high-volume dubbing, audio content creation, and expressive voice agents with fine-grained control over tone, pacing, and expressive nuance."[2] Google DeepMind's product page calls it "Best for everyday audio generation, like video dubbing or podcast creation," and says it "Balances efficiency with expressiveness to power large content libraries and high-traffic apps."[8]
The developer documentation is more specific. It lists the model's primary strength as "High throughput, low latency, and cost efficiency" and its best uses as "High-volume production, real-time voice agent cascades, read-aloud features, voice replication, everyday single-speaker generation."[1] A "voice agent cascade" here means the common pipeline in which a text model writes a reply and a separate TTS step speaks it; The New Stack described Flash-Lite as tuned for "cascaded voice agents that pair a text model with a separate speech step."[19] Google's guide for this use case recommends one TTS call per turn as text chunks arrive from the language model, a stored voice to carry speaker identity, and an empty or short constant style string.[4] TTS through the Gemini API is separate from the Live API, which Google recommends for interactive, unstructured audio conversations.[4]
Google's AI Studio developer guide calls Flash-Lite TTS "our recommended drop-in upgrade from gemini-3.1-flash-tts-preview" and says: "As a general rule, use Gemini 3.8 Flash-Lite TTS as your default drop-in replacement for gemini-3.1-flash-tts-preview."[10]
Technical specifications
| Property | Gemini 3.8 Flash-Lite TTS |
|---|---|
| Model code | gemini-3.8-flash-lite-tts[1] |
| Input | Text only[1] |
| Output | Audio only[1] |
| Input token limit | 8,192[1] |
| Output token limit | 16,384 (described as the "Gemini API serving limit"); the model card gives 64K output tokens[1][7] |
| Supported | Audio generation, caching[1] |
| Not supported | Code execution, file search, function calling, grounding with Google Maps, image generation, Live API, search grounding, structured outputs, thinking, URL context[1] |
| Consumption options | Batch API, Flex inference, Priority inference[1] |
| Default audio (unary requests) | WAV with RIFF header, 24 kHz, mono, 16-bit signed little-endian PCM[4] |
| Default audio (streaming) | Headerless raw linear PCM (audio/l16), 24 kHz, mono[4] |
| Other formats | Mu-law and A-law (G.711); sample rate selectable, for example 24,000, 16,000 or 8,000 Hz[4] |
| Base model | Gemini 3 Pro, per the Gemini 3.8 Audio model card[7] |
| Latest update | September 2026[1] |
The 8-bit mu-law and A-law options are the encodings used in telephone systems, which the documentation describes as commonly used in North American and Japanese telephony and IVR systems (mu-law) and in European and international telephony (A-law).[4] Unary requests return a complete WAV file by default, unlike gemini-3.1-flash-tts-preview and earlier Gemini TTS models, which returned headerless PCM.[1]
Capabilities
Prompting schema
Flash-Lite TTS "Shares the exact same API schema and structured prompting format as gemini-3.8-flash-tts, allowing you to switch models with a single parameter change."[1] Both 3.8 models treat the input text "strictly as a verbatim transcript," so stage directions typed into the text (for example "Say cheerfully: Hello!") may be read aloud.[1] Instead, sustained delivery instructions and speaker labels go into a speech_metadata annotation with style and speaker fields, and momentary vocal events go inline in angle brackets, such as <laugh>, <sigh>, <cough>, <breath> or <short pause>.[1][4] Google's recommended tag list also includes <gasp>, <chuckle>, <whispers>, <yawn> and <long pause>, and it advises keeping these tags in English even when the transcript is in another language.[4] In two-speaker dialogue, listener reactions wrapped in pipes, such as |mhm|, produce backchannels without a separate turn; the documentation says overlapping, interleaved speech built this way "works best with gemini-3.8-flash-tts."[4]
Voices
The model supports the four voice sources of the 3.8 TTS schema: 30 prebuilt studio voices (such as Zephyr, Puck, Kore and Charon), an Extended Voice Library queried through GET /v1beta/voices, custom voices created with Voice design, and voices cloned with Voice replication.[4][1] The Gemini API documentation states that both 3.8 TTS models support Voice design and Voice replication.[11][12] Stored custom voices are limited to 200 per project with a one-year retention period, while stateless replicated voice keys (voicekey_...) expire after seven days.[4]
Google's pages describe the size of the voice library differently: the release notes say "150+ prebuilt and custom voices," the TTS guide says "Hundreds of additional voices," DeepMind's page says "over a thousand" and "thousands," and the launch post and AI Studio guide say "2,000+."[3][4][8][2][10] Google DeepMind's speech-generation page also marks generative voice design as "Gemini 3.8 Flash TTS only," which conflicts with the Gemini API documentation that lists Voice design as supported on Flash-Lite TTS.[8][11] The launch blog likewise attributes voice design to Flash TTS.[2]
Multi-speaker audio
The Gemini API's supported-models table marks both single-speaker and multi-speaker generation as supported for Flash-Lite TTS.[4] A single request can include up to two speakers using prebuilt voices; to combine designed or replicated voices in a multi-character scene, each speaker's turn must be generated separately.[4] Google nonetheless steers "complex multi-speaker dialogue" toward Flash TTS and lists "everyday single-speaker generation" among Flash-Lite's best uses.[1]
Languages
The model detects the input language automatically.[1] Its model page gives "101 languages" in the comparison table and says it "supports over 100 languages," and the language table under that heading has 103 rows, some of which are script variants of one language (for example Chinese in Hans and Hant scripts, Standard Arabic in Arabic and Latin scripts, and Kashmiri in Arabic and Devanagari scripts).[1] Flash TTS is listed at 130 languages.[1][24] The TTS guide's side-by-side language table marks 29 entries as supported by Flash TTS but not Flash-Lite TTS, including Finnish, Swedish, Thai, Lithuanian, Slovenian, Swahili, Somali, Igbo, Burmese, Luxembourgish and Uyghur; no language is listed for Flash-Lite alone.[4] Google's launch post rounds the pair's coverage to "over 100 languages."[2]
Pricing
Google's pricing page sets launch prices "through December 31, 2026" and higher prices "starting January 1, 2027." Audio output tokens correspond to 25 tokens per second of audio.[5]
| Tier (paid, per 1M tokens, USD) | Text input, through 31 Dec 2026 | Audio output, through 31 Dec 2026 | Text input, from 1 Jan 2027 | Audio output, from 1 Jan 2027 |
|---|---|---|---|---|
| Standard | $0.50 | $6.00 (about $0.0015 per 10 s) | $1.00 | $12.00 (about $0.003 per 10 s) |
| Batch | $0.25 | $3.00 | $0.50 | $6.00 |
| Flex | $0.25 | $3.00 | $0.50 | $6.00 |
| Priority | $0.90 | $10.80 | $1.80 | $21.60 |
Source: Gemini Developer API pricing page.[5]
The Standard and Priority tiers are free of charge on the free tier; Batch and Flex are not available on the free tier.[5] Context caching is priced separately, for example $0.125 per million input tokens on Standard through 2026.[5] For comparison, Flash TTS costs $0.50 input and $9.00 output per million tokens on Standard through 2026, and the Gemini 3.1 Flash TTS Preview is listed at $1.00 input and $20.00 output.[5] LiteLLM summarized the same launch rates and noted that "Batch and Flex run at half these rates and Priority at 1.8x."[20]
In its launch-day post on 23 September 2026, Artificial Analysis converted prices to a per-character basis and put Flash-Lite TTS at $22.07 per million characters, against $32.98 for Flash TTS and $18.31 for Gemini 3.1 Flash TTS; by that calculation both new models cost more per character than 3.1 Flash TTS but much less than ElevenLabs' Eleven v3 at $100 per million characters.[17] Its live leaderboard later listed lower figures: on 30 September 2026 it showed $11.0 per million characters for Flash-Lite TTS and $16.5 for Flash TTS, both below the $18.3 it listed for Gemini 3.1 Flash TTS.[26] Google's per-token prices likewise put both 3.8 models below the 3.1 preview ($6.00 and $9.00 against $20.00 per million audio output tokens).[5]
Performance and benchmarks
Google-published results
Google's evaluation methodology document says all evaluations used "production checkpoints with default sampling settings and single-attempt generation (pass@1)" and drew on two external sources: Hume AI's text-to-speech benchmarks, which use human listener panels, and Voice Arena, described as "public crowdsourced, blind pairwise preference evaluations."[9] The document says the evaluations benchmark "acoustic quality, instruction following, latency, and multilingual breadth," but its published results pages contain only the Hume and Voice Arena tables.[9]
On Hume's text-to-speech quality benchmark as reported by Google, Flash-Lite TTS placed second behind Flash TTS on the overall score, and Google's launch post says the pair secured "the #1 and #2 spots respectively on Hume AI's Overall Quality Index."[9][2]
| Hume TTS quality benchmark (Google's table, selected columns, September 2026) | 3.8 Flash TTS | 3.8 Flash-Lite TTS | 3.1 Flash TTS | ElevenLabs v3 | Cartesia Sonic 3.6 | OpenAI gpt-4o-mini-tts |
|---|---|---|---|---|---|---|
| Overall (reliability x expressiveness) | 0.920 | 0.914 | 0.783 | 0.706 | 0.840 | 0.740 |
| Human-like variation | 4.58 | 4.51 | 3.95 | 5.00 | 3.40 | 4.22 |
| Multispeaker | 4.14 | 4.10 | 3.60 | 3.85 | n/a | n/a |
| Style tag control (single tag) | 4.34 | 4.32 | 4.31 | 3.89 | 3.37 | n/a |
Source: Google DeepMind evaluation methodology document and speech-generation page.[9][8]
In Google's Voice Arena table, Flash-Lite TTS scored higher than Flash TTS in three of the seven languages shown: English (1,087 against 1,061), Brazilian Portuguese (1,134 against 1,104) and Vietnamese (1,156 against 1,135). In English it also led every other model in the table, including Cartesia Sonic 3.6 (1,068).[9] Its other scores were 1,152 in Japanese, 1,181 in Arabic (MSA), 1,076 in Hindi and 1,146 in Mexican Spanish; in Hindi it placed behind Cartesia Sonic 3.6 (1,104), Flash TTS (1,106) and Gemini 3.1 Flash TTS (1,086).[9] The table compares only six models, so it does not show a full arena ranking.
Hume AI's evaluations
Hume published its own write-up on 24 September 2026, saying it had independently tested "preview versions" of both models with three blind human raters per clip. The post discloses that "Hume has a non-exclusive licensing agreement with Google" and says its held-out test prompts are not shared with model developers.[15] Alan Cowen, co-author of Google's launch post, was Hume's chief executive; PYMNTS, citing Wired, reported on 22 January 2026 that Google DeepMind had signed a licensing agreement with Hume and would hire Cowen and several of its engineers.[25][2]
Hume's findings for Flash-Lite TTS included:[15]
- A second-place score of 0.914 on Hume's new Expressivity-Reliability Frontier Score, behind Flash TTS (0.920) and ahead of Gemini 2.5 Pro TTS (0.880), Gemini 2.5 Flash TTS (0.861), Cartesia Sonic 3.6 (0.840) and Grok TTS (0.792).
- Extended long-form stability of 3.03 on a five-point scale, up from 1.22 for Gemini 3.1 Flash TTS and slightly above Flash TTS (2.89).
- On voice replication, seventh of 13 models on the same-speaker similarity measure, at 3.68 out of 5 against a field average of 3.69. Flash TTS ranked eleventh (3.53). Hume said both models scored more strongly on the naturalness of replicated speech, ranking second and fourth.
A second Hume post, on 29 September 2026, rated two-speaker dialogue across 48 scripts. Flash-Lite TTS scored 4.10 overall, against 4.11 for Flash TTS, 4.12 for Gemini 2.5 Pro TTS and 3.90 for ElevenLabs v3. It had the highest emotional-containment score of the models tested (4.31) and a seamlessness score of 4.21, against 3.97 for Flash TTS, but weaker speaker separation (3.88, against 4.17 for Flash TTS), especially for female-female (3.46) and male-male (3.31) pairs.[16]
Artificial Analysis
Artificial Analysis posted results on launch day. On its Provider Voice Arena, which ranks models by blind human comparisons of each model's own preset voices, Flash-Lite TTS debuted at sixth place with an Elo of 1,236, one point behind Simba 3.2 and ahead of Gemini 3.1 Flash TTS (1,199); Flash TTS debuted second at 1,263.[17] When the leaderboard was checked on 30 September 2026, Flash-Lite TTS stood seventh with an Elo of 1,240 and Flash TTS third at 1,268, behind ElevenLabs' Eleven v4 (1,316) and Cartesia's Sonic 3.6 (1,275).[26] On the Pronunciation Robustness Benchmark, Flash-Lite TTS scored 87.4%, fourth behind Flash TTS (89.5%), Gemini 3.1 Flash TTS (88.2%) and SpaceXAI TTS (87.6%).[17] In Artificial Analysis's Controlled Voice Arena, which uses the same replicated reference voices across models, Flash-Lite TTS debuted first in Japanese (Elo 1,212), second in Arabic (1,274) and third in German (1,250).[17]
Artificial Analysis measured Flash-Lite TTS at 40.2 characters per second, about 2.4 times faster than real time, which was slower than its measurement for Flash TTS (44.1 characters per second).[17] It added that both models "remain behind faster Text to Speech models we track, such as Falcon 2 at 204.9 characters per second and Luna TTS at 153.5 characters per second."[17] OpenRouter's listing showed a median (P50) latency of roughly 2.5 to 2.7 seconds for requests to Google AI Studio when checked on 30 September 2026; the figure is a rolling measurement that changes daily.[21]
Safety and consent
Google's launch post says "every audio clip generated by our Gemini Audio models is watermarked with SynthID," describing SynthID as an imperceptible watermark woven into the audio.[2] For voice replication, the post says users "must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created," and that replication is backed by consent verification, SynthID and C2PA credentials.[2] The API documentation requires a 10 to 30 second reference clip and a consent clip from the same adult speaker, who reads a fixed statement beginning "I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model."[12] The launch post, by contrast, describes replication from "just a 30-second audio sample."[2] A footnote to the post says voice replication through AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland and India.[2] These measures address the voice cloning and deepfake risks of consent-free replication.
The Gemini 3.8 Audio model card lists known limitations shared with foundation models, "such as hallucinations," and "occasional slowness or timeout issues."[7] For frontier safety it relies on Google DeepMind's Frontier Safety Framework evaluations of Gemini 3.7 Flash, which did not reach any tracked or critical capability levels; the card says the 3.8 Audio models, including Flash-Lite TTS, "do not have meaningful new capabilities or material increases in performance compared to Gemini 3.7 Flash."[7]
Comparison with Gemini 3.8 Flash TTS
| Gemini 3.8 Flash-Lite TTS | Gemini 3.8 Flash TTS | |
|---|---|---|
| Model ID | gemini-3.8-flash-lite-tts | gemini-3.8-flash-tts |
| Google's description of primary strength | "High throughput, low latency, and cost efficiency" | "Maximum voice fidelity, acting nuance, and dialect coverage" |
| Languages (model pages) | 101 | 130 |
| Replaces | gemini-3.1-flash-tts-preview | "New flagship creative tier" |
| Standard price through 2026 (per 1M tokens) | $0.50 in / $6.00 out | $0.50 in / $9.00 out |
| Token limits (Gemini API) | 8,192 in / 16,384 out | 8,192 in / 16,384 out |
| Consumer product | Google Vids | Gemini Notebook |
| Hume overall quality score (Google's table) | 0.914 | 0.920 |
| Artificial Analysis Provider Voice Arena, 23 Sep 2026 | #6, Elo 1,236 | #2, Elo 1,263 |
| Artificial Analysis Provider Voice Arena, 30 Sep 2026 | #7, Elo 1,240 | #3, Elo 1,268 |
| Artificial Analysis generation speed, 23 Sep 2026 | 40.2 characters/s | 44.1 characters/s |
Sources: Gemini API model pages, pricing page, launch post, DeepMind evaluation document and Artificial Analysis.[1][24][5][2][9][17][26]
The API documentation lists the same feature set for both models (single speaker, multi-speaker, voice design and voice replication), so the differences are in coverage and recommended use rather than switches in the API.[4] Flash-Lite TTS supports 29 fewer language entries, Google recommends Flash TTS for "complex multi-speaker dialogue, heavy vocal-burst acting, difficult pronunciation, regional dialects," and the documentation says overlapping backchannel speech works best on Flash TTS.[1][4] Google's DeepMind and blog pages present voice design as a Flash TTS feature, although the API documentation lists it for both.[8][2][11] Despite the "low latency" positioning, Artificial Analysis measured Flash-Lite as slightly slower than Flash TTS in characters per second at launch.[17]
Reception
Early coverage treated the two models as a pair and focused on voice replication and control. Android Authority said both models "can follow line-by-line directions for tone, pacing, dialect changes, whispers, laughs, sighs, and other conversational cues," and noted that ElevenLabs still scored higher in Hume's voice-quality and human-like variation categories.[18] The New Stack described Flash-Lite as "the faster, less expensive option and the direct replacement for gemini-3.1-flash-tts-preview," and called the verbatim-transcript rule "a breaking change for anyone who embedded stage directions in prompts to the 3.1 preview model."[19]
Among launch partners, Google named Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang as companies integrating the new TTS models, and Agora, LiveKit, Pipecat and Vercel as developer platforms supporting them.[2] Google DeepMind's page quotes Wondercraft CTO Mei Ki Yiu: "Gemini 3.8 Flash-Lite TTS brings the voice consistency and creative control Wondercraft needs, making it a powerful fit for producing standout podcasts, audio ads, and voice-driven content at scale."[8] Hume's evaluation summarized both 3.8 models as "among the leaders in Real World VoiceEQ" while noting that "Speaker similarity and the precision of some voice controls remain areas for improvement."[15]
See also
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24Google AI for Developers, "Gemini 3.8 Flash-Lite TTS" (model page), last updated 24 September 2026. ai.google.dev/...gemini-3.8-flash-lite-tts
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21Leland Rechis and Alan Cowen, "Gemini 3.8 text-to-speech says hello," Google blog (The Keyword), 23 September 2026. blog.google/...gemini-3-8-text-to-speech
- ^1 ^2 ^3 ^4 ^5 ^6Google AI for Developers, "Release notes" (Gemini API changelog), entries dated 22 September 2026, 15 April 2026 and 20 May 2025. ai.google.dev/...changelog
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17Google AI for Developers, "Text-to-speech generation (TTS)," last updated 24 September 2026. ai.google.dev/...speech-generation
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Google AI for Developers, "Gemini Developer API pricing," accessed 30 September 2026. ai.google.dev/...pricing
- ^1 ^2 ^3 ^4Google AI (@GoogleAI), "We're launching Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS," X, 23 September 2026. x.com/...2102781694730285427
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Google DeepMind, "Gemini 3.8 Audio - Model Card," published 15 September 2026. deepmind.google/...gemini-3-8-audio
- ^1 ^2 ^3 ^4 ^5 ^6Google DeepMind, "Gemini Audio - Speech generation," accessed 30 September 2026. deepmind.google/...speech-generation
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Google DeepMind, "Gemini 3.8 Audio (Flash TTS, Flash-Lite TTS) Model evaluation: Approach, methodology & results," September 2026. deepmind.google/...gemini-3-8-tts
- ^1 ^2fofr, "Gemini 3.8 Flash TTS: Developer Guide," Google AI Studio. aistudio.google.com/...8-flash-tts-developer-guide
- ^1 ^2 ^3Google AI for Developers, "Voice design," last updated 24 September 2026. ai.google.dev/...voice-design
- ^1 ^2Google AI for Developers, "Voice replication," last updated 24 September 2026. ai.google.dev/...voice-replication
- ^1 ^2Google AI for Developers, "Gemini deprecations," accessed 30 September 2026. ai.google.dev/...deprecations
- ^Google AI for Developers, "Models," accessed 30 September 2026. ai.google.dev/...models
- ^1 ^2 ^3Alice Baird, "Newly-released Google's Gemini 3.8 Flash TTS tops Hume's Real-World VoiceEQ leaderboard," Hume AI blog, 24 September 2026. hume.ai/...s-hume-s-real-world-voiceeq-leaderboard
- ^Sharath Rao, Kimberly Lo and Alice Baird, "Evaluating Google's multi-speaker TTS: A case study in why private evaluations matter," Hume AI blog, 29 September 2026. hume.ai/...evaluating-multi-speaker-tts
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Artificial Analysis (@ArtificialAnlys), "Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS" (thread), X, 23 September 2026. x.com/...2102784197853380647
- ^1 ^2Hillary Keverenge, "Gemini can now clone your voice and perform scripts like an actor," Android Authority, 24 September 2026. androidauthority.com/...speech-rolling-out-3714915
- ^1 ^2Amanda Caswell, "OpenAI makes you call sales for a custom voice. Google just made it self-serve.," The New Stack, 24 September 2026. thenewstack.io/gemini-tts-voice-replication-api
- ^1 ^2LiteLLM, "Day 0 support: Gemini 3.8 Flash TTS and Flash-Lite TTS," 23 September 2026. docs.litellm.ai/...gemini_3_8_flash_tts
- ^1 ^2OpenRouter, "Gemini 3.8 Flash Lite TTS - API Pricing & Benchmarks," accessed 30 September 2026. openrouter.ai/...gemini-3.8-flash-lite-tts
- ^Vercel, "Gemini 3.8 Flash-Lite TTS API & Pricing," Vercel AI Gateway, accessed 30 September 2026. vercel.com/...gemini-3.8-flash-lite-tts
- ^fal, "Gemini 3.8 Flash Lite TTS (Text to Speech) API on fal," accessed 30 September 2026. fal.ai/...gemini-3.8-flash-lite-tts
- ^1 ^2Google AI for Developers, "Gemini 3.8 Flash TTS" (model page), last updated 24 September 2026. ai.google.dev/...gemini-3.8-flash-tts
- ^PYMNTS, "Google Recruits Hume CEO Alan Cowen to Bolster Voice AI Efforts," 22 January 2026. pymnts.com/...-alan-cowen-bolster-voice-ai-efforts
- ^1 ^2 ^3Artificial Analysis, "Text to Speech Leaderboard" (Provider Voice Arena), accessed 30 September 2026. artificialanalysis.ai/...leaderboard
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 4,017 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent verification V6 (xg13, 30 Sep 2026): ~110 claims vs Gemini API docs/pricing/release notes, DeepMind eval PDF, Hume, Artificial Analysis, OpenRouter; 1 material (stale AA price comparison) + 8 minor fixed
Cite this page: AI Wiki. "Gemini 3.8 Flash-Lite TTS." aiwiki.ai, updated 30 Sept 2026, fact-checked 30 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/gemini_3_8_flash_lite_tts