Google is turning text-to-speech from a menu of presets into something closer to a recording studio. On Tuesday, the company introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two speech models that let users design entirely new voices from a natural-language description, direct the performance line by line, and replicate a voice from a 30-second audio sample.

Both models are rolling out immediately in the Gemini API and Google AI Studio. The launch arrives with striking benchmark claims — including a top score on the voice-industry benchmark run by Hume AI — a roster of launch partners, and, unusually for a cloning-capable model, an explicit consent gate: you must supply a verbal consent recording from the voice owner before replication is allowed.

Two models, one studio#

The pair splits the job the way Google's Flash line usually does: one model for craft, one for scale.

ModelPitched atRolling out in
Gemini 3.8 Flash TTSDeep creative direction and character design — gaming, audiobooks, podcasts, interactive mediaGemini API, Google AI Studio, Gemini Notebook (enterprise API coming soon)
Gemini 3.8 Flash-Lite TTSHigh-volume, cost-efficient speech — dubbing, content narration, expressive voice agentsGemini API, Google AI Studio, Google Vids (enterprise API coming soon)

The models sit in a growing Gemini Audio family that now includes 3.5 Live Translate, 3.5 Transcribe, and the 3.8 Live models. Developer platforms including Agora, LiveKit, Pipecat, and Vercel are integrating the speech generation, and Google named partners such as Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang as early adopters — a clear signal of where the first production use will land: dubbing, localized media, and voice agents.

Design a voice from a sentence#

The headline shift is in how voices come into being. Instead of picking from a preset list, users describe the voice they want — role, accent, character, general vibe — and the model builds it from that text alone. Google says Flash TTS scales "from 30 original voices to an infinite library," backed by more than 2,000 production-ready voices across 100-plus languages and dialects, including regional varieties like Mexican Spanish, Quebec French, and Scots English.

Created voices can be saved, managed across projects, and reused with minimal drift. A voice-remixing feature is promised next: pick any library voice and fine-tune its timbre, pitch, pace, or accent with prompts like "add subtle Southern US accent."

Extreme close-up of the woven metal grille of a studio condenser microphone
Photo by Lucasbosch, CC BY-SA 3.0, via Wikimedia Commons.

Direct the performance, line by line#

Once voices are chosen, both models offer what Google calls performance direction: write stage directions for individual lines, or let the model steer delivery from natural script cues — a calm customer-service agent one moment, a whispered suspense scene the next. Control covers acting cues, pacing, dialect shifts, and backchanneling.

The details aimed at audio producers are specific: long-form generation that holds voice quality, pacing, and character timbre across hours of continuous output; native two-speaker scene staging for multi-turn conversations with distinct voices; and scripted vocal bursts and non-verbal cues — <laughs>, <sigh>, <gasp>, plus active-listening interjections like |mhm| or |yeah| — for timing and reaction beats in dialogue.

A glowing cyan digital audio waveform flowing across a dark studio background with soft bokeh lights
Illustration generated with AI.

The benchmark claims#

Google is leading with third-party numbers, which is worth noting in itself. According to the announcement, Gemini 3.8 Flash TTS took the #1 overall spot on Hume AI's Voice Design Benchmark with a 71.4 score and led accent modeling at 60.8, while Flash TTS and Flash-Lite TTS ranked #1 and #2 respectively on Hume's Overall Quality Index. In blind human preference tests on Voice Arena, the models secured top positions in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.

There is a small irony in the source of the scores: the announcement is bylined by Group Product Manager Leland Rechis alongside Alan Cowen, a Director of Research Science credited on behalf of the Gemini Audio Team — Cowen is the former founder and CEO of Hume AI, the very company behind the benchmark Google is topping. As with all vendor-reported benchmarks, the numbers are a starting signal, not a verdict; independent testing will sort out how they hold up outside the lab.

Voice replication from a 30-second sample is the most sensitive capability here, and Google has wrapped it in process. Replication requires a verbal consent recording from the voice owner that matches the reference speaker before a voice profile can be created. Every clip generated by the Gemini Audio models carries a SynthID watermark woven into the audio, plus C2PA credentials for provenance.

The safeguards are the right shape, but the practical test — as with every cloning product — will be enforcement: how reliably the consent check resists spoofing, and whether watermarks survive the messy real world of re-recording and compression. Pricing and latency for Flash-Lite at true production scale remain unannounced details that matter for the dubbing and agent use cases Google is targeting.

What to watch#

  • Enterprise rollout and pricing. Gemini Enterprise API access is listed as "coming soon" — the cost curve decides whether Flash-Lite becomes the default for high-volume speech.
  • Consent verification under pressure. The first public attempts to defeat the cloning gate will be the real audit of the safeguards.
  • Independent benchmarks. Third-party re-testing of the Hume and Voice Arena claims, and comparisons against ElevenLabs, Hume's own models, and OpenAI's voice stack.
  • Voice remixing. The promised prompt-based fine-tuning of library voices would extend the "design studio" metaphor — and the cloning surface — further.

Text-to-speech has been good enough for narration for years; it has rarely been a creative instrument. Google's bet is that the next unlock is not better reading, but direction — and that voice design, like image generation before it, moves from selecting to prompting. The studio is open; now we find out who shows up to use it.