Google released Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23. The text-to-speech models speak more than 100 languages and cost about $0.54 to $0.81 per hour of audio until the end of the year. Both can clone a voice from a short sample after the voice’s owner records a consent statement.
New to this? Read it in simple words
- Google released two new Gemini models that turn written text into speech.
- They speak more than 100 languages. An hour of audio costs less than one dollar until the end of the year.
- They can copy a person’s voice from a short recording. That person must first read a consent statement aloud.
- Google says it marks the audio with a hidden watermark. Nobody outside Google has tested these protections yet.
- Text-to-speech
- Technology that reads written text aloud in a synthetic voice.
- Voice cloning
- Making a synthetic voice that sounds like a real person.
- Watermark
- A hidden signal added to content that shows it was made by AI.
Two voice models, priced by the hour
The models turn written text into speech. Google’s documentation lists 130 languages for Flash and 101 for Flash-Lite. Developers can also describe a new voice in words, and the model designs it.
Google’s blog says the models offer more than 2,000 voices. Its developer documentation lists 30 studio voices plus hundreds of others.
Until December 31, Flash costs $9 and Flash-Lite $6 per million tokens of audio. One second of audio uses 25 tokens, so an hour costs about $0.81 or $0.54. Google says both prices double on January 1, 2027.
Google says the models will also come to its Enterprise API, to Gemini Notebook, and to Google Vids.
Prices double on January 1, 2027. Cloning in AI Studio is not offered in the EEA, the UK, Switzerland, India, Illinois, or Texas.
Cloning needs a recorded consent
To clone a voice, a developer uploads 10 to 30 seconds of clean speech. The same speaker must also read a fixed consent statement aloud. An automatic check then compares that recording with the sample.
The consent statement exists in 30 language versions. A project can store up to 200 designed or cloned voices, and Google keeps them for one year.
Google’s blog says cloning in AI Studio is not available in Illinois, Texas, the European Economic Area, the UK, Switzerland, or India. It does not say why, and it is unclear whether the same limits apply in the API.
Google says every clip carries its SynthID watermark, and that cloned voices also carry C2PA content labels. Its documentation and model card do not describe these steps, and no outside test has checked them.
Strong scores, and one conflict of interest
Artificial Analysis, an independent testing company, ranks Flash second in its voice arena, behind Cartesia’s Sonic 3.6. Flash is one point ahead of Alibaba’s Qwen-Audio-3.0, which is within the margin of error. Flash-Lite ranks sixth.
In the same company’s pronunciation test, Flash ranks first with 89.5 percent.
Google’s blog also cites first places in benchmarks run by Hume AI. One author of the blog, Alan Cowen, founded Hume. In January, TechCrunch reported that Google had hired him and his team.
A reporter at The Decoder found the accents convincing but heard an occasional high-pitched whine. In one clip, the voice also changed at the end.
Sources
Every fact in this story comes from the sources below. Open them to check our work.
- 1
- 2
- 3
- 4
- 5Primary source · September 15, 2026Gemini 3.8 Audio (Live, Live Extended Thinking, Flash TTS, Flash-Lite TTS) Google DeepMind
- 6Research · September 23, 2026Google’s new Flash TTS models let you design AI voices from scratch using text descriptions The Decoder
- 7Research · September 23, 2026Google’s new Gemini TTS models can clone a voice from 30 seconds of audio The Next Web
- 8
- 9
- 10Research · September 23, 2026Google launches Gemini 3.8 voice models with Hume founder Alan Cowen credited RuntimeWire
We read Google’s blog post, developer documentation, pricing page, and model card, then compared The Decoder, The Next Web, and the Artificial Analysis leaderboard. Quality claims from Hume AI tests are Google’s own, and Hume’s founder co-wrote the post. The safeguards are Google’s description.