11 / 2383

Google Gemini 3.8 Flash TTS Adds 30-Second Voice Cloning

TL;DR

Google’s Gemini 3.8 Flash TTS models bring advanced features to text-to-speech technology, including voice cloning, style control, and multilingual support. These models cater to a range of use cases, such as creating expressive audiobooks or optimizing high-volume tasks like dubbing and voice agents.

Nauti's Take

Voice cloning from 30 seconds of audio cuts the cost of audiobooks, dubbing and voice agents sharply, a real advantage for small teams. The risk grows at the same pace: the easier voices are to copy, the easier scam calls and fake voice messages become.

Anyone using Gemini TTS in production should document speaker consent carefully and plan for labeling synthetic voices.

Video

Sources