An Ex-Apple Engineer Is Building an AI Agent That Lives Inside iMessage
We test the top AI tools for writing, video, images & music — so you don't have to.
Explore AI Tools
Google just released two new text-to-speech models in the Gemini family: Gemini 3.8 Flash TTS, built for high-fidelity creative work, and Gemini 3.8 Flash-Lite TTS, a cheaper tier designed for high-volume speech generation — dubbing, voice agents, and audio content at scale.
You can now describe a voice the way you would brief a casting director. Plain-English prompts control accent, role, pacing, and emotion — no training data or audio engineering required. The models cover 100+ languages and dialects, and Google has grown its library of ready-made production voices from 30 originals to over 2,000, including regional varieties like Mexican Spanish, Quebec French, and Scots English.
Voice cloning needs just a 30-second audio sample — but only with a verbal consent recording from the voice owner. Every generated clip carries a SynthID watermark plus C2PA credentials, so synthetic audio stays identifiable downstream. It is a strong identity-protection baseline for the industry.
You can direct the voice line by line, adding acting cues, pacing changes, tone shifts, and even dialect switches mid-script. Non-verbal cues like <laughs>, <sigh>, and <gasp> — plus active-listening interjections — make dialogue feel natural. Native two-speaker dialogue and long-form audio that keeps a voice's character consistent across hours open the door to podcasts and audiobooks. Voice remixing — adjusting timbre, pitch, pace, and accent through prompts — is coming soon.
On Hume AI's Voice Design Benchmark, Gemini 3.8 Flash TTS ranks #1 with 71.4 overall and 60.8 on accent modeling. Flash TTS takes #1 and Flash-Lite #2 on Hume's Overall Quality Index. In blind Voice Arena tests, the models hold top spots across Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.
Both models are available now through the Gemini API and Google AI Studio, which includes an audio playground with a dual-speaker screenplay editor. Flash TTS is also rolling out to Gemini Notebook users, while Flash-Lite TTS heads to Google Vids; a Gemini Enterprise API rollout follows later. Launch partners integrating the models include Agora, LiveKit, Pipecat, Vercel, Figma, HeyGen, and Wondercraft. One caveat: voice replication is not available in Illinois, Texas, the EEA, the UK, Switzerland, or India.
Google is turning voice generation into a programmable creative tool — and the consent-required, watermarked design is the safety bar other voice AI will be judged against.
Comments
Post a Comment