Gemini 3.8 Flash TTS

Direct AI voice acting with Gemini 3.8 Flash TTS: steer tone line by line, stage two speakers and output audio in 130 languages.

Gemini 3.8 Flash TTS
Shape expressive performances on the flagship tier, or keep spend low at volume with Flash-Lite TTS
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

What Sets the Gemini 3.8 Flash TTS Engine Apart

Announced on 23 September 2026, Google's audio release pairs an expressive flagship with a low-cost engine built for large batch workloads.

  • A Single Release, Two Distinct Jobs
    The flagship tier handles nuanced creative reads, while the Lite tier keeps per-minute costs down when speech is needed at volume.
  • Steering Delivery Instead of Choosing Presets
    Per-turn style notes, structured speech metadata and inline vocal cues govern tone, tempo, feeling and accent.
  • Designed Voices and Consented Replication
    Describe the timbre you want in plain words, or mirror a real person using a reference clip plus a matching consent recording.

Getting Your Prompts Right in Gemini 3.8 Flash TTS

Four habits that keep your transcript readable while the delivery metadata handles the acting.

What Gemini 3.8 Flash TTS Can Do, Feature by Feature

Performance control, voice design, cloning safeguards and multilingual reach — the full picture of Google's flagship speech model.

Line-by-Line Performance Control

Per-turn styling with inline laughs, sighs, coughs, breaths and pauses feels like coaching a performer rather than picking a stock voice.

Voice Design in Plain English

Describe age, personality, accent, texture and role in a prompt, drawing on more than 2,000 production voices through the Voices endpoint.

Cloning Behind a Consent Gate

You supply a clear reference take and a consent clip from that same adult voice, and every output carries SynthID watermarking plus C2PA credentials.

Dialogue Built for Two Voices

Write the exchange once and get podcasts, teaching dialogues, product walkthroughs or game scenes without splicing lines by hand.

Steady Voice Across Long Recordings

Google reports consistent identity, timbre, loudness and room tone across narrations and dialogues that run for several minutes.

130 Languages With Regional Accents

The flagship speaks 130 languages versus 101 on Lite, including regional accents, minority dialects and IPA overrides for tricky words.

FAQ

Gemini 3.8 Flash TTS: Questions People Ask

Quick answers on cost per minute, picking between tiers, benchmark scores and the rules around voice replication.

1

What does Gemini 3.8 Flash TTS charge per minute of audio?

Roughly 1.35 cents for every audio minute, based on launch pricing of $0.50 per million input tokens and $9 per million output tokens.

2

Which tier fits my project — Flash TTS or Flash-Lite TTS?

Choose the flagship when acting nuance and long-form audio matter; pick Lite for high-volume jobs and faster turnaround.

3

How does it score next to rival voice models?

Google cites a 71.4 score on Hume's Voice Design Benchmark, and Voice Arena ranks it second with 1,260 Elo.

4

Is it different from the Gemini 3.1 Flash TTS Preview?

The Lite tier takes the 3.1 preview's place and drops audio output pricing from $20 to $6 per million tokens.

5

Can I clone a voice, and what safeguards apply?

Yes, provided you supply a reference clip plus a consent recording from that same adult speaker.

6

Why is the model reading my stage directions aloud?

Everything you type is treated as spoken script, so shift lasting directions into the speech metadata field instead.

Run Gemini 3.8 Flash TTS on Scripts You Already Have

Try both tiers inside the Gemini API or Google AI Studio, where moving between them takes nothing more than a different model ID. Weigh batch against priority inference before you lock in a production budget.