Google introduced Gemini 3.1 Flash TTS, a text-to-speech model designed to deliver improved controllability, expressivity and quality for developers, enterprises and everyday users building AI-speech applications. The model began rolling out on April 15, 2026, across multiple platforms including the Gemini API, Google AI Studio, Vertex AI, and Google Workspace via Google Vids.

The new artificial intelligence model achieves an Elo score of 1,211 on the Artificial Analysis TTS leaderboard, a benchmark measuring thousands of blind human preferences. Gemini 3.1 Flash TTS ranks within the “most attractive quadrant” for its combination of high-quality speech generation and low cost, according to Artificial Analysis assessments.

Audio Tags Enable Granular Control

Gemini 3.1 Flash TTS introduces audio tags, allowing developers to control vocal style, pace and delivery through natural language commands embedded directly into text input. This feature provides improved granularity for steering AI-speech output without requiring complex technical configurations.

The model supports three primary control mechanisms within Google AI Studio. Scene direction enables developers to define environments and provide specific dialogue instructions, helping characters remain consistent across multiple turns. Speaker-level specificity allows casting of unique audio profiles with director’s notes to toggle pace, tone and accent, with inline tags enabling mid-sentence expression changes. Seamless export functionality permits developers to export exact parameters as Gemini API code, ensuring consistent voices across projects and platforms.

Multilingual Support and Global Scale

Gemini 3.1 Flash TTS delivers high-fidelity speech with precise control across more than 70 languages. The model includes native multi-speaker dialogue capabilities and supports major markets with advanced style, pacing and accent control. Early testers from developer and enterprise communities reported that audio tags provide new levels of creative precision, transforming text into high-fidelity vocal performances.

Safety Features and Watermarking

All audio generated by Gemini 3.1 Flash TTS includes watermarking via SynthID, an imperceptible watermark interwoven directly into audio output. This technology enables reliable detection of AI-generated content to help prevent misinformation. Google published a detailed model card outlining the company’s approach to safety and responsibility regarding the technology.

Availability and Developer Access

Developers can access Gemini 3.1 Flash TTS through preview availability on the Gemini API and Google AI Studio. Enterprise users gain preview access via Vertex AI, while Google Workspace subscribers can utilize the model through Google Vids. The Google AI Studio Playground provides a starting point for experimenting with high-fidelity speech generation and the new audio tag configurations.