by Google
Gemini 2.5 Flash Preview TTS is a lightweight, low-latency text-to-speech model designed for real-time voice generation. It produces natural, expressive speech with accurate control over tone, style, and pacing, while dynamically adjusting speaking speed based on context and instructions. The model also maintains consistent and distinguishable voices across multi-turn and multi-speaker conversations, making it well-suited for interactive and conversational applications that require stable, high-quality audio output.
Use gemini-2.5-flash-preview-tts via the AIHubMix unified API — one interface for every major LLM.