AI text to speech
Natural speech from any text — for videos, podcasts, audiobooks and greetings. Ready-made voices in English and many other languages, or a clone of your own voice from a recording.
Natural speech from any text — for videos, podcasts, audiobooks and greetings. Ready-made voices in English and many other languages, or a clone of your own voice from a recording.
In the Voice clone tab, upload 3–30 seconds of clean speech with no music (mp3, wav, m4a or mp4) — the model reads new text in that voice. For English text, both Fish Audio S2.1 Pro and Seed Audio 1.0 work; for most other languages, pick Fish Audio S2.1 Pro. You can clone your own voice or the voice of someone who has agreed to it.
Text to speech is paid in credits for every 1,000 characters. For example, Gemini 3.8 Flash TTS costs 3 credits per 1,000 characters, Fish Audio S2.1 Pro 5 credits per 1,000 characters, and MiniMax Speech 2.8 HD 10 credits per 1,000 characters. The price of the whole text shows on the Generate voice button before you start.
| Seed Audio 1.0 | 21 credits per 1,000 characters |
|---|---|
| Fish Audio S2.1 Pro | 5 credits per 1,000 characters |
| MiniMax Speech 2.8 HD | 10 credits per 1,000 characters |
| Gemini 3.8 Flash TTS | 3 credits per 1,000 characters |
| Qwen-Audio-3.0-TTS Plus | 2 credits per 1,000 characters |
| Gemini 3.1 Flash TTS | 6 credits per 1,000 characters |
| Fish Audio S2 Pro | 5 credits per 1,000 characters |
| MAI-Voice-2 | 3 credits per 1,000 characters |
| Grok Voice | 2 credits per 1,000 characters |
| MiniMax Speech 2.8 Turbo | 6 credits per 1,000 characters |
| Aura-2 | 3 credits per 1,000 characters |
| Deepgram Flux TTS | 5 credits per 1,000 characters |
Prices are in credits. The total for a generation shows on the button before you start.
Prices and capabilities checked against the studio catalog on October 7, 2026