AI text to speech

Natural speech from any text — for videos, podcasts, audiobooks and greetings. Ready-made voices in English and many other languages, or a clone of your own voice from a recording.

Voice models

  • Gemini 3.8 Flash TTS, Gemini 3.8 Flash Lite TTS and Gemini 3.1 Flash TTS by Google. The 3.8 models take a speaking style: calm, cheerful, whispering, like a news anchor.
  • MiniMax Speech 2.8 HD and MiniMax Speech 2.8 Turbo.
  • Fish Audio S2.1 Pro, S2 Pro and S1. Fish Audio S2.1 Pro also clones voices.
  • Seed Audio 1.0 by ByteDance speaks English and Chinese only, and it clones a voice from several samples.
  • MAI-Voice-2 and MAI-Voice-2-Flash by Microsoft speak English, Spanish, French and German, with an adjustable style and speed.
  • Grok Voice by xAI, and low-cost models such as Kokoro 82M (1 credit per 1,000 characters).

Voice cloning

In the Voice clone tab, upload 3–30 seconds of clean speech with no music (mp3, wav, m4a or mp4) — the model reads new text in that voice. For English text, both Fish Audio S2.1 Pro and Seed Audio 1.0 work; for most other languages, pick Fish Audio S2.1 Pro. You can clone your own voice or the voice of someone who has agreed to it.

What text to speech costs

Text to speech is paid in credits for every 1,000 characters. For example, Gemini 3.8 Flash TTS costs 3 credits per 1,000 characters, Fish Audio S2.1 Pro 5 credits per 1,000 characters, and MiniMax Speech 2.8 HD 10 credits per 1,000 characters. The price of the whole text shows on the Generate voice button before you start.

How to start

  1. Open the Text to speech tab and pick a model.
  2. Paste your text and choose a voice.
  3. Press Generate voice — the price shows on the button.

Prices

Seed Audio 1.021 credits per 1,000 characters
Fish Audio S2.1 Pro5 credits per 1,000 characters
MiniMax Speech 2.8 HD10 credits per 1,000 characters
Gemini 3.8 Flash TTS3 credits per 1,000 characters
Qwen-Audio-3.0-TTS Plus2 credits per 1,000 characters
Gemini 3.1 Flash TTS6 credits per 1,000 characters
Fish Audio S2 Pro5 credits per 1,000 characters
MAI-Voice-23 credits per 1,000 characters
Grok Voice2 credits per 1,000 characters
MiniMax Speech 2.8 Turbo6 credits per 1,000 characters
Aura-23 credits per 1,000 characters
Deepgram Flux TTS5 credits per 1,000 characters

Prices are in credits. The total for a generation shows on the button before you start.

Prices and capabilities checked against the studio catalog on October 7, 2026

FAQ

Is text to speech free?
No: text to speech costs credits on every plan. Credits come with paid plans and packs from 300 ₽ (≈ $3.51).
How much text can I convert at once?
It depends on the model: Gemini 3.8 Flash TTS takes up to 20,000 characters, MiniMax Speech 2.8 HD up to 5,000. A long book is easier to do chapter by chapter.
Can I set the tone and speed?
Gemini 3.8 Flash TTS, Gemini 3.8 Flash Lite TTS, MAI-Voice-2 and MAI-Voice-2-Flash take a speaking style. The two MAI-Voice models also change the speed; they speak English, Spanish, French and German.
Can I use my own voice?
Yes, in the Voice clone tab. You need a 3–30 second recording of clean speech. For English, use Fish Audio S2.1 Pro or Seed Audio 1.0.
Can I use the audio in my videos?
AstraGPT makes no claim to the result. If the audio uses someone else’s voice, you need that person’s consent.