Setting things up…
Setting things up…
Text-to-speech and voice generation you can ship
Compare and run the best AI voice and text-to-speech models in one place. Generate in your browser or over a simple API — pay per run, credits refunded on failed runs.
16 models ready to run

Coqui XTTS-v2: Multilingual Text To Speech Voice Cloning

The fastest open source TTS model without sacrificing quality.

Text-to-Audio (T2A) that offers voice synthesis, emotional expression, and multilingual capabilities. Optimized for high-fidelity applications like voiceovers and audiobooks.

Text-to-Audio (T2A) that offers voice synthesis, emotional expression, and multilingual capabilities. Designed for real-time applications with low latency

Voice cloning + text-to-speech — clone a voice from a short sample and make it say anything, multilingual.

ElevenLabs eleven-v3 is a text-to-speech model available as a hosted endpoint; requests cost $0.1 per 1000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

ElevenLabs Multilingual V2 is a multilingual text-to-speech model; cost $0.1 per 1000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

ElevenLabs Music generates original songs from text descriptions. Create instrumentals or full compositions with customizable duration. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Google Lyria 3 Pro generates high-quality music tracks from text prompts and optional image input.

Mirelo SFX1.6 Text To Audio generates sound effects or ambient audio directly from a text prompt, with optional seamless ambience looping.

mureka ai / mureka v9 / generate song via Mureka official API.

Music 2.6 generates complete songs with vocals and instrumentals from text prompts and lyrics.

Alibaba Qwen3 TTS Flash: Low-latency Text-to-Speech for English and Chinese with multiple voices, ideal for real-time dialogue. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Seed Audio 1.0 generates natural speech and audio from a prompt, with optional voice, reference audio, or reference image guidance.

Seed Speech TTS 2.0 converts text into natural speech with multilingual voices, delivery controls, and MP3 or Opus output.

's high-definition text-to-speech model with natural pronunciation and clear articulation.
The best choice depends on whether you need expressive multilingual text-to-speech or realistic voice cloning. Run the models below on your own script to compare quality and latency.
Yes — every model is callable over a simple REST API with one key, billed per run in credits. See the developer docs to start.
Each model shows its exact per-run credit cost before you run it, and failed runs are refunded automatically.