Explore AI models on Synexa — 70 models, one API

Audio & Speech

Music, speech and sound effects — plus transcription back to text.

Sort
mmaudio

zsxkib/mmaudio

Add sound to video using the MMAudio V2 model. An advanced AI model that synthesizes high-quality audio from video content, enabling seamless video-to-audio transformation.

Video to Audio$0.01/ request
suno-latest

suno/suno-latest

Suno v5.5 generates full-length, high-fidelity AI songs with vocals and instrumentation from a single text prompt, producing 44.1 kHz stereo MP3 output.

Text to Audio$0.10/ request
ace-step

ace-step/ace-step

ACE-Step generates music with sung lyrics from a list of genre tags, extremely cheaply.

Text to Audio$0.0002/ second
kokoro

hexgrad/kokoro

Kokoro is a lightweight American-English text-to-speech model that is fast and very cheap to run.

Text to Audio$0.02/ 1k characters
lyria-2

google/lyria-2

Lyria 2 generates 30 seconds of 48kHz music from a descriptive text prompt.

Text to Audio$0.10/ request
mimo-tts

xiaomi/mimo-tts

Xiaomi MiMo V2.5 text-to-speech. Speaks text with one of 9 built-in voices, clones a voice from a reference clip, or invents a new voice from a written description.

Text to Audio$0.02/ request
music-3

minimax/music-3

MiniMax Music 3 writes and performs a complete song from lyrics and a style description.

Text to Audio$0.002/ second
music-generator

cassetteai/music-generator

A very fast instrumental music generator: a 30-second sample in under two seconds.

Text to Audio$0.02/ minute
seed-audio-1.0

bytedance/seed-audio-1.0

Seed Audio 1.0 generates natural speech, and can clone a voice from short reference clips.

Text to Audio$0.1875/ minute
stable-audio-2.5

stability-ai/stable-audio-2.5

Stable Audio 2.5 generates music and sound effects from a text description.

Text to Audio$0.20/ request
tts-v3

elevenlabs/tts-v3

Eleven v3 reads text aloud with expressive, emotionally aware delivery across dozens of languages.

Text to Audio$0.10/ 1k characters
voice-changer

elevenlabs/voice-changer

ElevenLabs speech-to-speech voice changer. Keeps the words, timing and delivery of your source recording but speaks them in a voice cloned from a second audio clip.

Audio to Audio$1.00/ request
voice-changer

fish-audio/voice-changer

Fish Audio speech-to-speech voice changer. Keeps the words, timing and delivery of your source recording but speaks them in a voice cloned from a second audio clip.

Audio to Audio$1.00/ request
whisper-v3

openai/whisper-v3

Whisper large v3 transcribes or translates speech from an audio file, in 99 languages.

Speech to Text$0.004/ request