Explore AI models on Synexa — 70 models, one API
Audio & Speech
Music, speech and sound effects — plus transcription back to text.

zsxkib/mmaudio
Add sound to video using the MMAudio V2 model. An advanced AI model that synthesizes high-quality audio from video content, enabling seamless video-to-audio transformation.

suno/suno-latest
Suno v5.5 generates full-length, high-fidelity AI songs with vocals and instrumentation from a single text prompt, producing 44.1 kHz stereo MP3 output.

ace-step/ace-step
ACE-Step generates music with sung lyrics from a list of genre tags, extremely cheaply.

hexgrad/kokoro
Kokoro is a lightweight American-English text-to-speech model that is fast and very cheap to run.

google/lyria-2
Lyria 2 generates 30 seconds of 48kHz music from a descriptive text prompt.

xiaomi/mimo-tts
Xiaomi MiMo V2.5 text-to-speech. Speaks text with one of 9 built-in voices, clones a voice from a reference clip, or invents a new voice from a written description.

minimax/music-3
MiniMax Music 3 writes and performs a complete song from lyrics and a style description.

cassetteai/music-generator
A very fast instrumental music generator: a 30-second sample in under two seconds.

bytedance/seed-audio-1.0
Seed Audio 1.0 generates natural speech, and can clone a voice from short reference clips.

stability-ai/stable-audio-2.5
Stable Audio 2.5 generates music and sound effects from a text description.

elevenlabs/tts-v3
Eleven v3 reads text aloud with expressive, emotionally aware delivery across dozens of languages.

elevenlabs/voice-changer
ElevenLabs speech-to-speech voice changer. Keeps the words, timing and delivery of your source recording but speaks them in a voice cloned from a second audio clip.

fish-audio/voice-changer
Fish Audio speech-to-speech voice changer. Keeps the words, timing and delivery of your source recording but speaks them in a voice cloned from a second audio clip.

openai/whisper-v3
Whisper large v3 transcribes or translates speech from an audio file, in 99 languages.