Explore AI models on Synexa — 70 models, one API
All models
Every model available on Synexa — 70 in total.

google/nano-banana-pro
Google's state of the art image generation and editing model 🍌🍌

black-forest-labs/flux-kontext-pro
A state-of-the-art text-based image editing model that delivers high-quality outputs with excellent prompt following and consistent results for transforming images through natural language

tencent/hunyuan3d-2
Hunyuan3D 2.0, an advanced large-scale 3D synthesis system for generating high-resolution textured 3D assets

google/veo3.1
New and improved version of Veo 3, with higher-fidelity video, context-aware audio, reference image and last frame support

tongyi/z-image-turbo
Z-Image Turbo is a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.

black-forest-labs/flux-2-klein-9b
4-step distilled version of FLUX.2 [klein], a 9B foundation image model offering maximum flexibility and control for fast text-to-image and image-to-image generation.

openai/gpt-image-2
OpenAI's state-of-the-art image generation model. Create and edit images from text with strong instruction following, sharp text rendering, and detailed editing.

kling/kling-motion-control
Kling 3.0 motion control: transfer motion from a reference video to any character image with improved consistency and quality.

zsxkib/mmaudio
Add sound to video using the MMAudio V2 model. An advanced AI model that synthesizes high-quality audio from video content, enabling seamless video-to-audio transformation.

tongyi/qwen-image-edit
Qwen's image editing model with multi-image editing, identity preservation, and ControlNet support for precise text, people, and product editing.

bytedance/seedance-2.0
ByteDance's advanced video generation model with native audio, keyframe control, and multimodal references for consistent character and style.

suno/suno-latest
Suno v5.5 generates full-length, high-fidelity AI songs with vocals and instrumentation from a single text prompt, producing 44.1 kHz stereo MP3 output.

tripo3d/tripo
State-of-the-art single-image to 3D object generation. Produces production-ready GLB meshes with PBR textures in under a second.

tongyi/wan2.2
Generate 5s 480p videos using Wan 2.2 14B. A comprehensive video foundation models that pushes the boundaries of video generation.

black-forest-labs/flux-1.1-pro
Faster, better FLUX Pro. Text-to-image model with excellent image quality, prompt adherence, and output diversity.

black-forest-labs/flux-kontext-dev
Open-weight version of FLUX.1 Kontext

black-forest-labs/flux-schnell
The fastest image generation model tailored for local development and personal use

tencent/hunyuan3d-3.1
Generate high-quality 3D models with accurate geometry and realistic textures from input images.

meshy/meshy-latest
Advanced AI-powered 3D model generator that transforms images into production-ready 3D assets with high-quality textures and multiple export formats including GLB, FBX, USDZ, and OBJ.

microsoft/trellis2
TRELLIS.2 is a state-of-the-art 4B parameter 3D generative model for high-fidelity image-to-3D generation with complex topologies, sharp features, and full PBR materials.

fofr/face-to-many
Turn a face into 3D, emoji, pixel art, video game, claymation or toy

vetkastar/fooocus
Image generation, Added: inpaint_strength loras_custom_urls

bytedance/sdxl-lightning-4step
SDXL-Lightning by ByteDance: a fast text-to-image model that makes high-quality images in 4 steps

stability-ai/stable-diffusion
A latent text-to-image diffusion model capable of generating photo-realistic images given any text input

ace-step/ace-step
ACE-Step generates music with sung lyrics from a list of genre tags, extremely cheaply.

aura/aura-sr
AuraSR upscales images 4x with a GAN-based model that is fast and consistent.
heygen/avatar-4
Avatar 4 turns a photo into a talking avatar that speaks your text or lip-syncs to an audio file.

clarity-ai/clarity-upscaler
Clarity Upscaler enlarges images while adding believable detail, guided by an optional prompt.

sczhou/codeformer
Robust face restoration algorithm for old photos / AI-generated faces

recraft/crisp-upscale
Crisp Upscale raises an image's resolution while sharpening small details and edges.

veed/fabric-1.0
Fabric 1.0 turns a photo plus an audio track into a talking-head video.

black-forest-labs/flux-3-video
FLUX 3 animates a single still image into video, with audio, following a prompt that describes how the scene unfolds.

tencentarc/gfpgan
Practical face restoration algorithm for *old photos* or *AI-generated faces*

xai/grok-imagine-image
Grok Imagine edits images from a text instruction, following the wording closely and keeping the rest of the frame intact.

xai/grok-imagine-video
Grok Imagine Video animates a still image into a short clip with audio.

minimax/hailuo-h3
Hailuo H3 builds video from a mix of reference images, clips and audio, addressed by position in the prompt.

ideogram/ideogram-v3
Ideogram V3 generates posters, logos and illustrations with reliable, correctly spelled typography.

supavisual/image-upscale
Professional-grade image upscaling. Increase image resolution up to 4x with adjustable creativity and resemblance controls.

kling/kling-image-o3
Kling Omni 3 generates images from reference pictures with strong subject consistency, up to 4K.

kling/kling-video-v3-pro
Kling V3 Pro turns a still image into cinematic video with strong motion consistency, and can generate native audio in the same pass.

hexgrad/kokoro
Kokoro is a lightweight American-English text-to-speech model that is fast and very cheap to run.

krea/krea-2-large
Krea 2 Large generates high-fidelity images from text with a distinctly photographic, non-plasticky look.

yorickvp/llava-13b
Visual instruction tuning towards large language and vision models with GPT-4 level capabilities

lightricks/ltx-2.5-audio-to-video
LTX-2.5 Pro generates video timed to a supplied audio track, optionally starting from an image.

lightricks/ltx-2.5-pro
LTX-2.5 Pro animates a still image into video with synchronised audio and optional camera moves.

google/lyria-2
Lyria 2 generates 30 seconds of 48kHz music from a descriptive text prompt.

synexa/merge-videos
Merges two or more videos into one, end to end, normalising frame rate and resolution.

xiaomi/mimo-tts
Xiaomi MiMo V2.5 text-to-speech. Speaks text with one of 9 built-in voices, clones a voice from a reference clip, or invents a new voice from a written description.

minimax/music-3
MiniMax Music 3 writes and performs a complete song from lyrics and a style description.

cassetteai/music-generator
A very fast instrumental music generator: a 30-second sample in under two seconds.

falcons-ai/nsfw_image_detection
Fine-Tuned Vision Transformer (ViT) for NSFW Image Classification

pixverse/pixverse-v6
PixVerse V6 animates a still image into stylised or realistic video, with optional audio.

nightmareai/real-esrgan
Real-ESRGAN with optional face correction and adjustable upscale

recraft/recraft-v3
Recraft V3 generates brand-consistent images, vector art and long passages of readable text.

danielgatis/rembg
A fast, no-frills background remover that returns the subject on a transparent background.

ideogram/remove-background
Ideogram's background remover isolates the subject cleanly against transparency.

supavisual/remove-background
Remove the background from an image, producing a clean cutout with a transparent background.

bria/rmbg-2.0
RMBG 2.0 cuts the subject out of an image and returns it on a transparent background.

meta/sam-3
SAM 3 segments anything you can name: give it an image and a word, and it returns the mask.

stability-ai/sdxl
A text-to-image generative AI model that creates beautiful images

bytedance/seed-audio-1.0
Seed Audio 1.0 generates natural speech, and can clone a voice from short reference clips.

bytedance/seedance-2.5
Seedance 2.5 generates a single-shot video of up to 30 seconds from a text prompt, with synchronised audio.

bytedance/seedream-5-pro-edit
Seedream 5.0 Pro is a region-precise image editor that changes one element of a picture while leaving the rest untouched.

bytedance/seedvr2-upscale
SeedVR2 restores and upscales images, recovering detail rather than simply interpolating pixels.

stability-ai/stable-audio-2.5
Stable Audio 2.5 generates music and sound effects from a text description.

microsoft/trellis
TRELLIS converts a single image into a textured 3D mesh.

elevenlabs/tts-v3
Eleven v3 reads text aloud with expressive, emotionally aware delivery across dozens of languages.

elevenlabs/voice-changer
ElevenLabs speech-to-speech voice changer. Keeps the words, timing and delivery of your source recording but speaks them in a voice cloned from a second audio clip.

fish-audio/voice-changer
Fish Audio speech-to-speech voice changer. Keeps the words, timing and delivery of your source recording but speaks them in a voice cloned from a second audio clip.

openai/whisper-v3
Whisper large v3 transcribes or translates speech from an audio file, in 99 languages.