Explore AI models on Synexa — 70 models, one API

All models

Every model available on Synexa — 70 in total.

Sort
nano-banana-pro

google/nano-banana-pro

Google's state of the art image generation and editing model 🍌🍌

Text to Image$0.10/ request
flux-kontext-pro

black-forest-labs/flux-kontext-pro

A state-of-the-art text-based image editing model that delivers high-quality outputs with excellent prompt following and consistent results for transforming images through natural language

Text to Image$0.02/ request
hunyuan3d-2

tencent/hunyuan3d-2

Hunyuan3D 2.0, an advanced large-scale 3D synthesis system for generating high-resolution textured 3D assets

Image to 3D$0.025/ request
veo3.1

google/veo3.1

New and improved version of Veo 3, with higher-fidelity video, context-aware audio, reference image and last frame support

Text to Video$0.80/ request
z-image-turbo

tongyi/z-image-turbo

Z-Image Turbo is a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.

Text to Image$0.002/ request
flux-2-klein-9b

black-forest-labs/flux-2-klein-9b

4-step distilled version of FLUX.2 [klein], a 9B foundation image model offering maximum flexibility and control for fast text-to-image and image-to-image generation.

Image to Image$0.01/ request
gpt-image-2

openai/gpt-image-2

OpenAI's state-of-the-art image generation model. Create and edit images from text with strong instruction following, sharp text rendering, and detailed editing.

Text to Image$0.10/ request
kling-motion-control

kling/kling-motion-control

Kling 3.0 motion control: transfer motion from a reference video to any character image with improved consistency and quality.

Image to Video$0.112/ request
mmaudio

zsxkib/mmaudio

Add sound to video using the MMAudio V2 model. An advanced AI model that synthesizes high-quality audio from video content, enabling seamless video-to-audio transformation.

Video to Audio$0.01/ request
qwen-image-edit

tongyi/qwen-image-edit

Qwen's image editing model with multi-image editing, identity preservation, and ControlNet support for precise text, people, and product editing.

Image to Image$0.01/ request
seedance-2.0

bytedance/seedance-2.0

ByteDance's advanced video generation model with native audio, keyframe control, and multimodal references for consistent character and style.

Text to Video$0.3034/ request
suno-latest

suno/suno-latest

Suno v5.5 generates full-length, high-fidelity AI songs with vocals and instrumentation from a single text prompt, producing 44.1 kHz stereo MP3 output.

Text to Audio$0.10/ request
tripo

tripo3d/tripo

State-of-the-art single-image to 3D object generation. Produces production-ready GLB meshes with PBR textures in under a second.

Image to 3D$0.16/ request
wan2.2

tongyi/wan2.2

Generate 5s 480p videos using Wan 2.2 14B. A comprehensive video foundation models that pushes the boundaries of video generation.

Image to Video$0.20/ request
flux-1.1-pro

black-forest-labs/flux-1.1-pro

Faster, better FLUX Pro. Text-to-image model with excellent image quality, prompt adherence, and output diversity.

Text to Image$0.02/ request
flux-kontext-dev

black-forest-labs/flux-kontext-dev

Open-weight version of FLUX.1 Kontext

Text to Image$0.01/ request
flux-schnell

black-forest-labs/flux-schnell

The fastest image generation model tailored for local development and personal use

Text to Image$0.0015/ request
hunyuan3d-3.1

tencent/hunyuan3d-3.1

Generate high-quality 3D models with accurate geometry and realistic textures from input images.

Image to 3D$0.16/ request
meshy-latest

meshy/meshy-latest

Advanced AI-powered 3D model generator that transforms images into production-ready 3D assets with high-quality textures and multiple export formats including GLB, FBX, USDZ, and OBJ.

Image to 3D$0.40/ request
trellis2

microsoft/trellis2

TRELLIS.2 is a state-of-the-art 4B parameter 3D generative model for high-fidelity image-to-3D generation with complex topologies, sharp features, and full PBR materials.

Image to 3D$0.25/ request
face-to-many

fofr/face-to-many

Turn a face into 3D, emoji, pixel art, video game, claymation or toy

Image to Image$0.004/ request
fooocus

vetkastar/fooocus

Image generation, Added: inpaint_strength loras_custom_urls

Text to Image$0.015/ request
sdxl-lightning-4step

bytedance/sdxl-lightning-4step

SDXL-Lightning by ByteDance: a fast text-to-image model that makes high-quality images in 4 steps

Text to Image$0.001/ request
stable-diffusion

stability-ai/stable-diffusion

A latent text-to-image diffusion model capable of generating photo-realistic images given any text input

Text to Image$0.0007/ request
ace-step

ace-step/ace-step

ACE-Step generates music with sung lyrics from a list of genre tags, extremely cheaply.

Text to Audio$0.0002/ second
aura-sr

aura/aura-sr

AuraSR upscales images 4x with a GAN-based model that is fast and consistent.

Super Resolution$0.006/ request
avatar-4

heygen/avatar-4

Avatar 4 turns a photo into a talking avatar that speaks your text or lip-syncs to an audio file.

Image to Video$0.10/ second
clarity-upscaler

clarity-ai/clarity-upscaler

Clarity Upscaler enlarges images while adding believable detail, guided by an optional prompt.

Super Resolution$0.03/ megapixel
codeformer

sczhou/codeformer

Robust face restoration algorithm for old photos / AI-generated faces

Super Resolution$0.0025/ request
crisp-upscale

recraft/crisp-upscale

Crisp Upscale raises an image's resolution while sharpening small details and edges.

Super Resolution$0.004/ request
fabric-1.0

veed/fabric-1.0

Fabric 1.0 turns a photo plus an audio track into a talking-head video.

Image to Video$0.15/ second
flux-3-video

black-forest-labs/flux-3-video

FLUX 3 animates a single still image into video, with audio, following a prompt that describes how the scene unfolds.

Image to Video$0.12/ second
gfpgan

tencentarc/gfpgan

Practical face restoration algorithm for *old photos* or *AI-generated faces*

Super Resolution$0.0008/ request
grok-imagine-image

xai/grok-imagine-image

Grok Imagine edits images from a text instruction, following the wording closely and keeping the rest of the frame intact.

Image to Image$0.022/ image
grok-imagine-video

xai/grok-imagine-video

Grok Imagine Video animates a still image into a short clip with audio.

Image to Video$0.07/ second
hailuo-h3

minimax/hailuo-h3

Hailuo H3 builds video from a mix of reference images, clips and audio, addressed by position in the prompt.

Image to Video$0.13/ second
ideogram-v3

ideogram/ideogram-v3

Ideogram V3 generates posters, logos and illustrations with reliable, correctly spelled typography.

Text to Image$0.06/ image
image-upscale

supavisual/image-upscale

Professional-grade image upscaling. Increase image resolution up to 4x with adjustable creativity and resemblance controls.

Super Resolution$0.05/ request
kling-image-o3

kling/kling-image-o3

Kling Omni 3 generates images from reference pictures with strong subject consistency, up to 4K.

Image to Image$0.028/ image
kling-video-v3-pro

kling/kling-video-v3-pro

Kling V3 Pro turns a still image into cinematic video with strong motion consistency, and can generate native audio in the same pass.

Image to Video$0.168/ second
kokoro

hexgrad/kokoro

Kokoro is a lightweight American-English text-to-speech model that is fast and very cheap to run.

Text to Audio$0.02/ 1k characters
krea-2-large

krea/krea-2-large

Krea 2 Large generates high-fidelity images from text with a distinctly photographic, non-plasticky look.

Text to Image$0.06/ image
llava-13b

yorickvp/llava-13b

Visual instruction tuning towards large language and vision models with GPT-4 level capabilities

OCR & Vision$0.0005/ request
ltx-2.5-audio-to-video

lightricks/ltx-2.5-audio-to-video

LTX-2.5 Pro generates video timed to a supplied audio track, optionally starting from an image.

Audio to Video$0.17/ second
ltx-2.5-pro

lightricks/ltx-2.5-pro

LTX-2.5 Pro animates a still image into video with synchronised audio and optional camera moves.

Image to Video$0.17/ second
lyria-2

google/lyria-2

Lyria 2 generates 30 seconds of 48kHz music from a descriptive text prompt.

Text to Audio$0.10/ request
merge-videos

synexa/merge-videos

Merges two or more videos into one, end to end, normalising frame rate and resolution.

Video to Video$0.004/ request
mimo-tts

xiaomi/mimo-tts

Xiaomi MiMo V2.5 text-to-speech. Speaks text with one of 9 built-in voices, clones a voice from a reference clip, or invents a new voice from a written description.

Text to Audio$0.02/ request
music-3

minimax/music-3

MiniMax Music 3 writes and performs a complete song from lyrics and a style description.

Text to Audio$0.002/ second
music-generator

cassetteai/music-generator

A very fast instrumental music generator: a 30-second sample in under two seconds.

Text to Audio$0.02/ minute
nsfw_image_detection

falcons-ai/nsfw_image_detection

Fine-Tuned Vision Transformer (ViT) for NSFW Image Classification

Utilities$0.0001/ request
pixverse-v6

pixverse/pixverse-v6

PixVerse V6 animates a still image into stylised or realistic video, with optional audio.

Image to Video$0.025/ second
real-esrgan

nightmareai/real-esrgan

Real-ESRGAN with optional face correction and adjustable upscale

Super Resolution$0.0015/ request
recraft-v3

recraft/recraft-v3

Recraft V3 generates brand-consistent images, vector art and long passages of readable text.

Text to Image$0.04/ image
rembg

danielgatis/rembg

A fast, no-frills background remover that returns the subject on a transparent background.

Image to Image$0.01/ request
remove-background

ideogram/remove-background

Ideogram's background remover isolates the subject cleanly against transparency.

Image to Image$0.01/ request
remove-background

supavisual/remove-background

Remove the background from an image, producing a clean cutout with a transparent background.

Image to Image$0.01/ request
rmbg-2.0

bria/rmbg-2.0

RMBG 2.0 cuts the subject out of an image and returns it on a transparent background.

Image to Image$0.018/ request
sam-3

meta/sam-3

SAM 3 segments anything you can name: give it an image and a word, and it returns the mask.

Image to Image$0.005/ request
sdxl

stability-ai/sdxl

A text-to-image generative AI model that creates beautiful images

Text to Image$0.002/ request
seed-audio-1.0

bytedance/seed-audio-1.0

Seed Audio 1.0 generates natural speech, and can clone a voice from short reference clips.

Text to Audio$0.1875/ minute
seedance-2.5

bytedance/seedance-2.5

Seedance 2.5 generates a single-shot video of up to 30 seconds from a text prompt, with synchronised audio.

Text to Video$0.22/ second
seedream-5-pro-edit

bytedance/seedream-5-pro-edit

Seedream 5.0 Pro is a region-precise image editor that changes one element of a picture while leaving the rest untouched.

Image to Image$0.0675/ image
seedvr2-upscale

bytedance/seedvr2-upscale

SeedVR2 restores and upscales images, recovering detail rather than simply interpolating pixels.

Super Resolution$0.001/ megapixel
stable-audio-2.5

stability-ai/stable-audio-2.5

Stable Audio 2.5 generates music and sound effects from a text description.

Text to Audio$0.20/ request
trellis

microsoft/trellis

TRELLIS converts a single image into a textured 3D mesh.

Image to 3D$0.02/ request
tts-v3

elevenlabs/tts-v3

Eleven v3 reads text aloud with expressive, emotionally aware delivery across dozens of languages.

Text to Audio$0.10/ 1k characters
voice-changer

elevenlabs/voice-changer

ElevenLabs speech-to-speech voice changer. Keeps the words, timing and delivery of your source recording but speaks them in a voice cloned from a second audio clip.

Audio to Audio$1.00/ request
voice-changer

fish-audio/voice-changer

Fish Audio speech-to-speech voice changer. Keeps the words, timing and delivery of your source recording but speaks them in a voice cloned from a second audio clip.

Audio to Audio$1.00/ request
whisper-v3

openai/whisper-v3

Whisper large v3 transcribes or translates speech from an audio file, in 99 languages.

Speech to Text$0.004/ request