hexgrad/kokoro

Kokoro is a lightweight American-English text-to-speech model that is fast and very cheap to run.

$0.02 / 1k characters
GPU: A100

Model Inputs

string

Text to read aloud. American English only

string

Voice ID for the desired voice.

number

Speed of the generated audio. Default is 1.0.

Result

Idle
Waiting for your input...

Pricing

hexgrad/

kokoro

You are billed for what this model actually produces — by the 1k characters — so a shorter request costs less and you never pay for idle time.

Rate
$0.02 / 1k characters
or
50 1k characters / $1

For example, a request of 10 1k characters costs about $0.20.

You are charged for what the request actually produces, so the final amount scales with the 1k characters you use.

ProviderPrice ($)Saving (%)
Synexa$0.01-
fal$0.01same price

Comparison prices are for an equivalent request.

Readme

Kokoro is a small text-to-speech model that reaches quality comparable to far larger ones while being significantly faster and cheaper.

Twenty American-English voices are available - af_* are female, am_* are male. speed scales the delivery from 0.1x to 5x. This endpoint is American English only.