hexgrad/kokoro
Kokoro is a lightweight American-English text-to-speech model that is fast and very cheap to run.
Model Inputs
Text to read aloud. American English only
Voice ID for the desired voice.
Speed of the generated audio. Default is 1.0.
Result
Pricing
hexgrad/
kokoro
You are billed for what this model actually produces — by the 1k characters — so a shorter request costs less and you never pay for idle time.
For example, a request of 10 1k characters costs about $0.20.
You are charged for what the request actually produces, so the final amount scales with the 1k characters you use.
| Provider | Price ($) | Saving (%) |
|---|---|---|
| Synexa | $0.01 | - |
| fal | $0.01 | same price |
Comparison prices are for an equivalent request.
Readme
Kokoro is a small text-to-speech model that reaches quality comparable to far larger ones while being significantly faster and cheaper.
Twenty American-English voices are available - af_* are female, am_* are male. speed scales the delivery from 0.1x to 5x. This endpoint is American English only.