API reference

Base URL /v1 · JSON in, audio out · $ per million input characters.

Quickstart

  1. Create an account — you get free characters.
  2. Create a key on the API keys page.
  3. Make a request:
curl "{API}/audio/speech" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "{MODEL}", "voice": "alba", "input": "Hello! Your order ships today."}' \
  --output hello.mp3

Authentication

Send your key as a bearer token: Authorization: Bearer sk_live_.... X-API-Key is also accepted. Keys are secret; call the API from your server, never from a browser or mobile app.

sk_test_ keys return real audio for free, limited to 100 characters per request and 100 requests per day.

Using the OpenAI SDKs

The speech endpoint is wire-compatible with OpenAI's. Point the SDK's base URL at us; tts-1, tts-1-hd and gpt-4o-mini-tts are accepted as model names, and OpenAI voice names (alloy, nova, onyx, ...) map to similar stock voices. instructions is accepted and ignored (reported in X-Ignored-Parameters).

from openai import OpenAI

client = OpenAI(base_url="{API}", api_key="sk_live_...")
client.audio.speech.create(model="{MODEL}", voice="alba", input="Hi!").write_to_file("hi.mp3")

Create speech

POST/v1/audio/speech

ParameterTypeDefaultDescription
inputstring | string[]requiredText to speak. Up to 10,000 characters (20,000 with stream). Pass an array to keep your own chunk boundaries (e.g. sentences from an LLM).
voicestringalbaStock voice id, your clone id (cloned_...) or an OpenAI voice name.
modelstring{MODEL}Optional. OpenAI model names are accepted as aliases.
response_formatstringmp3mp3, opus (Ogg), aac, flac, wav, pcm (24 kHz, 16-bit, mono, little-endian).
speednumber1.00.5 to 2.0. Pitch is preserved.
normalizebooleantrueTurn text into speakable words (see normalization). Disable if you pre-process text yourself.
pronunciationsarray-Up to 100 {"word","respelling"} overrides, applied before the built-in dictionary.
streambooleanfalseStream audio as it is generated. mp3, opus, aac, pcm. stream_format: "audio" is an alias.
deliverystringbytesbytes: audio body. url: JSON with a download link valid for 24 hours. json: base64 audio plus usage.
include_costbooleanfalseAdds X-Cost-USD and X-Balance-USD headers. JSON deliveries always include a usage object.
return_normalized_textbooleanfalseEcho the exact text spoken, per chunk (json/url delivery).
seedintegerrandomRepeatable output for the same input and settings.
chunk_silence_msinteger80Pause inserted between chunks, 0-2000.

Every response carries X-Request-Id, X-Billable-Characters and X-Voice-Id; non-streamed responses also include X-Audio-Duration-Ms.

Streaming

With "stream": true the first bytes arrive after the first sentence is synthesized, so playback can start long before the full clip is ready. Use pcm for the lowest latency in voice agents (play samples directly), or mp3 for broad player support.

curl -N "{API}/audio/speech" -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
  -d '{"input": "A long story...", "voice": "george", "response_format": "pcm", "stream": true}' \
  | ffplay -f s16le -ar 24000 -ac 1 -nodisp -autoexit -
If your client disconnects mid-stream, generation stops and the request is still billed. If our engine fails mid-stream, the connection is closed abnormally and the request is refunded.

Delivery options

Audio is never stored by default. With "delivery": "url" we keep the file for 24 hours and return:

{
  "object": "audio.speech",
  "id": "req_...",
  "voice": "alba",
  "format": "mp3",
  "duration_ms": 5120,
  "audio": { "url": "{API}/files/aud_...mp3?exp=...&sig=...", "expires_at": 1791763200, "bytes": 41210 },
  "usage": { "billable_characters": 84, "normalized_characters": 97, "cost_usd": 0.00042, "balance_usd": 12.34558 }
}

Anyone with the link can download the file until it expires, so treat links like secrets. Not available with zero data retention.

Text normalization

On by default. Numbers, currency ($1,299.99), dates (3/14), times (9:30pm), ordinals, temperatures, phone-style digit runs, acronyms, Roman numerals, markdown, URLs and emoji are converted into words the voice reads naturally. Preview the result for free:

POST/v1/normalize

{ "input": "Dr. Smith paid $42.50 on 3/14.", "pronunciations": [{ "word": "Smith", "respelling": "Smyth" }] }
-> { "normalized": "Doctor Smyth paid forty two dollars and fifty cents on March fourteenth.", "billable_characters": 30, "estimated_cost_usd": 0.00015 }

List voices

GET/v1/voices   GET/v1/voices/{id}

Returns your clones first, then the stock catalogue. Filter with ?type=clone or ?type=stock. Every stock voice has a preview_url you can play without authentication.

Voice cloning

POST/v1/voices   PATCH/v1/voices/{id}   DELETE/v1/voices/{id}

Accept the cloning terms once in the dashboard, then upload 3-60 seconds of clean speech (10-20 s is ideal). Accepted formats: wav, mp3, m4a/aac, ogg/opus, webm, flac, up to 12 MB. Up to 20 voices per account; creating them is free and they cost the same to use as stock voices.

curl "{API}/voices" -H "Authorization: Bearer $API_KEY" \
  -F name="Sam narrator" -F file=@sam.m4a -F consent=true
-> { "id": "cloned_3f9c...", "object": "voice", "name": "Sam narrator", "type": "clone", ... }

consent=true is your attestation, for this specific recording, that it is your voice or that you have the speaker's documented permission. Your recording is decoded in memory and discarded; we keep only an encrypted voice profile.

Balance & usage

GET/v1/balance

{ "object": "balance", "balance_usd": 12.34, "characters_remaining": 2468000,
  "auto_topup": { "enabled": true, "threshold_usd": 5, "amount_usd": 25, "pending": false, "last_failure": null } }

GET/v1/usage?start=2026-10-01&end=2026-10-31&group_by=day — group_by is day, key or none. Totals include requests, errors, characters, audio seconds and cost.

Warmup & cold starts

Most requests are served by a warm instance. A cold instance takes up to ~11 seconds to start. If a request must be fast (a live call is about to begin, say), warm up a few seconds beforehand. It is free and limited to once per minute per key.

POST/v1/warmup?wait=true

{ "object": "warmup", "status": "warm", "waited_ms": 8420, "instance_uptime_s": 9.1 }

Without wait it returns immediately with "warming" or "warm". Instances stay warm for several minutes after their last request. Under bursty load a new request may still land on a fresh instance.

Zero data retention

Turn on ZDR for your whole account (Profile) or per key. With ZDR we write no request logs and store none of your text. Character counts and charges are still recorded because they are your billing record, and /v1/usage keeps working. delivery: "url" is rejected with zdr_conflict.

Errors

Errors use OpenAI's shape and always include a stable code and the request_id:

{ "error": { "message": "This request costs $0.000420 but your balance is $0.000100. Add credits or enable auto top-up.",
  "type": "billing_error", "code": "insufficient_balance", "param": null, "request_id": "req_8fK2...", "doc_url": "..." } }

If a request fails on our side (5xx), you are not charged. 429 and 503 responses include Retry-After.

HTTPcodeWhat to do
400invalid_request, invalid_jsonFix the parameter named in param.
400empty_inputThe input had nothing speakable (only emoji or symbols?).
400unsupported_formatUse a listed format; streaming needs mp3, opus, aac or pcm.
400model_not_foundUse {MODEL} or an OpenAI TTS model name.
400invalid_pronunciationWords and respellings may contain letters, spaces, apostrophes and hyphens only.
400zdr_conflictUse bytes or json delivery with ZDR.
400test_key_limitTest keys allow 100 characters per request.
401missing_api_key, invalid_api_key, api_key_revokedCheck the key or create a new one.
402insufficient_balanceAdd credits or turn on auto top-up.
402monthly_cap_reachedRaise or remove the key's monthly cap.
403account_suspendedContact support.
403test_key_not_allowedUse a live key for cloning.
404voice_not_foundList valid ids with GET /v1/voices.
409idempotency_conflict, idempotency_in_progress, idempotency_replaySee retries.
409voice_limit_reached, key_limit_reachedDelete an existing voice or key.
413input_too_long, audio_file_too_largeSplit the text, or use stream for up to 20,000 characters.
415unsupported_audio_type, unsupported_media_typeUpload a supported audio format; send JSON with the right Content-Type.
422clone_audio_too_short, clone_audio_too_long, clone_audio_unusableUpload 3-60 s of clear speech.
429rate_limited, concurrency_limited, test_quota_exhaustedWait for Retry-After seconds.
500internal_errorRetry; quote the request id if it persists. Not billed.
502engine_errorRetry. Not billed.
503engine_unavailable, engine_cold_start_timeoutRetry after Retry-After, or warm up first. Not billed.
504upstream_timeoutRetry with shorter input. Not billed.

Retries & idempotency

Send an Idempotency-Key header (any unique string, e.g. a UUID) to make retries safe for 24 hours:

Retry 429, 500, 502, 503 and 504 with exponential backoff. Don't retry other 4xx errors without changing the request.

Limits

Characters per request10,000 (20,000 streaming)
Requests per minute120 per key (configurable up to 600)
Concurrent requests10 per account
Cloned voices20 per account
Clone reference audio3-60 s, 12 MB
Download links24 hours
Request logs30 days (none with ZDR)

Need more? Contact us.

OpenAPI spec

The machine-readable spec is at /v1/openapi.json. Import it into Postman, Insomnia or an SDK generator.