How it works
- Create an API key in the ReadAloud console and store it as READALOUD_API_KEY on your server. Never put it in a browser or mobile app.
- Before each call, POST https://api.readaloudai.org/tts/authorize with {"key": ..., "engine": "live"}. You get back a token and a WebSocket url. The token lasts 60 seconds, so authorize again for every new connection.
- Open the WebSocket at wss://.../tts?token=<token> and keep it open for the whole call. Send {"type":"synthesize","text":"...","voice":"default","format":"mulaw_8000"} for each sentence your model produces.
- Read chunk_meta frames, each followed by one binary audio frame, and forward the audio to your telephony stream. Each chunk_meta reports the format and sample_rate of the audio that follows.
- When your voice-activity detector reports the caller is talking, send {"type":"stop"}. The server replies with cancelled and you discard any queued audio.
- If you build on a framework, skip the socket code: pip install pipecat-readaloud and use ReadAloudHttpTTSService with sample_rate=8000, or pip install livekit-plugins-readaloud and pass readaloud.TTS(voice='readaloud-default') to AgentSession.
- Handle a 1013 close or an 'at capacity' error by retrying after a short backoff, and handle 402 by prompting the account owner to add a payment method.
A sample text
An illustrative example written for this page. It is the text you would send to the API, not a recording, and it makes no claim about how the audio sounds.
Thanks for calling Harbor Dental. I can book, move or cancel an appointment. Which would you like to do today?
Choose ReadAloud Live with format mulaw_8000 for a PSTN or SIP trunk, or pcm_24000 for a web call. Stream over the WebSocket and send one sentence per synthesize message so playback starts after the first sentence rather than the whole reply. Keep the connection open between turns.
When ReadAloud fits
- Your agent talks over a phone line and you want 8 kHz G.711 audio straight from the API.
- You prefer to pay per character of speech rather than per minute of call time.
- You use Pipecat or LiveKit Agents and want a maintained, MIT-licensed plugin rather than custom socket code.
- One American English voice is acceptable, and you control the wording so the script stays short.
When another service may suit you
- You need several distinct voices or personas in one product; Live has one voice, and Studio's English voice list is the only broader option here.
- Callers speak languages other than English; look for a vendor that documents each language and voice you need.
- You want laughter, whispers or controlled emotion; ReadAloud ignores style instructions and does not interpret SSML or audio tags.
- You need a platform that already bundles telephony, recognition and the model into one hosted product with a support contract.
Features used
- WebSocket streaming: Audio arrives per sentence over a connection you keep open, so the first sentence can play while the model writes the rest.
- stop message: Sending stop ends the current speech at once, which is how barge-in is implemented; only finished requests are billed.
- Telephony formats: mulaw_8000, alaw_8000 and pcm_8000 are available per request, so phone audio needs no transcoding.
- Pipecat and LiveKit plugins: pipecat-readaloud and livekit-plugins-readaloud read READALOUD_API_KEY and can output 8 kHz audio.
- Vapi custom voice route: POST /v1/vapi/custom-voice returns raw 16-bit mono PCM at the rate Vapi requests; see the Vapi integration page for what was and was not tested.
Things to consider
- Some jurisdictions and carriers expect you to tell callers they are speaking with an automated system. That disclosure is your responsibility.
- Each Live server accepts a limited number of simultaneous streams. Plan for retries and test your peak concurrency before launch.
- Each synthesize request is limited to 5,000 characters; agent replies are normally far shorter, but cap them anyway.
- Latency depends on your region and connection. Measure from your own servers, using the method described on the developers page, before you commit.
- ReadAloud speaks English today. More languages are on the roadmap, and the site says so rather than list languages it cannot yet do well.
- ReadAloud does not claim HIPAA, SOC 2, GDPR or other certifications. If you need one of them, ask the vendors you are considering what they can show you.
What ReadAloud costs
- ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
- ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.
Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).
Common questions
- Can callers interrupt the agent?
- Yes. Send {"type":"stop"} on the open WebSocket. The server confirms with a cancelled message, and a cancelled request is not billed.
- What audio formats work for phone calls?
- On ReadAloud Live you can request mulaw_8000, alaw_8000 or pcm_8000 per synthesize message. The default is pcm_24000, which suits web audio.
- Does ReadAloud work with Vapi?
- There is a custom-voice endpoint at /v1/vapi/custom-voice that returns raw PCM at the requested sample rate and accepts your key as x-vapi-secret. One real Vapi assistant call has been run, with the inline secret; the Vapi integration page lists what was and was not tested, so read it before relying on the setup.
- Is there more than one voice for agents?
- ReadAloud Live has one American English voice, readaloud-default. ReadAloud Studio offers a broader list of English voices but is positioned for read-aloud content rather than live calls.
- How do I test without paying?
- Every account gets free credits worth $0.10, roughly 10,000 characters, shared across all of its keys. When they run out the API returns 402 until a payment method is added.
Something here does not match what you see? Tell us at support@readaloudai.org.