How it works
- Rewrite your bot's system prompt for speech: replies of one to three sentences, no lists, no URLs, no markdown, and numbers spelled the way they should be said.
- Pick a framework or build directly. With Pipecat, pip install pipecat-readaloud and add ReadAloudHttpTTSService; with LiveKit Agents, pip install livekit-plugins-readaloud and pass readaloud.TTS() to AgentSession.
- If you call the API yourself, authorize with POST https://api.readaloudai.org/tts/authorize and engine live, open the WebSocket, and send each sentence as a synthesize message with voice default.
- Cancel speech when the customer interrupts: send {"type":"stop"} and discard audio still queued.
- Add a handoff rule: after a second failed answer, or a direct request, say a short handoff line and transfer to a person with the conversation summary.
- Log the text of each spoken reply, not the audio, so reviewers can read what the bot said and spot wrong answers.
- Test with real support transcripts. Check reading of order numbers, email addresses and amounts, and rewrite the prompt for any that are awkward.
A sample text
An illustrative example written for this page. It is the text you would send to the API, not a recording, and it makes no claim about how the audio sounds.
I have reopened your case. Your reference number is A, four, seven, two, nine. I will say that once more. A, four, seven, two, nine. A specialist will reply within one working day.
Use ReadAloud Live over the WebSocket, pcm_24000 for web or mulaw_8000 for phone. Send sentence by sentence so the customer hears the first one while the next is being written. Spell out reference numbers.
When ReadAloud fits
- Your bot already works in text and you want to add an English voice without changing the rest of the stack.
- You use Pipecat or LiveKit Agents and want a maintained plugin.
- Cost per answer matters: 300 characters of reply is about $0.0012 on Live.
- A single voice without personality controls is acceptable for factual answers.
When another service may suit you
- Customers write or speak in many languages; look for documented per-language voices and test your top languages with fluent reviewers.
- Your brand needs a recognisable voice, or empathetic delivery in difficult conversations; ReadAloud does not control emotion.
- You want one vendor to run telephony, recognition, language model and speech with a service-level agreement.
- Regulated conversations, such as payments or health data, need audited vendors and recorded consent flows, which you must set up and check.
Features used
- WebSocket streaming: Replies play sentence by sentence, so the customer hears the beginning while the model finishes.
- stop message: Barge-in is a single message, and cancelled requests are not billed.
- Pipecat and LiveKit plugins: Maintained packages with an MIT licence that read READALOUD_API_KEY.
- Telephony output: 8 kHz mu-law, A-law and pcm are available per request.
- Per-character pricing: About $0.004 per 1,000 characters on Live, so pricing follows reply length.
Things to consider
- Tell customers they are talking to an automated assistant where law or policy requires it, and offer a route to a person.
- Do not record or store calls without the consents your jurisdiction requires. This is separate from the speech provider.
- Do not read out sensitive data such as full card numbers or health details.
- Keep a text channel as a fallback for outages, 402 responses and capacity errors.
- ReadAloud speaks English today. More languages are on the roadmap, and the site says so rather than list languages it cannot yet do well.
- ReadAloud does not claim HIPAA, SOC 2, GDPR or other certifications. If you need one of them, ask the vendors you are considering what they can show you.
What ReadAloud costs
- ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
- ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.
Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).
Common questions
- Does ReadAloud include speech recognition for live calls?
- Its speech to text is batch only, with no live streaming transcription. For live calls, use a separate streaming recognizer.
- Can the bot sound apologetic or warm?
- Not through settings. Style instructions are ignored. Write empathetic wording, and keep it short.
- How should the bot read an email address?
- Spell it out in the text, for example 'sam at example dot com', since the engine reads exactly what it receives.
- What happens when the service is at capacity?
- You get an 'at capacity' error and a 1013 close on the WebSocket. Retry with a brief backoff and have a text fallback.
- Can I measure latency before committing?
- Yes. The developers page describes a method and its results; repeat it from your own region with your own sentences.
Something here does not match what you see? Tell us at support@readaloudai.org.