How it works
- On the server, extract the readable text of a page: headings, paragraphs and list items, without navigation, scripts or hidden elements. Keep the text identical to what is visible.
- Split the text at paragraph boundaries into pieces under 5,000 characters, since that is the per-request limit.
- For each piece, call POST https://api.readaloudai.org/v1/audio/speech with Authorization: Bearer <key>, response_format mp3 and a voice id. Use readaloud-default for Live or an English voice id from GET /v1/voices for Studio.
- Store the audio files under a key made from the page URL and a hash of the text, so any edit to the article invalidates the cached audio.
- Render a native audio element, or a small player, with a visible label, keyboard focus, play and pause, and a speed control. Test with a screen reader.
- Pre-generate audio when content is published instead of when a visitor clicks, so the first listener does not wait.
- Show a short note that the voice is synthetic, and keep the text version one click away.
A sample text
An illustrative example written for this page. It is the text you would send to the API, not a recording, and it makes no claim about how the audio sounds.
How to reset your password. First, open Settings and choose Security. Next, select Change password and enter your current password. Finally, enter a new password twice and save.
Choose ReadAloud Studio when the page is long-form prose and you want a broader English voice list, or Live for short help pages. Request mp3 through the OpenAI-compatible route in one-shot mode and cache the result; streaming adds nothing for pre-generated audio.
When ReadAloud fits
- You publish English content and want an optional listen control without running your own speech software.
- You can generate audio at publish time and cache it, which keeps cost proportional to new content rather than page views.
- You already use an OpenAI client library and want a base-URL change rather than a new integration.
- You want a straightforward per-character price: Live is about $0.004 and Studio about $0.01 per 1,000 characters.
When another service may suit you
- Your audience reads many languages; look for a provider that lists supported languages and voices for each, and test the ones you publish in.
- You need word-by-word highlighting synced to audio; ReadAloud's compatibility routes return no alignment data, so you would need a service that supplies timestamps.
- You want a turnkey widget with analytics and a hosted player instead of building one.
- You need the browser to read text locally with no server round trip; the built-in browser speech feature may be enough.
Features used
- OpenAI-compatible speech route: Server code can use an existing OpenAI client with a changed base URL and returns mp3 for an audio element.
- Studio voice list: ReadAloud Studio offers a broader set of English voices than Live's single voice; GET /v1/voices lists the ids.
- mp3 and opus output: mp3 plays in the audio element of every common browser; Ogg Opus support varies by browser, so test it on the ones you target.
- 5,000-character requests: Paragraph-level chunking keeps each request within the documented limit.
- Server-side keys: Clients take your API key, so synthesis runs on your server and the browser only receives audio.
Things to consider
- Conformance to WCAG or any other accessibility standard is the site owner's responsibility. A listen button does not replace semantic HTML, captions, contrast or keyboard support.
- Do not intercept or disable a visitor's own screen reader. Offer the audio as an extra, not a replacement.
- Disclose that the voice is synthetic where your policy or local rules require it.
- Keep spoken text identical to visible text, and re-generate audio whenever an article is edited.
- ReadAloud speaks English today. More languages are on the roadmap, and the site says so rather than list languages it cannot yet do well.
- ReadAloud does not claim HIPAA, SOC 2, GDPR or other certifications. If you need one of them, ask the vendors you are considering what they can show you.
What ReadAloud costs
- ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
- ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.
Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).
Common questions
- Can I call ReadAloud directly from the browser?
- No. The clients take your API key, so keep it on a server. Have the server return audio or a URL to a stored file.
- Does this make my site WCAG compliant?
- No. Audio can help some readers, but compliance depends on markup, contrast, navigation and many other factors that you must assess yourself.
- How do I handle a very long article?
- Split it at paragraph boundaries under 5,000 characters, request each part, and play them in order. The audiobook endpoints chain chapters, but they use a session token or the MCP tools rather than a plain API key.
- Can the page highlight words as they are spoken?
- Not with the data ReadAloud returns on its compatibility routes. Alignment fields are null there, so you would have to estimate timing yourself.
- What if a visitor needs another language?
- Only describe languages you have tested. The documentation commits to English for Live and Studio voices, so this page does not promise others.
Something here does not match what you see? Tell us at support@readaloudai.org.