How it works
- On publish, take the final article text, not the rendered web page. Remove link lists, footers, boilerplate and image credits, and write out abbreviations you want read in full.
- Add a one-line spoken header such as the title and the date, then the body, so the audio stands alone.
- Split the text at paragraph breaks into parts below 5,000 characters, the per-request limit.
- For each part, POST https://api.readaloudai.org/v1/audio/speech with a Studio voice id from GET /v1/voices, response_format mp3 and your key as a Bearer token.
- Concatenate the parts in order, for example with ffmpeg, and write the file to object storage with a stable URL.
- Link the file from the article page and from your RSS feed as an enclosure, with text noting that the narration is synthetic.
- Log characters per issue so you can compare spend with listens, and re-render only when the source text changes.
A sample text
An illustrative example written for this page. It is the text you would send to the API, not a recording, and it makes no claim about how the audio sounds.
This week in the Harbor Report: the city council delays a vote on the ferry schedule, two local bakeries merge, and a short guide to winter parking rules. First, the ferry.
Use ReadAloud Studio and mp3. Newsletters are batch jobs, so render each part with a one-shot request and join them afterwards. The HTTP route is simpler than a socket here.
When ReadAloud fits
- You publish regular English text and want an audio edition without scheduling recording time.
- Cost scales with characters: a 6,000-character issue is about $0.06 on Studio at list price.
- You prefer to run the pipeline yourself, with cleanup, chunking and hosting under your control.
- You want a voice choice from several English voices rather than a single fixed voice.
When another service may suit you
- You need a named, recognisable host voice with a performance; hire a narrator or look for a service with performance-focused voices.
- Your audience reads multiple languages; look for a vendor that documents per-language voices and test samples in each.
- You need chapter markers and ready-made podcast hosting; ReadAloud has an audiobook workflow, but it has not been tested end to end on a long book, so check it yourself before relying on it.
- Your text includes a lot of tables, code or formulas that do not read well aloud.
Features used
- ReadAloud Studio voices: A broader list of English voices, available through the same API as Live.
- OpenAI-compatible speech route: One POST per chunk with an existing client library, returning mp3.
- mp3 output at 24 kHz: A standard format that podcast apps, RSS enclosures and email links accept.
- Per-character pricing: About $0.01 per 1,000 characters on Studio, billed per request with no plan to manage.
- Audiobook endpoints: POST /api/audiobooks chains chapters and exports one mp3 with chapter markers, through a session token or the MCP tools.
Things to consider
- Check that you hold the rights to the text you convert, including guest writers and quoted material, before publishing audio.
- Tell listeners that the narration is synthetic. Some platforms and regions also require that.
- Review a rendered issue by ear. Product names, acronyms and non-English words may be mispronounced and need respelling in the source.
- Requests are limited to 5,000 characters, so splitting and joining is your code's job on the plain speech routes.
- ReadAloud speaks English today. More languages are on the roadmap, and the site says so rather than list languages it cannot yet do well.
- ReadAloud does not claim HIPAA, SOC 2, GDPR or other certifications. If you need one of them, ask the vendors you are considering what they can show you.
What ReadAloud costs
- ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
- ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.
Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).
Common questions
- Can I get an RSS feed with audio?
- You build the feed yourself. ReadAloud returns audio files; you host them and reference them in an enclosure tag.
- Which tier should a newsletter use?
- Studio, in most cases, since latency is irrelevant and it offers more English voices. Live is cheaper if one voice is acceptable.
- How do I stop the audio sounding choppy at chunk joins?
- Split on paragraph boundaries, not mid-sentence, and keep a short silence between parts when you join them.
- Does it read links and images?
- It reads whatever text you send. Strip URLs and captions in your cleanup step or they will be spoken.
- Is there a limit on issue length?
- Each request is capped at 5,000 characters, but you can send as many requests as you like, subject to your account and the server's capacity.
Something here does not match what you see? Tell us at support@readaloudai.org.