How it works
- Write narration as a script file with one entry per slide or screen, each with an id and the exact words. Keep sentences short and spell out numbers and acronyms the way you want them spoken.
- Call GET https://api.readaloudai.org/v1/voices with your key and choose one ReadAloud Studio voice for the whole course, so learners hear a consistent narrator.
- For each entry, POST /v1/audio/speech with that voice, response_format mp3 and your Bearer key. Save the response as <slide-id>.mp3.
- Write a manifest mapping each slide id to its audio file and the hash of its text. Your player reads the manifest to start audio when a slide appears.
- In your build pipeline, compare hashes and re-render only entries whose text changed. Commit scripts and manifests, not the audio, if your storage policy prefers it.
- Listen to every file once. Fix mispronunciations by respelling in the script, then re-render that entry.
- Add a transcript or captions from the same script so the course does not depend on audio alone.
A sample text
An illustrative example written for this page. It is the text you would send to the API, not a recording, and it makes no claim about how the audio sounds.
Slide four. When a customer reports a delayed shipment, first confirm the tracking number. Then check the carrier status. If the package is more than three days late, offer a replacement or a refund.
Use ReadAloud Studio with a single English voice for the course, mp3, one request per slide. This is batch work, so skip streaming.
When ReadAloud fits
- Course text changes often enough that re-recording would be a recurring cost or delay.
- Your modules are in English and one narrator voice is acceptable.
- You want the narration pipeline in your build system, with scripts in version control.
- Short per-slide requests keep each file small and each re-render cheap: a 400-character slide is about $0.004 on Studio.
When another service may suit you
- You teach in several languages; look for a provider that documents languages and voices separately and let a native speaker review the output.
- Learner engagement depends on a warm human instructor; use a real presenter.
- You need controlled emphasis, dramatized scenarios or varied characters; ReadAloud does not offer style control.
- Your authoring tool has built-in narration you are already satisfied with.
Features used
- Studio voice list: GET /v1/voices lists English voices, so a course can pick one and keep it.
- OpenAI-compatible speech route: A short loop over script entries is enough; no special SDK is required.
- mp3 output: Plays in every common course player and LMS without conversion.
- Per-character billing: Cost follows script length, so a small edit costs a small re-render.
- Speed field: speed between 0.25 and 4.0 lets you offer slower narration for dense slides.
Things to consider
- Rights and consent: if your course teaches a regulated topic, you remain responsible for accuracy; the voice only reads what you write.
- Disclose synthetic narration where your institution, accreditor or local rules require it.
- Accessibility standards, including captions and transcripts, remain the course owner's responsibility.
- Keep voice choice stable across a course; switching mid-course, or after a service change, can be noticeable to learners.
- ReadAloud speaks English today. More languages are on the roadmap, and the site says so rather than list languages it cannot yet do well.
- ReadAloud does not claim HIPAA, SOC 2, GDPR or other certifications. If you need one of them, ask the vendors you are considering what they can show you.
What ReadAloud costs
- ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
- ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.
Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).
Common questions
- Can I use different voices for different characters?
- Studio has multiple English voices, so you can assign one per speaker. There is no emotion or style control, so you cannot direct how a line is delivered.
- How do I fix a mispronounced term?
- Respell it in the script, for example by writing it phonetically, and re-render that slide. There is no pronunciation dictionary on the compatibility routes.
- Can narration sync to animations?
- Not automatically. You get a file per slide, so time animations to the file length in your player.
- Does it support SCORM or xAPI?
- No. ReadAloud produces audio files only. Your authoring tool or LMS handles packaging.
- Is there a free way to try it?
- New accounts get free credits worth $0.10, about 10,000 characters, which is enough for a few slides before adding a payment method.
Something here does not match what you see? Tell us at support@readaloudai.org.