Narrate course modules from scripts and re-render when they change

Course content changes. A regulation is updated, a product screen is redesigned, a wrong figure is spotted in lesson four, and every change to a recorded lesson means booking the narrator again or living with a mismatch between slides and audio. Narration generated from a script keeps the two in step: the script is the master, the audio is a build artifact, and an edit is a re-render. That suits a team that maintains many short modules, each with a slide or screen and a few sentences of speech. A workable structure is one script entry per slide, one audio file per entry, and a manifest your course player reads. Because every entry is short, a request stays well under the 5,000-character limit and you re-render only what changed. Studio, with its broader English voice list, is the sensible default for learners who will hear an hour of it. The honest limits are that there is no control over emphasis or delivery, and anything with heavy jargon needs respelling and a listen-through before release.

How it works

  1. Write narration as a script file with one entry per slide or screen, each with an id and the exact words. Keep sentences short and spell out numbers and acronyms the way you want them spoken.
  2. Call GET https://api.readaloudai.org/v1/voices with your key and choose one ReadAloud Studio voice for the whole course, so learners hear a consistent narrator.
  3. For each entry, POST /v1/audio/speech with that voice, response_format mp3 and your Bearer key. Save the response as <slide-id>.mp3.
  4. Write a manifest mapping each slide id to its audio file and the hash of its text. Your player reads the manifest to start audio when a slide appears.
  5. In your build pipeline, compare hashes and re-render only entries whose text changed. Commit scripts and manifests, not the audio, if your storage policy prefers it.
  6. Listen to every file once. Fix mispronunciations by respelling in the script, then re-render that entry.
  7. Add a transcript or captions from the same script so the course does not depend on audio alone.

A sample text

An illustrative example written for this page. It is the text you would send to the API, not a recording, and it makes no claim about how the audio sounds.

Single slide narration
Slide four. When a customer reports a delayed shipment, first confirm the tracking number. Then check the carrier status. If the package is more than three days late, offer a replacement or a refund.

Use ReadAloud Studio with a single English voice for the course, mp3, one request per slide. This is batch work, so skip streaming.

When ReadAloud fits

  • Course text changes often enough that re-recording would be a recurring cost or delay.
  • Your modules are in English and one narrator voice is acceptable.
  • You want the narration pipeline in your build system, with scripts in version control.
  • Short per-slide requests keep each file small and each re-render cheap: a 400-character slide is about $0.004 on Studio.

When another service may suit you

  • You teach in several languages; look for a provider that documents languages and voices separately and let a native speaker review the output.
  • Learner engagement depends on a warm human instructor; use a real presenter.
  • You need controlled emphasis, dramatized scenarios or varied characters; ReadAloud does not offer style control.
  • Your authoring tool has built-in narration you are already satisfied with.

Features used

  • Studio voice list: GET /v1/voices lists English voices, so a course can pick one and keep it.
  • OpenAI-compatible speech route: A short loop over script entries is enough; no special SDK is required.
  • mp3 output: Plays in every common course player and LMS without conversion.
  • Per-character billing: Cost follows script length, so a small edit costs a small re-render.
  • Speed field: speed between 0.25 and 4.0 lets you offer slower narration for dense slides.

Things to consider

  • Rights and consent: if your course teaches a regulated topic, you remain responsible for accuracy; the voice only reads what you write.
  • Disclose synthetic narration where your institution, accreditor or local rules require it.
  • Accessibility standards, including captions and transcripts, remain the course owner's responsibility.
  • Keep voice choice stable across a course; switching mid-course, or after a service change, can be noticeable to learners.
  • ReadAloud speaks English today. More languages are on the roadmap, and the site says so rather than list languages it cannot yet do well.
  • ReadAloud does not claim HIPAA, SOC 2, GDPR or other certifications. If you need one of them, ask the vendors you are considering what they can show you.

What ReadAloud costs

  • ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
  • ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.

Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).

Common questions

Can I use different voices for different characters?
Studio has multiple English voices, so you can assign one per speaker. There is no emotion or style control, so you cannot direct how a line is delivered.
How do I fix a mispronounced term?
Respell it in the script, for example by writing it phonetically, and re-render that slide. There is no pronunciation dictionary on the compatibility routes.
Can narration sync to animations?
Not automatically. You get a file per slide, so time animations to the file length in your player.
Does it support SCORM or xAPI?
No. ReadAloud produces audio files only. Your authoring tool or LMS handles packaging.
Is there a free way to try it?
New accounts get free credits worth $0.10, about 10,000 characters, which is enough for a few slides before adding a payment method.

Something here does not match what you see? Tell us at support@readaloudai.org.