Spoken alerts and notifications from your own backend events

Most alerts are text that nobody reads in time. A pager message gets buried, a dashboard sits on a wall screen with the sound off, a warehouse worker has gloves on and cannot look at a tablet. Speaking the alert changes who notices it. The engineering problem is small but specific: events arrive in bursts, each message is short, wording must stay consistent, and you rarely want to keep a connection open all day for a message that fires twice an hour. That shape favours plain HTTP requests over a long-lived socket. Your service builds a sentence from structured fields, sends it to the speech API, and plays or forwards the returned audio. Because the text is generated from templates, you can also cache clips for messages that repeat. ReadAloud Live charges per character, so a forty-character alert costs a fraction of a cent. The limits to plan around are a single English voice on Live, no control over urgency or tone, and a capacity ceiling per server that means a burst of alerts should be queued, not fired all at once.

How it works

  1. Define alert templates with named fields, for example 'Disk usage on {host} is above {threshold}.' Keep each message to one or two sentences.
  2. Authorize with POST https://api.readaloudai.org/tts/authorize using engine live, then POST the text to the returned http_url with Authorization: Bearer <token>, or call POST /v1/audio/speech with your key if you prefer the OpenAI-compatible shape.
  3. Request mp3 for browsers and chat attachments or wav where a device needs a plain container. The response is audio you can store, play or attach.
  4. Cache clips by the hash of the final text so a recurring alert is synthesized once and replayed afterwards.
  5. Put events on a queue with a small concurrency limit. On a 429 or 503 response, wait for the Retry-After value or a short backoff and retry.
  6. Deliver the audio to its destination: a speaker, a phone call, a push notification attachment or a kiosk page.
  7. Keep a text fallback. If the speech request fails or credits run out (402), send the written alert as normal.

A sample text

An illustrative example written for this page. It is the text you would send to the API, not a recording, and it makes no claim about how the audio sounds.

Infrastructure alert
Warning. The payment service in the east region has returned errors for five minutes. The on-call engineer has been paged.

Use ReadAloud Live, format mp3, one-shot HTTP request. Live is the cheaper tier and the message is short, so there is no benefit in a socket. Cache the clip if the same wording repeats.

When ReadAloud fits

  • Alerts are short, templated English sentences generated by your own code.
  • You want a plain HTTP request per event instead of managing a streaming connection.
  • Volume is irregular, so paying per character with no plan to manage is a fit, unlike a monthly seat or tier.
  • You can cache repeated clips and queue bursts, which keeps usage predictable.

When another service may suit you

  • You need audibly different urgency levels, such as a calm notice versus an urgent alarm; ReadAloud ignores style instructions.
  • Alerts must be intelligible in noisy plants in several languages; look for documented languages and test on the real speaker hardware.
  • A hard real-time safety function depends on the message; use certified alerting equipment, not a cloud API.
  • You need an offline fallback with no network; look at an on-device speech engine.

Features used

  • HTTP streaming: POST to the http_url from authorize returns audio as each sentence is ready, with X-Sample-Rate and X-Audio-Format headers.
  • OpenAI-compatible route: POST /v1/audio/speech means existing OpenAI client code can generate alerts after a base URL change.
  • Per-character billing: Live is about $0.004 per 1,000 characters, so a short alert costs a small fraction of a cent.
  • Output formats: mp3, opus, wav and pcm cover browsers, chat attachments and embedded players.
  • Clear error codes: 401, 402, 429 and 503 tell your queue whether to retry, stop or fall back to text.

Things to consider

  • A synthesized voice is not a safety system. Anything life-critical needs redundant, certified alerting.
  • Do not put secrets, personal data or customer details into spoken alerts that play in shared rooms.
  • Each server handles a limited number of simultaneous streams; queue alerts and retry with backoff instead of sending bursts.
  • Requests are limited to 5,000 characters. Alerts should be far shorter, which is also better for the listener.
  • ReadAloud speaks English today. More languages are on the roadmap, and the site says so rather than list languages it cannot yet do well.
  • ReadAloud does not claim HIPAA, SOC 2, GDPR or other certifications. If you need one of them, ask the vendors you are considering what they can show you.

What ReadAloud costs

  • ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
  • ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.

Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).

Common questions

Should I use WebSocket or HTTP for alerts?
HTTP is usually simpler. The WebSocket interface suits live conversation where you keep a connection open across many turns; an alert fires once and can be a single request.
Can I vary the tone for critical alerts?
No. ReadAloud does not interpret emotion or style instructions. You can change wording and the speed field, but not the delivery style.
What happens if I run out of credits during an incident?
The API returns 402. Your code should fall back to a written notification so the alert is never dropped.
Can I pre-generate common alerts?
Yes. Render the fixed ones ahead of time and store the files; only render the ones that contain changing values at event time.
Does ReadAloud store the text of my alerts?
Check the privacy page on the site for the current policy. The MCP text_to_speech tool is documented as not saving the text, but this page does not make claims about other routes.

Something here does not match what you see? Tell us at support@readaloudai.org.