Use ReadAloud as a custom voice in Vapi

Vapi is a platform for building voice assistants that answer phone calls and web calls. Besides its built-in voice providers, it can call a server you control for every spoken line: with the custom-voice provider, Vapi POSTs the text to a URL and plays back the raw audio the server returns. ReadAloud serves that contract at https://api.readaloudai.org/v1/vapi/custom-voice, so ReadAloud Live can be the voice of a Vapi assistant without any code on your side. The request is Vapi's documented voice-request JSON, and the reply is a headerless stream of 16-bit signed little-endian mono PCM at exactly the sample rate Vapi asks for. The endpoint accepts 8000, 16000, 22050, 24000 and 44100 Hz and authenticates with the same ReadAloud key you use everywhere else, sent as x-vapi-secret or as a Bearer token. This is a different kind of integration from the SDK pages here: there is no package to install, only an assistant configuration. We checked the endpoint ourselves on 10 October 2026 with Vapi-shaped curl requests, and the ReadAloud repository records one real Vapi assistant call. The sections below separate those two sources and list what has not been tested.

Install

No package. 1) Create a ReadAloud key (https://readaloudai.org/developers) and run the curl check below. 2) In Vapi, create a saved assistant whose voice block is:

{
  "provider": "custom-voice",
  "server": {
    "url": "https://api.readaloudai.org/v1/vapi/custom-voice?voice=readaloud-default",
    "secret": "<your ReadAloud key>",
    "timeoutSeconds": 30
  }
}

Vapi sends the secret as the X-VAPI-SECRET header. Vapi also supports Custom Credentials (Bearer Token) for the same purpose; the endpoint accepts Authorization: Bearer, but we have not run that flow in Vapi. Vapi's own page on custom TTS is the reference for the assistant fields.

Quickstart

Quickstart (bash)
curl -s -D - -X POST 'https://api.readaloudai.org/v1/vapi/custom-voice' \
  -H "x-vapi-secret: $READALOUD_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"message":{"type":"voice-request","text":"Hello from ReadAloud.","sampleRate":16000,"timestamp":1}}' \
  -o out.pcm
# out.pcm is raw 16-bit signed little-endian mono PCM at 16000 Hz, no header
ffmpeg -f s16le -ar 16000 -ac 1 -i out.pcm out.wav

What was tested

Checked with ReadAloud /v1/vapi/custom-voice endpoint, called with curl using Vapi-shaped voice-request bodies; one real Vapi call recorded in the ReadAloud repository. Version tested: live API, gateway release v60. Date: October 10, 2026. This is one dated check against the live API. Other versions, and later releases, may behave differently.

What worked

  • Our own run on 10 October 2026: a Vapi-shaped voice-request POST with x-vapi-secret returned 200 with content type application/octet-stream at sampleRate 8000 (50,714 bytes), 16000 (103,284 bytes) and 24000 (151,582 bytes). Each body had an even byte count and decoded to about 3.2 seconds of mono audio at its rate.
  • Authorization: Bearer with the same key also returned 200. A wrong key returned 401, an unlisted sampleRate such as 12345 returned 422, and an unknown voice returned 404, each with a JSON body whose top-level error field is a string.
  • Blank text returned 200 with 20 ms of silence (640 bytes at 16000 Hz) and an x-vapi-blank-text: silence header, so a blank line does not read as a voice failure.
  • Recorded in the ReadAloud repository's Vapi docs, not re-run by us: one real assistant with the custom-voice provider, the inline secret and no fallback voice was called by an automated caller. The call lasted 24 seconds and ended normally, the assistant spoke five lines, and Vapi reported custom-voice latency of 266 ms and 493 ms on the two turns it measured.

Limits and what is not supported

  • Only one real Vapi call has been run, with the inline secret. The Custom Credential (credentialId) flow, the sample rate Vapi actually requested, and Vapi's behavior on 429 responses and timeouts have not been tested.
  • Nobody has listened to the phone-line audio of that call, so we make no statement about how it sounds over a phone.
  • ReadAloud Live has one English voice, readaloud-default. Each request is limited to 5,000 characters and the request body to 1 MiB.
  • The curl runs checked the response format and status codes. We did not run a load test; the repository notes parallel requests were checked only up to 14.

Limits that apply to every integration

  • 5,000 characters per request. Split longer text and send it in order.
  • Output formats on the OpenAI-compatible route: mp3 (default, 24 kHz), opus (48 kHz, Ogg), wav (24 kHz) and pcm (24 kHz, 16-bit, mono); aac and flac return 400.
  • No SSML and no audio tags.
  • Beyond capacity the API returns 429 (HTTP) or an "at capacity" error with close code 1013 (WebSocket); retry with a short backoff.

What ReadAloud costs

  • ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
  • ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.

Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).

Common questions

Is there anything to install?
No. You paste the endpoint URL and your ReadAloud key into a Vapi assistant's custom-voice settings. The curl command on this page checks the endpoint first.
How is the voice chosen?
Vapi's request does not name a voice, so the URL does. With no query string the endpoint uses readaloud-default; adding ?voice=readaloud-default makes it explicit. An unknown voice returns 404.
Why use a saved assistant?
Vapi's authentication documentation says a Custom Credential is attached only when the server URL comes from your saved organization configuration, not from a URL inside an API request body. The inline secret was what the recorded real call used.
Should I set a fallback voice?
Vapi's documentation says that without a fallback plan a failed voice request ends the call. The recorded real call used no fallback and had no failures, but one call does not show how often requests fail, so a fallback is a reasonable precaution.

Something here does not match what you see? Tell us at support@readaloudai.org.