Install
No package. 1) Create a ReadAloud key (https://readaloudai.org/developers) and run the curl check below. 2) In Vapi, create a saved assistant whose voice block is:
{
"provider": "custom-voice",
"server": {
"url": "https://api.readaloudai.org/v1/vapi/custom-voice?voice=readaloud-default",
"secret": "<your ReadAloud key>",
"timeoutSeconds": 30
}
}
Vapi sends the secret as the X-VAPI-SECRET header. Vapi also supports Custom Credentials (Bearer Token) for the same purpose; the endpoint accepts Authorization: Bearer, but we have not run that flow in Vapi. Vapi's own page on custom TTS is the reference for the assistant fields.Quickstart
curl -s -D - -X POST 'https://api.readaloudai.org/v1/vapi/custom-voice' \
-H "x-vapi-secret: $READALOUD_API_KEY" \
-H 'Content-Type: application/json' \
-d '{"message":{"type":"voice-request","text":"Hello from ReadAloud.","sampleRate":16000,"timestamp":1}}' \
-o out.pcm
# out.pcm is raw 16-bit signed little-endian mono PCM at 16000 Hz, no header
ffmpeg -f s16le -ar 16000 -ac 1 -i out.pcm out.wavWhat was tested
Checked with ReadAloud /v1/vapi/custom-voice endpoint, called with curl using Vapi-shaped voice-request bodies; one real Vapi call recorded in the ReadAloud repository. Version tested: live API, gateway release v60. Date: October 10, 2026. This is one dated check against the live API. Other versions, and later releases, may behave differently.
What worked
- Our own run on 10 October 2026: a Vapi-shaped voice-request POST with x-vapi-secret returned 200 with content type application/octet-stream at sampleRate 8000 (50,714 bytes), 16000 (103,284 bytes) and 24000 (151,582 bytes). Each body had an even byte count and decoded to about 3.2 seconds of mono audio at its rate.
- Authorization: Bearer with the same key also returned 200. A wrong key returned 401, an unlisted sampleRate such as 12345 returned 422, and an unknown voice returned 404, each with a JSON body whose top-level error field is a string.
- Blank text returned 200 with 20 ms of silence (640 bytes at 16000 Hz) and an x-vapi-blank-text: silence header, so a blank line does not read as a voice failure.
- Recorded in the ReadAloud repository's Vapi docs, not re-run by us: one real assistant with the custom-voice provider, the inline secret and no fallback voice was called by an automated caller. The call lasted 24 seconds and ended normally, the assistant spoke five lines, and Vapi reported custom-voice latency of 266 ms and 493 ms on the two turns it measured.
Limits and what is not supported
- Only one real Vapi call has been run, with the inline secret. The Custom Credential (credentialId) flow, the sample rate Vapi actually requested, and Vapi's behavior on 429 responses and timeouts have not been tested.
- Nobody has listened to the phone-line audio of that call, so we make no statement about how it sounds over a phone.
- ReadAloud Live has one English voice, readaloud-default. Each request is limited to 5,000 characters and the request body to 1 MiB.
- The curl runs checked the response format and status codes. We did not run a load test; the repository notes parallel requests were checked only up to 14.
Limits that apply to every integration
- 5,000 characters per request. Split longer text and send it in order.
- Output formats on the OpenAI-compatible route: mp3 (default, 24 kHz), opus (48 kHz, Ogg), wav (24 kHz) and pcm (24 kHz, 16-bit, mono); aac and flac return 400.
- No SSML and no audio tags.
- Beyond capacity the API returns 429 (HTTP) or an "at capacity" error with close code 1013 (WebSocket); retry with a short backoff.
What ReadAloud costs
- ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
- ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.
Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).
Common questions
- Is there anything to install?
- No. You paste the endpoint URL and your ReadAloud key into a Vapi assistant's custom-voice settings. The curl command on this page checks the endpoint first.
- How is the voice chosen?
- Vapi's request does not name a voice, so the URL does. With no query string the endpoint uses readaloud-default; adding ?voice=readaloud-default makes it explicit. An unknown voice returns 404.
- Why use a saved assistant?
- Vapi's authentication documentation says a Custom Credential is attached only when the server URL comes from your saved organization configuration, not from a URL inside an API request body. The inline secret was what the recorded real call used.
- Should I set a fallback voice?
- Vapi's documentation says that without a fallback plan a failed voice request ends the call. The recorded real call used no fallback and had no failures, but one call does not show how often requests fail, so a fallback is a reasonable precaution.
Something here does not match what you see? Tell us at support@readaloudai.org.