Migrate from ElevenLabs
The ReadAloud API serves the ElevenLabs text-to-speech routes, so code written for the ElevenLabs SDKs runs against https://api.readaloudai.org after you change three values. This page lists what we ran against the live service, what behaves differently, and what is not supported.
Three things to change
The earlier names piper-default, piper and kokoro keep working as aliases.
- Base URL:
https://api.readaloudai.orginstead ofhttps://api.elevenlabs.io. - API key: your ReadAloud key (it starts with
rtts_). Send it asxi-api-keyor asAuthorization: Bearer; both work. - Voice id: use
readaloud-default. ElevenLabs voice ids, for example21m00Tcm4TlvDq8ikWAM, do not exist here and return404 voice_not_found. We do not substitute a voice for you, and there is deliberately no table mapping ElevenLabs voices to ours.
Everything else in a typical text-to-speech call, such as model_id, output_format and voice_settings.speed, can stay as it is. readaloud-default is the only voice this guide covers. It does not sound like any ElevenLabs voice, and we make no claim that it matches ElevenLabs quality. Listen to it on your own text before you switch.
Before and after
curl
# before
curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/21m00Tcm4TlvDq8ikWAM?output_format=mp3_44100_128" \
-H "xi-api-key: $ELEVENLABS_API_KEY" -H "Content-Type: application/json" \
-d '{"text":"Thanks for calling.","model_id":"eleven_flash_v2_5"}' -o out.mp3
# after: new host, new key, new voice id
curl -X POST "https://api.readaloudai.org/v1/text-to-speech/readaloud-default?output_format=mp3_44100_128" \
-H "xi-api-key: $READALOUD_API_KEY" -H "Content-Type: application/json" \
-d '{"text":"Thanks for calling.","model_id":"eleven_flash_v2_5"}' -o out.mp3Python
# pip install elevenlabs (tested with 2.71.0)
import os
from elevenlabs.client import ElevenLabs
# before
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
# after: new key and one extra argument, base_url
client = ElevenLabs(
api_key=os.environ["READALOUD_API_KEY"],
base_url="https://api.readaloudai.org",
)
audio = client.text_to_speech.convert(
voice_id="readaloud-default", # was an ElevenLabs voice id
text="Thanks for calling.",
model_id="eleven_multilingual_v2", # accepted; the voice decides how audio is made
output_format="mp3_44100_128",
)
with open("out.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)JavaScript
// npm i @elevenlabs/elevenlabs-js (tested with 2.71.0)
import fs from "node:fs";
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
// before
// const client = new ElevenLabsClient({ apiKey: process.env.ELEVENLABS_API_KEY });
// after: new key and one extra option, baseUrl
const client = new ElevenLabsClient({
apiKey: process.env.READALOUD_API_KEY,
baseUrl: "https://api.readaloudai.org",
});
const audio = await client.textToSpeech.convert("readaloud-default", {
text: "Thanks for calling.",
modelId: "eleven_flash_v2_5",
outputFormat: "pcm_16000",
});
const chunks = [];
for await (const chunk of audio) chunks.push(chunk);
fs.writeFileSync("out.pcm", Buffer.concat(chunks)); // raw 16-bit mono, 16 kHzWebSocket
The Python SDK’s convert_realtime works against the live API. The current JavaScript SDK has no text-to-speech WebSocket client, so there is no JavaScript SDK result for this route; the raw protocol below was exercised with a plain WebSocket client.
# Python SDK, WebSocket (stream-input). Needs wss://, which api.readaloudai.org serves.
def text_chunks():
yield "Thanks for calling. "
yield "How can I help you today?"
with open("out.mp3", "wb") as f:
for chunk in client.text_to_speech.convert_realtime(
voice_id="readaloud-default",
text=text_chunks(),
model_id="eleven_flash_v2_5",
output_format="mp3_44100_128",
):
f.write(chunk)wss://api.readaloudai.org/v1/text-to-speech/readaloud-default/stream-input?output_format=mp3_44100_128
header: xi-api-key: <your ReadAloud key> (or send "xi_api_key" in the first message)
send: {"text": " "} first message
send: {"text": "Thanks for calling. "} then text, as it arrives
send: {"text": ""} end of input
receive: {"audio": "<base64>", ...} ... then {"isFinal": true} and a normal closeText is buffered and spoken at sentence boundaries, on flush, and at end of input. The multi-stream-input route (several contexts on one socket) also worked in our run. opus_* output is rejected on WebSocket routes. Word-level timing is not provided.
What we tested
On 10 October 2026 we ran each item below against the live API at api.readaloudai.org, with the official SDKs and with plain HTTP and WebSocket clients, using a temporary key that has since been revoked. Only items that worked are in this table. Differences are listed in the next section.
SDKs: elevenlabs (Python) 2.71.0 on Python 3.14, and @elevenlabs/elevenlabs-js 2.71.0 on Node 26. Only the base URL, the key and the voice id were changed from the ElevenLabs code.
| Route or call | Result in our run |
|---|---|
POST /v1/text-to-speech/{voice_id} | curl, Python text_to_speech.convert and JavaScript textToSpeech.convert all returned audio for mp3_44100_128, pcm_16000 and ulaw_8000. Checked with ffprobe: the MP3 was 44.1 kHz mono. Also returned audio for wav_24000 (a valid RIFF/WAVE file) and opus_48000_64 (Ogg Opus, 48 kHz). |
POST .../{voice_id}/stream | Python text_to_speech.stream and JavaScript textToSpeech.stream returned audio in many chunks. |
POST .../with-timestamps | Returns JSON with base64 audio. alignment and normalized_alignment are always null. |
WS .../stream-input | Python convert_realtime over wss:// returned a playable MP3. A plain client also worked with the key in the header and with the key in the first message. |
WS .../multi-stream-input | Plain client only: one context with flush, close_context and close_socket worked. |
GET /v1/voices, /v2/voices, /v1/voices/{id} | Python voices.get_all, voices.search, voices.get and the JavaScript equivalents worked with the ElevenLabs response shape. These routes use the ElevenLabs shape when the key is sent as xi-api-key. |
GET /v1/models | Both SDKs listed three model ids: eleven_flash_v2_5, eleven_turbo_v2_5, eleven_multilingual_v2. |
GET /v1/user/subscription | Both SDKs parsed the response. For a pay-as-you-go key character_limit is a very large placeholder number, not a real cap. |
voice_settings.speed | Changes the length of the audio: the same sentence came out about 3.3 s at 0.7, 2.7 s at 1.0 and 2.1 s at 1.5. Speed 0 is rejected with 422. |
| Error classes | Both SDKs raised their API error with status 404 (unknown voice), 401 (bad key), 400 (text too long) and 422 (empty text). |
| Auth headers | xi-api-key and Authorization: Bearer both worked on text to speech. |
Not tested against the live service: the other output_format names beyond those listed above (an unknown name returns a 422 that lists the names the API accepts, which follow the ElevenLabs list), the ElevenLabs deprecated npm package, voice-agent framework plugins for ElevenLabs (LiveKit, Pipecat), non-English text, and the 402 response when free credits run out. For LiveKit and Pipecat, use the ReadAloud plugins.
What is different or not supported
Different
- Voices. ElevenLabs voice ids return
404. Usereadaloud-default. Your own ElevenLabs voice clones and Voice Library voices cannot be used here. - Sound. The voice is not the same as any ElevenLabs voice and we do not claim equal quality. We publish no side-by-side audio on this page.
model_iddoes not choose the voice. Any string is accepted, including ids we do not list such aseleven_v3; the voice you pass decides how the audio is made. The response headerX-Model-Mappedreports how the id was treated.- Voice settings. Only
speedhas an effect.stability,similarity_boost,styleanduse_speaker_boostare accepted and ignored, as areseed,language_code,previous_text,next_textand a few other request fields. The ones you sent are echoed in thex-compat-ignoredresponse header, so nothing is dropped silently. - Language.
language_codedoes not select a language.readaloud-defaultis an American English voice. We did not test other languages through this route. - Sample rates above 24 kHz are upsampled. The audio is made at 24 kHz. Formats such as
mp3_44100_128have the sample rate you asked for but no content above 12 kHz. Response headersX-Sample-Rate,X-Audio-FormatandX-Source-Sample-Ratesay so. wav_*formats work only on the non-streaming endpoint (the streaming endpoint returns422), as in ElevenLabs.- Timestamps.
/with-timestampsworks but returnsnullalignment. We do not compute character timings. - Markup. No SSML and no v3 audio tags. According to the gateway code, text is passed through as written, so a tag such as
[laughs]may be read aloud. We did not test that live. - Quota status code. When free credits run out the API returns
402, not the401ElevenLabs uses. We did not trigger this in the live run.
Not supported on the ElevenLabs paths
These paths returned 404 in our run (or, for voice creation, a ReadAloud error rather than an ElevenLabs-style result), so do not expect the ElevenLabs SDK methods for them to work against ReadAloud:
- Speech to text (Scribe),
/v1/speech-to-text. ReadAloud has its own transcription API with a different interface; see the Voice API page. - Dubbing,
/v1/dubbing. ReadAloud’s dubbing is a separate API, not this one. - The Agents platform (Conversational AI),
/v1/convai/.... - Voice cloning with ElevenLabs semantics,
POST /v1/voices/add(instant and professional voice clones). Sending it with anxi-api-keyheader returned a ReadAloud authentication error, not a clone. ReadAloud cloning has its own API and needs a payment method. - Speech to speech, text to voice (voice design), audio isolation, sound generation, history, pronunciation dictionaries and projects.
- Not tested: the deprecated ElevenLabs npm package and the ElevenLabs plugins for voice-agent frameworks, see above.
We checked a list of common ElevenLabs paths, not every path in their API. Anything not in the table above is not verified.
Error handling
Errors use the ElevenLabs shape, so the SDK error classes work unchanged. These are real responses from our run.
# 404, unknown voice id (an ElevenLabs voice id)
{"detail":{"status":"voice_not_found","message":"Voice 21m00Tcm4TlvDq8ikWAM not found. ..."}}
# 401, missing or revoked key
{"detail":{"status":"invalid_api_key","message":"Invalid API key"}}
# 400, text over 5,000 characters
{"detail":{"status":"max_character_limit_exceeded","message":"text too long (max 5000 characters per request)"}}
# 422, empty text. Validation errors are a list, as in ElevenLabs' own API
{"detail":[{"loc":["body","text"],"msg":"text must be a non-empty string","type":"value_error"}]}| Status | detail.status | Meaning |
|---|---|---|
| 401 | invalid_api_key | Missing or revoked key. Seen live. |
| 402 | quota_exceeded | Free credits used up, or the request is larger than what remains. Add a payment method. Not triggered in our run. |
| 404 | voice_not_found | Unknown voice id. Seen live. |
| 400 | max_character_limit_exceeded | More than 5,000 characters. Seen live. |
| 422 | (a list) | Validation: empty or missing text, malformed JSON, an unknown output_format, speed of 0 or less. Seen live. |
| 429 | too_many_concurrent_requests | The service is at capacity. A Retry-After header is set. We saw this once during the run, on a different voice, and the retry a few seconds later succeeded. |
| 405 | method_not_allowed | A GET on a text-to-speech path. Use POST. |
Limits and credits
- 5,000 characters per request. A 5,000-character request succeeded; 5,001 returned
400. According to the gateway code the same limit applies to each generation on a WebSocket; we did not test that live. - Long text takes time to finish. In our run a 5,000-character request returned its first byte in about 0.2 s and the last byte after about 36 s (roughly 5.7 minutes of audio). Split long text into several requests, or use the streaming endpoint and play as you receive.
- Concurrency. Six simultaneous requests all succeeded. We do not publish a concurrency number on this page; the service returns
429withRetry-Afterwhen it is full. - Free credits. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech), shared across all of your keys. When they run out, requests return
402until you add a payment method; after that you pay as you go. See free credits and the pricing on the Voice API page.
FAQ
Do I only need to change the base URL?
No, three values: the base URL, the key and the voice id. Without a ReadAloud voice id the call returns 404.
Can I reuse my ElevenLabs voices?
No. ElevenLabs voice ids, including your own clones, do not exist here. We do not map them to ReadAloud voices.
Will it sound the same?
No. It is a different voice. We do not claim parity, so judge it on your own text.
Do I need to remove stability and similarity_boost?
No. They are accepted and ignored, and the x-compat-ignored response header tells you which ones were ignored. Only speed has an effect.
Does streaming work?
Yes: the /stream endpoint and the stream-input WebSocket both worked in our run.
Can I use ElevenLabs speech to text, dubbing or agents through this API?
No. Only text to speech is covered. See the list above.
What does it cost?
New accounts get a one-time grant of free credits; after that you pay as you go. Current prices are on the Voice API page.
Something here does not match what I see.
Tell us at support@readaloudai.org. This page records one dated test run, not a guarantee, and we will correct it.
ReadAloud AI is not affiliated with ElevenLabs. ElevenLabs and its product names belong to their owner and are used here only to describe compatibility.