Migrate from ElevenLabs

The ReadAloud API serves the ElevenLabs text-to-speech routes, so code written for the ElevenLabs SDKs runs against https://api.readaloudai.org after you change three values. This page lists what we ran against the live service, what behaves differently, and what is not supported.

Get an API key · Voice API reference

Three things to change

The earlier names piper-default, piper and kokoro keep working as aliases.

  1. Base URL: https://api.readaloudai.org instead of https://api.elevenlabs.io.
  2. API key: your ReadAloud key (it starts with rtts_). Send it as xi-api-key or as Authorization: Bearer; both work.
  3. Voice id: use readaloud-default. ElevenLabs voice ids, for example 21m00Tcm4TlvDq8ikWAM, do not exist here and return 404 voice_not_found. We do not substitute a voice for you, and there is deliberately no table mapping ElevenLabs voices to ours.

Everything else in a typical text-to-speech call, such as model_id, output_format and voice_settings.speed, can stay as it is. readaloud-default is the only voice this guide covers. It does not sound like any ElevenLabs voice, and we make no claim that it matches ElevenLabs quality. Listen to it on your own text before you switch.

Before and after

curl

# before
curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/21m00Tcm4TlvDq8ikWAM?output_format=mp3_44100_128" \
  -H "xi-api-key: $ELEVENLABS_API_KEY" -H "Content-Type: application/json" \
  -d '{"text":"Thanks for calling.","model_id":"eleven_flash_v2_5"}' -o out.mp3

# after: new host, new key, new voice id
curl -X POST "https://api.readaloudai.org/v1/text-to-speech/readaloud-default?output_format=mp3_44100_128" \
  -H "xi-api-key: $READALOUD_API_KEY" -H "Content-Type: application/json" \
  -d '{"text":"Thanks for calling.","model_id":"eleven_flash_v2_5"}' -o out.mp3

Python

# pip install elevenlabs   (tested with 2.71.0)
import os
from elevenlabs.client import ElevenLabs

# before
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

# after: new key and one extra argument, base_url
client = ElevenLabs(
    api_key=os.environ["READALOUD_API_KEY"],
    base_url="https://api.readaloudai.org",
)

audio = client.text_to_speech.convert(
    voice_id="readaloud-default",          # was an ElevenLabs voice id
    text="Thanks for calling.",
    model_id="eleven_multilingual_v2",  # accepted; the voice decides how audio is made
    output_format="mp3_44100_128",
)
with open("out.mp3", "wb") as f:
    for chunk in audio:
        f.write(chunk)

JavaScript

// npm i @elevenlabs/elevenlabs-js   (tested with 2.71.0)
import fs from "node:fs";
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";

// before
// const client = new ElevenLabsClient({ apiKey: process.env.ELEVENLABS_API_KEY });

// after: new key and one extra option, baseUrl
const client = new ElevenLabsClient({
  apiKey: process.env.READALOUD_API_KEY,
  baseUrl: "https://api.readaloudai.org",
});

const audio = await client.textToSpeech.convert("readaloud-default", {
  text: "Thanks for calling.",
  modelId: "eleven_flash_v2_5",
  outputFormat: "pcm_16000",
});
const chunks = [];
for await (const chunk of audio) chunks.push(chunk);
fs.writeFileSync("out.pcm", Buffer.concat(chunks)); // raw 16-bit mono, 16 kHz

WebSocket

The Python SDK’s convert_realtime works against the live API. The current JavaScript SDK has no text-to-speech WebSocket client, so there is no JavaScript SDK result for this route; the raw protocol below was exercised with a plain WebSocket client.

# Python SDK, WebSocket (stream-input). Needs wss://, which api.readaloudai.org serves.
def text_chunks():
    yield "Thanks for calling. "
    yield "How can I help you today?"

with open("out.mp3", "wb") as f:
    for chunk in client.text_to_speech.convert_realtime(
        voice_id="readaloud-default",
        text=text_chunks(),
        model_id="eleven_flash_v2_5",
        output_format="mp3_44100_128",
    ):
        f.write(chunk)
wss://api.readaloudai.org/v1/text-to-speech/readaloud-default/stream-input?output_format=mp3_44100_128

header:   xi-api-key: <your ReadAloud key>      (or send "xi_api_key" in the first message)
send:     {"text": " "}                          first message
send:     {"text": "Thanks for calling. "}       then text, as it arrives
send:     {"text": ""}                           end of input
receive:  {"audio": "<base64>", ...}  ...        then {"isFinal": true} and a normal close

Text is buffered and spoken at sentence boundaries, on flush, and at end of input. The multi-stream-input route (several contexts on one socket) also worked in our run. opus_* output is rejected on WebSocket routes. Word-level timing is not provided.

What we tested

On 10 October 2026 we ran each item below against the live API at api.readaloudai.org, with the official SDKs and with plain HTTP and WebSocket clients, using a temporary key that has since been revoked. Only items that worked are in this table. Differences are listed in the next section.

SDKs: elevenlabs (Python) 2.71.0 on Python 3.14, and @elevenlabs/elevenlabs-js 2.71.0 on Node 26. Only the base URL, the key and the voice id were changed from the ElevenLabs code.

Compatibility results from the 10 October 2026 live run
Route or callResult in our run
POST /v1/text-to-speech/{voice_id}curl, Python text_to_speech.convert and JavaScript textToSpeech.convert all returned audio for mp3_44100_128, pcm_16000 and ulaw_8000. Checked with ffprobe: the MP3 was 44.1 kHz mono. Also returned audio for wav_24000 (a valid RIFF/WAVE file) and opus_48000_64 (Ogg Opus, 48 kHz).
POST .../{voice_id}/streamPython text_to_speech.stream and JavaScript textToSpeech.stream returned audio in many chunks.
POST .../with-timestampsReturns JSON with base64 audio. alignment and normalized_alignment are always null.
WS .../stream-inputPython convert_realtime over wss:// returned a playable MP3. A plain client also worked with the key in the header and with the key in the first message.
WS .../multi-stream-inputPlain client only: one context with flush, close_context and close_socket worked.
GET /v1/voices, /v2/voices, /v1/voices/{id}Python voices.get_all, voices.search, voices.get and the JavaScript equivalents worked with the ElevenLabs response shape. These routes use the ElevenLabs shape when the key is sent as xi-api-key.
GET /v1/modelsBoth SDKs listed three model ids: eleven_flash_v2_5, eleven_turbo_v2_5, eleven_multilingual_v2.
GET /v1/user/subscriptionBoth SDKs parsed the response. For a pay-as-you-go key character_limit is a very large placeholder number, not a real cap.
voice_settings.speedChanges the length of the audio: the same sentence came out about 3.3 s at 0.7, 2.7 s at 1.0 and 2.1 s at 1.5. Speed 0 is rejected with 422.
Error classesBoth SDKs raised their API error with status 404 (unknown voice), 401 (bad key), 400 (text too long) and 422 (empty text).
Auth headersxi-api-key and Authorization: Bearer both worked on text to speech.

Not tested against the live service: the other output_format names beyond those listed above (an unknown name returns a 422 that lists the names the API accepts, which follow the ElevenLabs list), the ElevenLabs deprecated npm package, voice-agent framework plugins for ElevenLabs (LiveKit, Pipecat), non-English text, and the 402 response when free credits run out. For LiveKit and Pipecat, use the ReadAloud plugins.

What is different or not supported

Different

  • Voices. ElevenLabs voice ids return 404. Use readaloud-default. Your own ElevenLabs voice clones and Voice Library voices cannot be used here.
  • Sound. The voice is not the same as any ElevenLabs voice and we do not claim equal quality. We publish no side-by-side audio on this page.
  • model_id does not choose the voice. Any string is accepted, including ids we do not list such as eleven_v3; the voice you pass decides how the audio is made. The response header X-Model-Mapped reports how the id was treated.
  • Voice settings. Only speed has an effect. stability, similarity_boost, style and use_speaker_boost are accepted and ignored, as are seed, language_code, previous_text, next_text and a few other request fields. The ones you sent are echoed in the x-compat-ignored response header, so nothing is dropped silently.
  • Language. language_code does not select a language. readaloud-default is an American English voice. We did not test other languages through this route.
  • Sample rates above 24 kHz are upsampled. The audio is made at 24 kHz. Formats such as mp3_44100_128 have the sample rate you asked for but no content above 12 kHz. Response headers X-Sample-Rate, X-Audio-Format and X-Source-Sample-Rate say so.
  • wav_* formats work only on the non-streaming endpoint (the streaming endpoint returns 422), as in ElevenLabs.
  • Timestamps. /with-timestamps works but returns null alignment. We do not compute character timings.
  • Markup. No SSML and no v3 audio tags. According to the gateway code, text is passed through as written, so a tag such as [laughs] may be read aloud. We did not test that live.
  • Quota status code. When free credits run out the API returns 402, not the 401 ElevenLabs uses. We did not trigger this in the live run.

Not supported on the ElevenLabs paths

These paths returned 404 in our run (or, for voice creation, a ReadAloud error rather than an ElevenLabs-style result), so do not expect the ElevenLabs SDK methods for them to work against ReadAloud:

  • Speech to text (Scribe), /v1/speech-to-text. ReadAloud has its own transcription API with a different interface; see the Voice API page.
  • Dubbing, /v1/dubbing. ReadAloud’s dubbing is a separate API, not this one.
  • The Agents platform (Conversational AI), /v1/convai/....
  • Voice cloning with ElevenLabs semantics, POST /v1/voices/add (instant and professional voice clones). Sending it with an xi-api-key header returned a ReadAloud authentication error, not a clone. ReadAloud cloning has its own API and needs a payment method.
  • Speech to speech, text to voice (voice design), audio isolation, sound generation, history, pronunciation dictionaries and projects.
  • Not tested: the deprecated ElevenLabs npm package and the ElevenLabs plugins for voice-agent frameworks, see above.

We checked a list of common ElevenLabs paths, not every path in their API. Anything not in the table above is not verified.

Error handling

Errors use the ElevenLabs shape, so the SDK error classes work unchanged. These are real responses from our run.

# 404, unknown voice id (an ElevenLabs voice id)
{"detail":{"status":"voice_not_found","message":"Voice 21m00Tcm4TlvDq8ikWAM not found. ..."}}

# 401, missing or revoked key
{"detail":{"status":"invalid_api_key","message":"Invalid API key"}}

# 400, text over 5,000 characters
{"detail":{"status":"max_character_limit_exceeded","message":"text too long (max 5000 characters per request)"}}

# 422, empty text. Validation errors are a list, as in ElevenLabs' own API
{"detail":[{"loc":["body","text"],"msg":"text must be a non-empty string","type":"value_error"}]}
Statusdetail.statusMeaning
401invalid_api_keyMissing or revoked key. Seen live.
402quota_exceededFree credits used up, or the request is larger than what remains. Add a payment method. Not triggered in our run.
404voice_not_foundUnknown voice id. Seen live.
400max_character_limit_exceededMore than 5,000 characters. Seen live.
422(a list)Validation: empty or missing text, malformed JSON, an unknown output_format, speed of 0 or less. Seen live.
429too_many_concurrent_requestsThe service is at capacity. A Retry-After header is set. We saw this once during the run, on a different voice, and the retry a few seconds later succeeded.
405method_not_allowedA GET on a text-to-speech path. Use POST.

Limits and credits

  • 5,000 characters per request. A 5,000-character request succeeded; 5,001 returned 400. According to the gateway code the same limit applies to each generation on a WebSocket; we did not test that live.
  • Long text takes time to finish. In our run a 5,000-character request returned its first byte in about 0.2 s and the last byte after about 36 s (roughly 5.7 minutes of audio). Split long text into several requests, or use the streaming endpoint and play as you receive.
  • Concurrency. Six simultaneous requests all succeeded. We do not publish a concurrency number on this page; the service returns 429 with Retry-After when it is full.
  • Free credits. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech), shared across all of your keys. When they run out, requests return 402 until you add a payment method; after that you pay as you go. See free credits and the pricing on the Voice API page.

FAQ

Do I only need to change the base URL?

No, three values: the base URL, the key and the voice id. Without a ReadAloud voice id the call returns 404.

Can I reuse my ElevenLabs voices?

No. ElevenLabs voice ids, including your own clones, do not exist here. We do not map them to ReadAloud voices.

Will it sound the same?

No. It is a different voice. We do not claim parity, so judge it on your own text.

Do I need to remove stability and similarity_boost?

No. They are accepted and ignored, and the x-compat-ignored response header tells you which ones were ignored. Only speed has an effect.

Does streaming work?

Yes: the /stream endpoint and the stream-input WebSocket both worked in our run.

Can I use ElevenLabs speech to text, dubbing or agents through this API?

No. Only text to speech is covered. See the list above.

What does it cost?

New accounts get a one-time grant of free credits; after that you pay as you go. Current prices are on the Voice API page.

Something here does not match what I see.

Tell us at support@readaloudai.org. This page records one dated test run, not a guarantee, and we will correct it.

ReadAloud AI is not affiliated with ElevenLabs. ElevenLabs and its product names belong to their owner and are used here only to describe compatibility.