Install
pip install livekit-plugins-readaloud
export READALOUD_API_KEY=rtts_... # placeholder: use your own key
Requires Python 3.10+ and livekit-agents 1.8 or newer (below 2). Create a ReadAloud API key at https://readaloudai.org/developers and export it as READALOUD_API_KEY; new accounts start with free credits.Quickstart
import asyncio
import os
import wave
import aiohttp
from livekit.plugins import readaloud
async def main():
async with aiohttp.ClientSession() as session:
tts = readaloud.TTS(
api_key=os.environ["READALOUD_API_KEY"],
http_session=session,
voice="readaloud-default",
sample_rate=24000, # 24000 | 16000 | 8000
)
async with tts.synthesize("Thanks for calling. How can I help you today?") as stream:
frame = await stream.collect()
with wave.open("livekit_out.wav", "wb") as f:
f.setnchannels(frame.num_channels)
f.setsampwidth(2)
f.setframerate(frame.sample_rate)
f.writeframes(bytes(frame.data))
print(f"wrote livekit_out.wav ({frame.duration:.2f} s at {frame.sample_rate} Hz)")
asyncio.run(main())What was tested
Checked with livekit-plugins-readaloud (PyPI), with livekit-agents 1.8.6, aiohttp 3.14.5, Python 3.14.6. Version tested: 0.1.1. Date: October 10, 2026. This is one dated check against the live API. Other versions, and later releases, may behave differently.
What worked
- Installed livekit-plugins-readaloud 0.1.1 from PyPI into a clean virtual environment; it resolved livekit-agents 1.8.6, one patch above the 1.8.5 its README names as tested.
- readaloud.TTS(...).synthesize(...) followed by collect() returned one audio frame of 2.38 seconds at 24,000 Hz; the WAV we wrote from it was 114,212 bytes and ffprobe read it as 24 kHz mono 16-bit.
- tts.stream() with text pushed in three fragments and end_input() yielded 17 audio events, 37,128 bytes in total, at 8,000 Hz when the plugin was created with sample_rate=8000.
- The voice value readaloud-default was accepted by the plugin; no code change was needed beyond the key and the voice name.
Limits and what is not supported
- We did not run an AgentSession, a LiveKit server, speech-to-text or an LLM. The test called the TTS object directly.
- Barge-in, LiveKit's APIConnectOptions retries and the plugin's own retry on 429 and 503 are described in the README but were not exercised.
- ReadAloud Live offers one English voice, readaloud-default, and only that voice was tested. Each request is limited to 5,000 characters.
- The API returns no word timestamps, so LiveKit's aligned transcripts are not available.
Limits that apply to every integration
- 5,000 characters per request. Split longer text and send it in order.
- Output formats on the OpenAI-compatible route: mp3 (default, 24 kHz), opus (48 kHz, Ogg), wav (24 kHz) and pcm (24 kHz, 16-bit, mono); aac and flac return 400.
- No SSML and no audio tags.
- Beyond capacity the API returns 429 (HTTP) or an "at capacity" error with close code 1013 (WebSocket); retry with a short backoff.
What ReadAloud costs
- ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
- ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.
Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).
Common questions
- Where do I pass the plugin?
- As the tts argument of an AgentSession, for example AgentSession(tts=readaloud.TTS(voice="readaloud-default", sample_rate=24000)). We tested the TTS object on its own, not inside a session.
- What is the difference between synthesize() and stream()?
- synthesize() takes one piece of text and returns the audio for it. stream() accepts text in fragments, splits it into sentences and synthesizes each one, which is the path a streaming LLM reply uses. We ran both.
- Which sample rates are supported?
- The plugin documents 24000, 16000 and 8000. We ran 24000 and 8000 and did not run 16000.
- Do I need a LiveKit server to try it?
- Not to hear the voice. The quickstart creates the TTS object, synthesizes a sentence and writes a WAV file with only a ReadAloud key.
Something here does not match what you see? Tell us at support@readaloudai.org.