Install
pip install pipecat-readaloud
export READALOUD_API_KEY=rtts_... # placeholder: use your own key
Requires Python 3.10+ and pipecat-ai 1.12 or newer (below 2). Create a ReadAloud API key at https://readaloudai.org/developers and export it as READALOUD_API_KEY; new accounts start with free credits.Quickstart
import asyncio
import os
import wave
import aiohttp
from pipecat.frames.frames import EndFrame, Frame, TTSAudioRawFrame, TTSSpeakFrame
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.worker import PipelineParams, PipelineWorker
from pipecat.processors.frame_processor import FrameDirection, FrameProcessor
from pipecat.workers.runner import WorkerRunner
from pipecat_readaloud import ReadAloudHttpTTSService
RATE = 24000
class WavWriter(FrameProcessor):
def __init__(self, path):
super().__init__()
self._path, self._chunks = path, []
async def process_frame(self, frame: Frame, direction: FrameDirection):
await super().process_frame(frame, direction)
if isinstance(frame, TTSAudioRawFrame):
self._chunks.append(frame.audio)
elif isinstance(frame, EndFrame):
with wave.open(self._path, "wb") as f:
f.setnchannels(1)
f.setsampwidth(2)
f.setframerate(RATE)
f.writeframes(b"".join(self._chunks))
print("frames:", len(self._chunks), "bytes:", sum(map(len, self._chunks)))
await self.push_frame(frame, direction)
async def main():
async with aiohttp.ClientSession() as session:
tts = ReadAloudHttpTTSService(
api_key=os.environ["READALOUD_API_KEY"],
aiohttp_session=session,
settings=ReadAloudHttpTTSService.Settings(voice="readaloud-default", speed=1.0),
)
worker = PipelineWorker(
Pipeline([tts, WavWriter("pipecat_out.wav")]),
params=PipelineParams(audio_out_sample_rate=RATE),
)
await worker.queue_frames([TTSSpeakFrame("Thanks for calling. How can I help you today?"), EndFrame()])
runner = WorkerRunner(handle_sigint=False)
await runner.add_workers(worker)
await runner.run()
asyncio.run(main())What was tested
Checked with pipecat-readaloud (PyPI), with pipecat-ai 1.12.0, aiohttp 3.14.5, Python 3.14.6. Version tested: 0.1.1. Date: October 10, 2026. This is one dated check against the live API. Other versions, and later releases, may behave differently.
What worked
- Installed pipecat-readaloud 0.1.1 from PyPI into a clean virtual environment; it pulled in pipecat-ai 1.12.0 and the quickstart on this page ran unchanged.
- A TTSSpeakFrame through ReadAloudHttpTTSService with voice readaloud-default produced 28 audio frames, 113,128 bytes, which ffprobe read as a 24 kHz mono 16-bit WAV of 2.36 seconds.
- With sample_rate=8000 the same pipeline produced 7 frames, 37,526 bytes, read by ffprobe as 8 kHz mono audio of 2.35 seconds.
- The package's own default voice value, "default", also synthesized audio (2.29 seconds in our run), so existing code that uses it keeps working.
- A wrong API key surfaced as a Pipecat ErrorFrame reading '401 invalid_api_key', not as a hang or a crash.
Limits and what is not supported
- We ran a TTS-only pipeline that writes a WAV file. We did not run a full agent with speech-to-text, an LLM and a telephony or WebRTC transport.
- Interruption handling, runtime voice and speed updates with TTSUpdateSettingsFrame, and the retry path for 429 and 503 responses are described in the package README; we did not exercise them.
- ReadAloud Live has one English voice, readaloud-default, and we used only that voice. Each request is limited to 5,000 characters.
- The API returns no word timestamps, so the text frames Pipecat emits are not time-aligned with the audio.
Limits that apply to every integration
- 5,000 characters per request. Split longer text and send it in order.
- Output formats on the OpenAI-compatible route: mp3 (default, 24 kHz), opus (48 kHz, Ogg), wav (24 kHz) and pcm (24 kHz, 16-bit, mono); aac and flac return 400.
- No SSML and no audio tags.
- Beyond capacity the API returns 429 (HTTP) or an "at capacity" error with close code 1013 (WebSocket); retry with a short backoff.
What ReadAloud costs
- ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
- ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.
Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).
Common questions
- Which voice should I pass?
- Use readaloud-default, the single English voice on ReadAloud Live. The package's built-in value "default" also worked in our run. We only tested this one voice.
- Can I use it on a phone call?
- The service can request 8 kHz audio, which is what Twilio and Telnyx transports use, and our 8 kHz run produced a valid 8 kHz mono file. We did not place a phone call through Pipecat.
- What happens if the user interrupts the bot?
- The README says Pipecat cancels the in-flight synthesis and the connection is closed, which stops generation on the server. We did not test an interruption, so confirm it in your own pipeline.
- Does it need an LLM or a transport to try it?
- No. The quickstart here is a pipeline of the TTS service and a small processor that writes a WAV file, so a ReadAloud key is the only credential you need.
Something here does not match what you see? Tell us at support@readaloudai.org.