ReadAloud speech provider for the Vercel AI SDK

The Vercel AI SDK gives TypeScript applications one function, generateSpeech, for text to speech, and provider packages decide which service answers. ai-sdk-provider-readaloud is a provider for ReadAloud. It implements the SpeechModelV4 interface that AI SDK 7 uses, so you call generateSpeech with model: readaloud.speech() and get back the audio bytes, the media type, warnings and response metadata in the SDK's usual shape. It has a single runtime dependency, @ai-sdk/provider, ships TypeScript types and is ESM only. Output can be MP3, Opus in an Ogg container, WAV or raw 24 kHz PCM. Options the service cannot honor, such as instructions, do not fail silently: the provider returns an AI SDK warning that says the parameter was ignored. Errors from the API arrive as the SDK's APICallError with the status code, so the SDK's retry logic and your own error handling work as they do with other providers. The provider reads READALOUD_API_KEY when a request is made, and createReadAloud lets you set the base URL, headers or a custom fetch. We installed the published package with ai 7.0.137 from npm on 10 October 2026 and ran generateSpeech against the live API; the results and the untested parts are below.

Install

npm install ai ai-sdk-provider-readaloud
export READALOUD_API_KEY=rtts_...   # placeholder: use your own key

Targets AI SDK 7 (specificationVersion v4); ESM only. Create a ReadAloud API key at https://readaloudai.org/developers and export it as READALOUD_API_KEY; new accounts start with free credits.

Quickstart

Quickstart (javascript)
import { writeFile } from "node:fs/promises";
import { generateSpeech } from "ai";
import { readaloud } from "ai-sdk-provider-readaloud";

const { audio, warnings, providerMetadata } = await generateSpeech({
  model: readaloud.speech(),
  text: "Thanks for calling. How can I help you today?",
  voice: "readaloud-default",
  outputFormat: "mp3",
});
await writeFile("vercel_out.mp3", audio.uint8Array);
console.log(audio.mediaType, audio.uint8Array.length, "bytes", JSON.stringify(warnings), JSON.stringify(providerMetadata));

What was tested

Checked with ai-sdk-provider-readaloud (npm), with ai 7.0.137, on Node 26.8.2. Version tested: 0.1.1. Date: October 10, 2026. This is one dated check against the live API. Other versions, and later releases, may behave differently.

What worked

  • npm installed ai-sdk-provider-readaloud 0.1.1 next to ai 7.0.137 without peer-dependency errors, and the quickstart ran as written.
  • generateSpeech with outputFormat mp3 returned mediaType audio/mpeg and 39,213 bytes; ffprobe read the saved file as 24 kHz mono MP3 of 2.45 seconds. providerMetadata.readaloud reported characterCount 45, sampleRate 24000 and audioFormat mp3_24000_128.
  • The other formats returned the right container: opus gave audio/ogg starting with OggS (6,374 bytes), wav gave audio/wav starting with RIFF (37,382 bytes), and pcm gave audio/pcm (34,552 bytes).
  • Passing instructions returned a warning, 'ReadAloud speech models do not support instructions. The parameter was ignored.', and the call still succeeded.
  • A wrong API key threw the SDK's AI_APICallError with statusCode 401 and the message 'Invalid API key'.

Limits and what is not supported

  • Tested on AI SDK 7 only, on Node 26.8.2; the package README lists Node 22 and 24 as its tested versions. The README's AI SDK 6 route through @ai-sdk/openai was not run by us.
  • ReadAloud Live has one English voice, readaloud-default, and only that voice was used. language and instructions are not supported. Text is limited to 5,000 characters per call.
  • We did not test speed clamping, abort signals, custom fetch or headers, or use in an edge runtime.
  • The package README and its default voice still use a legacy voice id that the API accepts as an alias; pass voice: 'readaloud-default' explicitly as shown.

Limits that apply to every integration

  • 5,000 characters per request. Split longer text and send it in order.
  • Output formats on the OpenAI-compatible route: mp3 (default, 24 kHz), opus (48 kHz, Ogg), wav (24 kHz) and pcm (24 kHz, 16-bit, mono); aac and flac return 400.
  • No SSML and no audio tags.
  • Beyond capacity the API returns 429 (HTTP) or an "at capacity" error with close code 1013 (WebSocket); retry with a short backoff.

What ReadAloud costs

  • ReadAloud Live: $4 per 1M characters ($0.004 per 1,000), the low-latency tier, built for live calls and voice agents.
  • ReadAloud Studio: $10 per 1M characters ($0.01 per 1,000), the higher-priced tier for read-aloud and narration, with several English voices.

Billing is per character of speech that finishes; cancelled requests are not billed, and there are no minimums. Every account gets a one-time grant of free credits (worth $0.10, about 10,000 characters of speech). It is one capped pool per account, shared across all of the account's keys and across speech, transcription, dubbing, voice conversion and voice design, and it is used first. When it runs out, requests return 402 until a payment method is added, then billing is pay as you go. Current prices are on the Voice API page (/developers).

Common questions

Does it work with AI SDK 6?
This package targets AI SDK 7. Its README describes an AI SDK 6 route using @ai-sdk/openai with ReadAloud's OpenAI-compatible endpoint; we did not run that route.
Does the model id choose the voice?
No. The voice option does. The model id is forwarded but does not select a voice, so readaloud.speech() with no argument is fine.
How are unsupported options handled?
They appear in the warnings array of the result. In our run, instructions produced a warning and the audio was still generated.
What are the output formats?
mp3, opus (Ogg), wav and pcm (24 kHz, 16-bit mono). We ran all four and checked the file signature of each.

Something here does not match what you see? Tell us at support@readaloudai.org.