Voice for your AI tools

Add one URL to Claude, Cursor, VS Code or any other MCP client, and your assistant can turn text into speech with ReadAloud AI voices. Nothing to install or host.

https://readaloudai.org/mcp

1. Connect with a login (OAuth)

The easiest way. Add the URL, sign in to your ReadAloud AI account in the browser, and approve. No key to copy or paste.

  • claude.ai: Settings, then Connectors, then Add custom connector. Enter https://readaloudai.org/mcp and follow the sign-in window.
  • Claude Code: run claude mcp add --transport http readaloud https://readaloudai.org/mcp, then type /mcp and choose readaloud to sign in.
  • Cursor and VS Code: add the URL as a remote MCP server. The app opens a sign-in page when it first connects.

When you approve, we create an API key named “MCP connector (app name)” for that app and it uses your free or paid characters. To disconnect, revoke that key in the developer console. The app’s next request fails and asks you to reconnect. Access tokens last one hour and are renewed automatically for up to 30 days; reconnecting replaces the previous key for the same app.

2. Or use an API key header

For scripts and clients without login support. Sign in on the Voice API page and create a key (new keys include 10,000 free characters), then pick your client and replace YOUR_API_KEY. Your client sends the key with every request, so keep it out of chat messages and public repos.

claude mcp add --transport http readaloud https://readaloudai.org/mcp \
  --header "Authorization: Bearer YOUR_API_KEY"

Run this once in your terminal. Add --scope user to make it available in every project.

Try it

Ask your assistant something like:

  • “Say ‘Your order has shipped’ using the Kokoro voice af_heart.”
  • “What voices are available? Read each one a short greeting so I can compare.”
  • “Turn this paragraph into audio at 1.2x speed.”

The audio comes back inside the tool result as a WAV clip. Whether you can hear it inline depends on the client; some only show a download or attachment.

Tools

ToolWhat it does
text_to_speechSpeaks up to 1,000 characters and returns a WAV clip (24 kHz, mono) plus the duration and time to first audio. Inputs: text, engine (piper or kokoro), voice, speed (0.5 to 2).
list_voicesLists the voices you can use for each engine. Input: engine (optional).
get_api_statusShows whether the Piper engine is up and how many of its simultaneous streams are in use. Useful after a capacity error.
speech_to_textTranscribes spoken audio with Whisper large-v3-turbo. Inputs: audio_base64 (up to 25 MB decoded; WAV, FLAC, OGG, MP3, M4A or WEBM), language (optional), word_timestamps (optional). Returns the transcript, detected language and duration. Batch only.
dub_audioRe-voices a spoken audio clip into another language: transcribe, translate, then resynthesize each segment with approximate timing. Audio in, audio out only (no video, no subtitles), and the timing is an approximation, not true alignment. Returns a job_id. Inputs: audio_base64 (up to 25 MB), target_language, source_language (optional). Uses your free credits, then needs a payment method.
get_dub_statusPolls a dub_audio job. Returns processing, ready (with an audio_url) or failed. Input: job_id.
isolate_voiceSeparates vocals from the rest of a clip (Demucs) and can also return the instrumental. Returns a job_id. Inputs: audio_base64 (up to 8 MB decoded), confirms_rights (must be true), want_instrumental (optional). Uses your free credits, then needs a payment method, and needs an isolation container deployed for your account first (deploy it at /isolate-voice while signed in). Quality on noisy real-world speech is limited: it is music-oriented separation, not a speech denoiser.
get_voice_isolationPolls an isolate_voice job. Returns the status, and the separated WAV once it is done. Inputs: job_id, stem (vocals or instrumental).
create_audiobookTurns long text or an ePub into a chaptered audiobook. Returns an audiobook_id right away; chapters are synthesized in the background.
get_audiobook_statusShows the status and duration of each chapter of an audiobook. Input: audiobook_id.
export_audiobookJoins the finished chapters into one MP3 with chapter markers and returns a download URL. Every chapter must be ready. Input: audiobook_id.
design_voiceGenerates a new synthetic voice from a text description and a sample sentence. Returns a job_id. Inputs: description (10 to 800 characters), text (up to 500 characters). Uses your free credits, then needs a payment method.
get_voice_designPolls a design_voice job. Returns the status, and the WAV sample once it is ready. Input: job_id.
convert_voiceSpeech-to-speech conversion: re-speaks a source clip in the voice of a target reference clip. Returns a job_id. Inputs: source_audio_base64, target_audio_base64 (up to 4 MB each), MIME types, confirms_rights (must be true). The first call on an account sets up a private converter (about 3 minutes) and returns a temporary capacity error; call again afterwards.
get_voice_conversionPolls a convert_voice job. Returns the status, and the converted WAV once it is done. Input: job_id.
create_voice_cloneStarts a custom cloned voice (Piper fine-tune) and records the speaker’s consent. Returns a voice_id. Inputs: speaker_name, attested_by, consent (true), consent_statement (the exact wording). Needs a billing-enabled key.
upload_voice_clone_datasetUploads the recordings for a voice as a ZIP. Inputs: voice_id, and either zip_base64 (up to 8 MB) or zip_url (public https, up to 48 MB). Does not start training or bill.
commit_voice_clone_datasetStarts training (30 to 60 minutes) and bills $2.50 per voice. Inputs: voice_id, confirms_charge (must be true).
get_voice_clone_statusShows a cloned voice’s status: created, training, ready or rejected. Input: voice_id.
deploy_voice_cloneMakes a ready cloned voice usable. Returns the voice name to pass to text_to_speech, as custom:<voice_id>. Input: voice_id.
delete_voice_clonePermanently deletes a cloned voice and its recordings. Input: voice_id.

Voices: Piper has one voice, default. Kokoro has 28 English voices such as af_heart, am_adam and bf_emma; ask for list_voices for all of them. Custom trained voices use custom:<id>.

Most tools return a job id or a URL rather than inline audio, and the longer ones (dubbing, isolation, conversion, audiobooks, cloning) are polled with a matching get_… tool. The tools that need an account (dubbing, sound effects, isolation, conversion, voice design, audiobooks) bill the account your key belongs to; a key that is not tied to an account gets a payment-required error. The REST endpoints behind them do not take your key directly. This server is the only place an API key works for them. Web versions: transcribe, dub, sound effects, isolate voice, audiobooks, design voice and convert voice. Prices are on the Voice API page; audio-based usage shows on your invoice as character equivalents.

Limits

  • text_to_speech: up to 1,000 characters and about 20 seconds of audio per call. Longer audio is cut off, so split long text into several calls. Other tools have their own size limits, listed above.
  • Each call times out after 30 seconds.
  • 15 speech calls per minute per key, and 120 requests per minute per IP address. Over the limit you get a retry message.
  • Speech uses your key’s characters and pricing, the same as the WebSocket API. When the free characters run out, tools return a “payment required” error.
  • If the voice server is busy, the tool says so and tells the assistant to retry in a few seconds.

Security

  • Keys belong to one person. Do not share yours or commit it. If it leaks, or you want to disconnect an app you connected with a login, revoke its key in the developer console.
  • Our server uses your key only to authorize each request and to work out which account it belongs to. We do not log or store it.
  • text_to_speech returns audio inline and we do not save it or keep the text you send. Tools that run jobs (dubbing, sound effects, isolation, conversion, audiobooks, voice design, cloning) necessarily store your input and results on our servers while the job runs and so you can fetch it.
  • The text you ask to speak is sent to the voice engine that generates it.

Prefer the raw API? See the WebSocket reference for streaming and lower latency.