Moving from OpenAI Text to Speech to ReadAloud

A step-by-step checklist for moving text-to-speech calls from OpenAI Text to Speech to ReadAloud, based on OpenAI Text to Speech's public pages as of October 10, 2026.

Last verified October 10, 2026.

Details about OpenAI Text to Speech come from its public pages as they were on October 10, 2026. OpenAI Text to Speech may have changed them since, and a page that was unclear to us may be clear to you. Corrections: support@readaloudai.org. We make ReadAloud, so read this page with that in mind and check anything important with OpenAI Text to Speech directly.

Moving from OpenAI Text to Speech: what kind of move this is

How OpenAI Text to Speech positions itself: OpenAI's Audio API speech endpoint, offering steerable gpt-4o-mini-tts and the earlier tts-1 and tts-1-hd models behind a single request shape that many tools and SDKs already support.

OpenAI Text to Speech is a model vendor, and speech is one endpoint among its other services. Moving speech alone leaves the rest of your integration in place.

OpenAI Text to Speech's pages suggest it fits: developers already on the OpenAI SDK who want steerable speech or a per-character model without adding another vendor, and who mainly synthesize English.

Before you leave OpenAI Text to Speech

  • Collect the exact text you send to OpenAI Text to Speech today: the longest, the shortest, and the ones with names, numbers and abbreviations.
  • Conditions on OpenAI Text to Speech's side: OpenAI usage policies require a clear disclosure to end users that the TTS voice is AI-generated and not a human voice. Check what they mean for audio you already generated. (source: developers.openai.com)
  • List where your code uses Node.js / TypeScript, Python, Go or Java so you can replace each call site.
  • Condition listed by OpenAI Text to Speech: Voices are stated to be optimized for English. (source: developers.openai.com)
  • Condition listed by OpenAI Text to Speech: Custom voices are limited to eligible customers and capped at 20 voices per organization. (source: developers.openai.com)
  • Condition listed by OpenAI Text to Speech: The instructions parameter and sse streaming are not supported on tts-1 and tts-1-hd. (source: developers.openai.com)
  • Save any audio you need to keep from OpenAI Text to Speech before you cancel, and note which plan covers the right to keep using it.

Steps to move off OpenAI Text to Speech

These steps come from OpenAI Text to Speech's public pages. Sources are listed at the end.

  1. Find the two settings that point an OpenAI client at a service: the base URL (for example https://api.openai.com/v1) and the Bearer key. Both are normally set once, in the client. (source: developers.openai.com)
  2. List the model values (tts-1, tts-1-hd, gpt-4o-mini-tts) and voice names (alloy, coral, marin and the others) that your call sites pass. (source: developers.openai.com)
  3. List the calls that pass instructions: this steering text is a gpt-4o-mini-tts feature. (source: developers.openai.com)
  4. List the response_format values you request (mp3, opus, aac, flac, wav, pcm) and the sample rate your playback path expects; OpenAI pcm is 24 kHz 16-bit mono. (source: developers.openai.com)
  5. Check whether you use stream_format=sse or chunked audio streaming. (source: developers.openai.com)
  6. OpenAI custom voice ids do not transfer to another service. (source: developers.openai.com)

Map each OpenAI Text to Speech request field and endpoint

OpenAI Text to Speech fields and endpoints and what they become at ReadAloud
In OpenAI Text to SpeechAt ReadAloud
POST /v1/audio/speechPOST /v1/audio/speech (the OpenAI-compatible route), which returns the audio in the response. The request shape is the same, so the OpenAI SDK works against the ReadAloud base URL.
inputinput on the OpenAI-compatible route (text on the WebSocket), up to 5,000 characters per request; split longer text at sentence boundaries
model (tts-1, tts-1-hd, gpt-4o-mini-tts)model is accepted and ignored: the voice decides how the audio is made
voice (alloy, coral, marin and the other stock names)voice: readaloud-default on ReadAloud Live, or a ReadAloud Studio voice listed by GET /v1/voices. Vendor voice names and ids do not exist at ReadAloud. The OpenAI stock voice names are accepted and all map to readaloud-default.
instructions (a gpt-4o-mini-tts parameter)No equivalent. ReadAloud has no acting instructions, style prompts, SSML or audio tags
response_format (mp3, opus, aac, flac, wav, pcm)response_format on the OpenAI-compatible route: mp3, opus, wav or pcm (aac and flac return 400); the WebSocket also offers 8 kHz mu-law and A-law
speed (0.25 to 4.0)speed, 0.25 to 4.0 on the OpenAI-compatible route
stream_format (sse or audio)Accepted and ignored by the OpenAI-compatible route

What changes in your code and what does not carry over

  • Endpoint: OpenAI Text to Speech lists an OpenAI-compatible speech endpoint, so changing the base URL to https://api.readaloudai.org/v1 and the key may be enough.
  • Voice: OpenAI Text to Speech voices do not exist at ReadAloud. Live: one American English voice. Studio: several English voices.
  • Languages, as OpenAI Text to Speech's pages put it: Follows the language support of OpenAI's speech-to-text model (about 57 languages listed); voices are stated to be optimized for English. ReadAloud: English today. (source: developers.openai.com)
  • Formats: OpenAI Text to Speech lists mp3 (default), opus, aac, flac, wav and pcm (24 kHz, 16-bit signed little-endian). ReadAloud returns mp3, opus, wav and pcm on at least one route. It does not return aac and flac, so convert on your side if you need them.
  • Request size: OpenAI Text to Speech states 4096 characters per request (API reference). ReadAloud accepts 5,000 characters per request, so split longer text at sentence boundaries. (source: developers.openai.com)
  • Custom voices, as OpenAI Text to Speech's pages describe them: Custom voices are limited to eligible customers (contact sales). Creating one needs a consent recording reading a prescribed phrase plus a sample of 30 seconds or less; maximum 20 voices per organization; the Text-to-Speech Supplemental Agreement applies. Those voices cannot be exported to ReadAloud. (source: developers.openai.com)

Where ReadAloud does not match OpenAI Text to Speech

  • Expressive control: OpenAI Text to Speech's pages list controls over emotion, acting or style. ReadAloud does not offer emotion or expressive control: it has no acting instructions, no style prompts, no SSML and no audio tags, and the instructions field of the OpenAI route is ignored. If you need that, ReadAloud is not the right fit for that part of your product; evaluate other vendors that list it. (source: developers.openai.com)

Steps on the ReadAloud side

  1. Create a key (/developers#get-started) and run the test call below.
  2. Pick ReadAloud Live (the low-latency tier, built for live calls and voice agents) or ReadAloud Studio (the higher-priced tier for read-aloud and narration, with several English voices).
  3. Keep the base URL, key and voice in a setting, and move a small share of traffic first.
A first test call to ReadAloud (works from any language that can send an HTTP request)
# ReadAloud through its OpenAI-compatible speech route (/v1/audio/speech).
# Replace YOUR_KEY with a ReadAloud API key.
curl -X POST "https://api.readaloudai.org/v1/audio/speech" \
  -H "Authorization: Bearer YOUR_KEY" -H "Content-Type: application/json" \
  -d '{"model":"tts-1","voice":"readaloud-default","input":"Paste a sentence you send to OpenAI Text to Speech today.","response_format":"mp3"}' \
  -o test.mp3

What it costs to test ReadAloud

A one-time grant of free credits worth $0.10 (about 10,000 characters of speech), shared by all keys on the account.

Every per-character list price we found for OpenAI Text to Speech, tts-1 at $15 and tts-1-hd at $30, is higher than both ReadAloud Live ($4) and ReadAloud Studio ($10) per 1M characters.

These are list prices as of October 10, 2026. The tiers cover different things, so price your own volume rather than reading across.

What we could not confirm about OpenAI Text to Speech

  • Per-character equivalent cost for gpt-4o-mini-tts (token-billed)
  • Whether the speech endpoint offers WebSocket streaming (the separate Realtime API does)
  • SSML support (no SSML parameter in the API reference read)
  • Commercial-use statement for generated audio (not found on pages read)
  • SOC 2, HIPAA and GDPR statements (not on pages read)
  • Vendor-stated latency figures (none quoted)
  • tts-1-hd price confirmed on pricing page and model page quick comparison as a table row; rate limit rows for tts-1-hd not read

Sources