Last verified October 10, 2026.
Details about OpenAI Text to Speech come from its public pages as they were on October 10, 2026. OpenAI Text to Speech may have changed them since, and a page that was unclear to us may be clear to you. Corrections: support@readaloudai.org. We make ReadAloud, so read this page with that in mind and check anything important with OpenAI Text to Speech directly.
About OpenAI Text to Speech
How OpenAI Text to Speech positions itself: OpenAI's Audio API speech endpoint, offering steerable gpt-4o-mini-tts and the earlier tts-1 and tts-1-hd models behind a single request shape that many tools and SDKs already support.
Price per 1 million characters
USD per one million characters, by tier. Other units are shown as the vendor writes them, not converted.
| Option | Listed price | What it covers | Against ReadAloud Live |
|---|---|---|---|
| ReadAloud Live | $4 per 1M characters | the low-latency tier, built for live calls and voice agents | Reference point |
| ReadAloud Studio | $10 per 1M characters | the higher-priced tier for read-aloud and narration, with several English voices | $6 above ReadAloud Live |
| OpenAI Text to Speech: tts-1 | $15 per 1M characters (source: developers.openai.com) | Described in the OpenAI guide as the lower-latency option next to tts-1-hd. Supports 9 voices. Does not accept the instructions parameter. | $11 higher than ReadAloud Live |
| OpenAI Text to Speech: tts-1-hd | $30 per 1M characters (source: developers.openai.com) | Listed above tts-1 on the OpenAI pricing page. | $26 higher than ReadAloud Live |
| OpenAI Text to Speech: gpt-4o-mini-tts | Vendor's own unit: USD per 1M tokens: $0.60 text input, $12.00 audio output; not converted (source: developers.openai.com) | Token-billed, so cost per character depends on text and generated audio length; OpenAI states audio output is metered as tokens. Supports the instructions parameter and 13 voices; model page states max input is 2000 tokens. | Not comparable: OpenAI Text to Speech lists a different unit |
Every per-character list price we found for OpenAI Text to Speech, tts-1 at $15 and tts-1-hd at $30, is higher than both ReadAloud Live ($4) and ReadAloud Studio ($10) per 1M characters.
OpenAI Text to Speech's own headline, as of October 10, 2026: tts-1 is $15 per 1M characters; tts-1-hd is $30 per 1M characters; gpt-4o-mini-tts is $0.60 per 1M text input tokens and $12.00 per 1M audio output tokens.
Free tier at OpenAI Text to Speech: No free character or credit allowance stated on pages read. The tts-1 model page lists a Free rate-limit tier of 3 requests per minute and 200 per day. ReadAloud: A one-time grant of free credits worth $0.10 (about 10,000 characters of speech), shared by all keys on the account.
Tiers cover different things, so compare on your own monthly volume, and check OpenAI Text to Speech's pricing page before you decide. Current ReadAloud prices: /developers.
Side by side
"Not stated on the pages we reviewed" marks what OpenAI Text to Speech's pages did not say. We print no head-to-head latency number.
| Topic | ReadAloud | OpenAI Text to Speech |
|---|---|---|
| Voices | Live: one American English voice. Studio: several English voices. | 13 built-in voices for gpt-4o-mini-tts (alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin, cedar); 9 for tts-1 and tts-1-hd (source: developers.openai.com) |
| Languages | English today | Follows the language support of OpenAI's speech-to-text model (about 57 languages listed); voices are stated to be optimized for English (source: developers.openai.com) |
| Custom voices or cloning | API, with a consent step and a payment method | Custom voices are limited to eligible customers (contact sales). Creating one needs a consent recording reading a prescribed phrase plus a sample of 30 seconds or less; maximum 20 voices per organization; the Text-to-Speech Supplemental Agreement applies. (source: developers.openai.com) |
| Streaming | WebSocket and HTTP streaming | WebSocket: not stated; HTTP streaming: yes (source: developers.openai.com) |
| Time to first audio | ReadAloud Live: about 250 ms (our measurement, median to first audio byte, warm connection, San Jose; October 10, 2026; /developers#vs-elevenlabs) | Not stated on the pages we reviewed |
| Output formats | mp3, opus, wav, pcm; 8 kHz mu-law and A-law over WebSocket | mp3 (default), opus, aac, flac, wav, pcm (24 kHz, 16-bit signed little-endian) |
| SSML | No SSML and no audio tags. | Not stated on the pages we reviewed |
| OpenAI-compatible speech endpoint | Yes, /v1/audio/speech | Yes. This is the OpenAI speech request itself: POST /v1/audio/speech with model, input, voice, optional instructions, response_format, speed (0.25 to 4.0) and stream_format (sse or audio; sse not supported on tts-1 and tts-1-hd). |
| Limit per request | 5,000 characters | 4096 characters per request (API reference) (source: developers.openai.com) |
| Concurrency | Up to 12 Live streams per server; more start under load | Rate limits by usage tier. tts-1: Free 3 RPM / 200 RPD, Build 2,500 RPM, Launch 7,500 RPM, Grow 10,000 RPM. gpt-4o-mini-tts: Build 2,000 RPM / 150,000 TPM, Launch 10,000 RPM, Grow 10,000 RPM. (source: developers.openai.com) |
| SDKs and plugins | Python and JavaScript libraries (readaloud), Pipecat and LiveKit plugins, any OpenAI SDK | Node.js / TypeScript, Python, Go, Java, C# / .NET, Ruby and OpenAI CLI |
| HIPAA | Not claimed | Not stated on the pages we reviewed |
| SOC 2 | Not claimed | Not stated on the pages we reviewed |
| GDPR | Not claimed | Not stated on the pages we reviewed |
About OpenAI Text to Speech's compliance wording: No compliance statements were found on the speech documentation or pricing pages read; OpenAI publishes security and privacy material elsewhere that was not reviewed.
What OpenAI Text to Speech lists as strengths
- gpt-4o-mini-tts accepts natural-language instructions to steer accent, emotional range, intonation, speed, tone and whispering. (source: developers.openai.com)
- A single endpoint with first-party SDKs in many languages and support in many third-party tools. (source: developers.openai.com)
- Per-character pricing for tts-1 ($15) and tts-1-hd ($30) alongside a token-priced steerable model. (source: developers.openai.com)
- Streaming through chunked transfer and, on gpt-4o-mini-tts, SSE; wav or pcm recommended by OpenAI for low-latency responses. (source: developers.openai.com)
- Custom voices with a consent-recording requirement for eligible customers. (source: developers.openai.com)
Conditions to check with OpenAI Text to Speech
Limits or conditions from the pages we reviewed.
- Voices are stated to be optimized for English. (source: developers.openai.com)
- Custom voices are limited to eligible customers and capped at 20 voices per organization. (source: developers.openai.com)
- The instructions parameter and sse streaming are not supported on tts-1 and tts-1-hd. (source: developers.openai.com)
OpenAI Text to Speech usage terms
From the vendor pages we reviewed; not legal advice.
- Attribution or conditions: OpenAI usage policies require a clear disclosure to end users that the TTS voice is AI-generated and not a human voice. (source: developers.openai.com)
Where ReadAloud is different
- Voices: OpenAI Text to Speech lists 13 built-in voices for gpt-4o-mini-tts (alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin, cedar); 9 for tts-1 and tts-1-hd. ReadAloud: Live: one American English voice. Studio: several English voices. (source: developers.openai.com)
- Custom voices, as OpenAI Text to Speech's pages describe them: Custom voices are limited to eligible customers (contact sales). Creating one needs a consent recording reading a prescribed phrase plus a sample of 30 seconds or less; maximum 20 voices per organization; the Text-to-Speech Supplemental Agreement applies. (source: developers.openai.com)
- Switching: both list an OpenAI-compatible speech endpoint, so a base URL, a key and a voice name may be all that changes. The OpenAI stock voice names all map to readaloud-default.
Which one fits
OpenAI Text to Speech may suit you if this describes you: developers already on the OpenAI SDK who want steerable speech or a per-character model without adding another vendor, and who mainly synthesize English.
ReadAloud may suit you for streaming English speech at a low price per character.
What we could not confirm about OpenAI Text to Speech
- Per-character equivalent cost for gpt-4o-mini-tts (token-billed)
- Whether the speech endpoint offers WebSocket streaming (the separate Realtime API does)
- SSML support (no SSML parameter in the API reference read)
- Commercial-use statement for generated audio (not found on pages read)
- SOC 2, HIPAA and GDPR statements (not on pages read)
- Vendor-stated latency figures (none quoted)
- tts-1-hd price confirmed on pricing page and model page quick comparison as a table row; rate limit rows for tts-1-hd not read
Sources
- OpenAI API pricing (retrieved October 10, 2026)
- Text to speech guide (retrieved October 10, 2026)
- Create speech (API reference) (retrieved October 10, 2026)
- TTS-1 model page (retrieved October 10, 2026)
- GPT-4o Mini TTS model page (retrieved October 10, 2026)
- Custom voices (retrieved October 10, 2026)