Last verified October 10, 2026.
Details about Amazon Polly come from its public pages as they were on October 10, 2026. Amazon Polly may have changed them since, and a page that was unclear to us may be clear to you. Corrections: support@readaloudai.org. We make ReadAloud, so read this page with that in mind and check anything important with Amazon Polly directly.
Moving from Amazon Polly: what kind of move this is
How Amazon Polly positions itself: AWS's managed text-to-speech service offering Standard, Neural, Long-Form and Generative voice engines with SSML, lexicons and speech marks.
Amazon Polly is a cloud provider speech service, so speech calls usually sit inside a larger cloud account. Expect to change the client library and the credentials, and check whether other parts of your account depend on the same project or billing setup.
Amazon Polly's pages suggest it fits: teams on AWS who need high-volume speech for IVR, notifications or accessibility, plus speech marks metadata and HIPAA-eligible handling.
Before you leave Amazon Polly
- Collect the exact text you send to Amazon Polly today: the longest, the shortest, and the ones with names, numbers and abbreviations.
- Commercial use at Amazon Polly: Features page lists the ability to securely store and redistribute speech in standard formats; pricing page states generated speech can be cached and replayed at no additional cost. (source: aws.amazon.com)
- Condition listed by Amazon Polly: SynthesizeSpeech input is capped at 3,000 billed characters (6,000 total) and 10 minutes of audio per call. (source: docs.aws.amazon.com)
- Condition listed by Amazon Polly: The limits page lists the bidirectional StartSpeechSynthesisStream operation under Generative voice. (source: docs.aws.amazon.com)
- Condition listed by Amazon Polly: Brand Voice is offered through an engagement with AWS rather than self-serve. (source: aws.amazon.com)
- Save any audio you need to keep from Amazon Polly before you cancel, and note which plan covers the right to keep using it.
Steps to move off Amazon Polly
These steps come from Amazon Polly's public pages. Sources are listed at the end.
- Polly SynthesizeSpeech takes JSON {Engine, VoiceId, OutputFormat, Text, TextType, SampleRate, LexiconNames} signed with AWS SigV4; find each call site. (source: docs.aws.amazon.com)
- List the Polly VoiceIds you use (Joanna, Matthew, Ruth...) and the Engine you pick for each. (source: docs.aws.amazon.com)
- List where you use lexicons and SSML (TextType=ssml). (source: docs.aws.amazon.com)
- If you use OutputFormat=json (word, sentence and viseme speech marks), list what reads them, such as highlighting or lip-sync. (source: docs.aws.amazon.com)
- OutputFormat values are mp3, ogg_vorbis, ogg_opus, pcm, mulaw and alaw; list which ones you request. (source: docs.aws.amazon.com)
- Polly accepts up to 3,000 billed characters per SynthesizeSpeech call, so your code may already split long text. (source: docs.aws.amazon.com)
Map each Amazon Polly request field and endpoint
| In Amazon Polly | At ReadAloud |
|---|---|
| SynthesizeSpeech (signed with AWS SigV4) | POST /v1/audio/speech (the OpenAI-compatible route), which returns the audio in the response. Polly signs requests with AWS credentials where ReadAloud uses an API key. |
| Text | input on the OpenAI-compatible route (text on the WebSocket), up to 5,000 characters per request; split longer text at sentence boundaries |
| TextType=ssml and LexiconNames | No equivalent. ReadAloud does not accept SSML; send plain text and remove markup |
| VoiceId (Joanna, Matthew, Ruth...) with Engine | voice: readaloud-default on ReadAloud Live, or a ReadAloud Studio voice listed by GET /v1/voices. Vendor voice names and ids do not exist at ReadAloud |
| OutputFormat (mp3, ogg_vorbis, ogg_opus, pcm, mulaw, alaw) | response_format on the OpenAI-compatible route: mp3, opus, wav or pcm (aac and flac return 400); the WebSocket also offers 8 kHz mu-law and A-law |
| OutputFormat=json (speech marks: word, sentence, viseme) | No equivalent is documented. ReadAloud documents audio plus a per-sentence chunk_meta message on the WebSocket; no word or viseme marks are documented |
| SampleRate | No matching field is documented; leave it out and test the result |
What changes in your code and what does not carry over
- Endpoint: Amazon Polly's request shape is its own, so rewrite the call; /v1/audio/speech is one plain HTTP request.
- Voice: Amazon Polly voices do not exist at ReadAloud. Live: one American English voice. Studio: several English voices.
- Languages, as Amazon Polly's pages put it: 42 language and language-variant rows in the available-voices table. ReadAloud: English today. (source: aws.amazon.com)
- Formats: Amazon Polly lists mp3, ogg_vorbis, ogg_opus, pcm (16-bit mono little-endian), mulaw and alaw. ReadAloud returns mp3, opus, pcm, mulaw and alaw on at least one route.
- SSML: Amazon Polly lists SSML support. ReadAloud does not accept it; strip or rewrite markup before sending text.
- Request size: Amazon Polly states SynthesizeSpeech: up to 3,000 billed characters (6,000 total; SSML tags not billed); output stream limited to 10 minutes. Longer text uses asynchronous speech synthesis tasks. ReadAloud accepts 5,000 characters per request, so split longer text at sentence boundaries. (source: docs.aws.amazon.com)
- Custom voices, as Amazon Polly's pages describe them: Brand Voice: custom voice built with AWS on request via an AWS account manager or contact form; cost and timeline scoped per engagement. No self-serve cloning stated on pages read. Those voices cannot be exported to ReadAloud. (source: aws.amazon.com)
Steps on the ReadAloud side
- Create a key (/developers#get-started) and run the test call below.
- Pick ReadAloud Live (the low-latency tier, built for live calls and voice agents) or ReadAloud Studio (the higher-priced tier for read-aloud and narration, with several English voices).
- Keep the base URL, key and voice in a setting, and move a small share of traffic first.
# ReadAloud through its OpenAI-compatible speech route (/v1/audio/speech).
# Replace YOUR_KEY with a ReadAloud API key.
curl -X POST "https://api.readaloudai.org/v1/audio/speech" \
-H "Authorization: Bearer YOUR_KEY" -H "Content-Type: application/json" \
-d '{"model":"tts-1","voice":"readaloud-default","input":"Paste a sentence you send to Amazon Polly today.","response_format":"mp3"}' \
-o test.mp3What it costs to test ReadAloud
A one-time grant of free credits worth $0.10 (about 10,000 characters of speech), shared by all keys on the account.
Amazon Polly's Standard at $4 matches ReadAloud Live at $4 per 1M characters.
Amazon Polly's Neural at $16, Generative at $30 and Long-Form at $100 are higher than ReadAloud Live.
Amazon Polly's Standard at $4 is lower than ReadAloud Studio at $10 per 1M characters, so on list price per character Amazon Polly costs less there.
Amazon Polly's Neural at $16, Generative at $30 and Long-Form at $100 are higher than ReadAloud Studio.
These are list prices as of October 10, 2026. The tiers cover different things, so price your own volume rather than reading across.
What we could not confirm about Amazon Polly
- Total voice count (table lists voices per language; not summed)
- SDK languages (AWS SDK list not read)
- WebSocket support (bidirectional streaming described over HTTP/2)
- Quantified latency
- SOC 2 and GDPR statements
- Whether quotas are adjustable
- OpenAI-compatible mode (none found)
Sources
- Amazon Polly pricing (retrieved October 10, 2026)
- AWS Price List: Amazon Polly (us-east-1, published 2026-09-11) (retrieved October 10, 2026)
- Amazon Polly FAQs (retrieved October 10, 2026)
- Amazon Polly features (retrieved October 10, 2026)
- Limits in Amazon Polly (retrieved October 10, 2026)
- SynthesizeSpeech API reference (retrieved October 10, 2026)
- StartSpeechSynthesisStream API reference (retrieved October 10, 2026)