Moving from Amazon Polly to ReadAloud

A step-by-step checklist for moving text-to-speech calls from Amazon Polly to ReadAloud, based on Amazon Polly's public pages as of October 10, 2026.

Last verified October 10, 2026.

Details about Amazon Polly come from its public pages as they were on October 10, 2026. Amazon Polly may have changed them since, and a page that was unclear to us may be clear to you. Corrections: support@readaloudai.org. We make ReadAloud, so read this page with that in mind and check anything important with Amazon Polly directly.

Moving from Amazon Polly: what kind of move this is

How Amazon Polly positions itself: AWS's managed text-to-speech service offering Standard, Neural, Long-Form and Generative voice engines with SSML, lexicons and speech marks.

Amazon Polly is a cloud provider speech service, so speech calls usually sit inside a larger cloud account. Expect to change the client library and the credentials, and check whether other parts of your account depend on the same project or billing setup.

Amazon Polly's pages suggest it fits: teams on AWS who need high-volume speech for IVR, notifications or accessibility, plus speech marks metadata and HIPAA-eligible handling.

Before you leave Amazon Polly

  • Collect the exact text you send to Amazon Polly today: the longest, the shortest, and the ones with names, numbers and abbreviations.
  • Commercial use at Amazon Polly: Features page lists the ability to securely store and redistribute speech in standard formats; pricing page states generated speech can be cached and replayed at no additional cost. (source: aws.amazon.com)
  • Condition listed by Amazon Polly: SynthesizeSpeech input is capped at 3,000 billed characters (6,000 total) and 10 minutes of audio per call. (source: docs.aws.amazon.com)
  • Condition listed by Amazon Polly: The limits page lists the bidirectional StartSpeechSynthesisStream operation under Generative voice. (source: docs.aws.amazon.com)
  • Condition listed by Amazon Polly: Brand Voice is offered through an engagement with AWS rather than self-serve. (source: aws.amazon.com)
  • Save any audio you need to keep from Amazon Polly before you cancel, and note which plan covers the right to keep using it.

Steps to move off Amazon Polly

These steps come from Amazon Polly's public pages. Sources are listed at the end.

  1. Polly SynthesizeSpeech takes JSON {Engine, VoiceId, OutputFormat, Text, TextType, SampleRate, LexiconNames} signed with AWS SigV4; find each call site. (source: docs.aws.amazon.com)
  2. List the Polly VoiceIds you use (Joanna, Matthew, Ruth...) and the Engine you pick for each. (source: docs.aws.amazon.com)
  3. List where you use lexicons and SSML (TextType=ssml). (source: docs.aws.amazon.com)
  4. If you use OutputFormat=json (word, sentence and viseme speech marks), list what reads them, such as highlighting or lip-sync. (source: docs.aws.amazon.com)
  5. OutputFormat values are mp3, ogg_vorbis, ogg_opus, pcm, mulaw and alaw; list which ones you request. (source: docs.aws.amazon.com)
  6. Polly accepts up to 3,000 billed characters per SynthesizeSpeech call, so your code may already split long text. (source: docs.aws.amazon.com)

Map each Amazon Polly request field and endpoint

Amazon Polly fields and endpoints and what they become at ReadAloud
In Amazon PollyAt ReadAloud
SynthesizeSpeech (signed with AWS SigV4)POST /v1/audio/speech (the OpenAI-compatible route), which returns the audio in the response. Polly signs requests with AWS credentials where ReadAloud uses an API key.
Textinput on the OpenAI-compatible route (text on the WebSocket), up to 5,000 characters per request; split longer text at sentence boundaries
TextType=ssml and LexiconNamesNo equivalent. ReadAloud does not accept SSML; send plain text and remove markup
VoiceId (Joanna, Matthew, Ruth...) with Enginevoice: readaloud-default on ReadAloud Live, or a ReadAloud Studio voice listed by GET /v1/voices. Vendor voice names and ids do not exist at ReadAloud
OutputFormat (mp3, ogg_vorbis, ogg_opus, pcm, mulaw, alaw)response_format on the OpenAI-compatible route: mp3, opus, wav or pcm (aac and flac return 400); the WebSocket also offers 8 kHz mu-law and A-law
OutputFormat=json (speech marks: word, sentence, viseme)No equivalent is documented. ReadAloud documents audio plus a per-sentence chunk_meta message on the WebSocket; no word or viseme marks are documented
SampleRateNo matching field is documented; leave it out and test the result

What changes in your code and what does not carry over

  • Endpoint: Amazon Polly's request shape is its own, so rewrite the call; /v1/audio/speech is one plain HTTP request.
  • Voice: Amazon Polly voices do not exist at ReadAloud. Live: one American English voice. Studio: several English voices.
  • Languages, as Amazon Polly's pages put it: 42 language and language-variant rows in the available-voices table. ReadAloud: English today. (source: aws.amazon.com)
  • Formats: Amazon Polly lists mp3, ogg_vorbis, ogg_opus, pcm (16-bit mono little-endian), mulaw and alaw. ReadAloud returns mp3, opus, pcm, mulaw and alaw on at least one route.
  • SSML: Amazon Polly lists SSML support. ReadAloud does not accept it; strip or rewrite markup before sending text.
  • Request size: Amazon Polly states SynthesizeSpeech: up to 3,000 billed characters (6,000 total; SSML tags not billed); output stream limited to 10 minutes. Longer text uses asynchronous speech synthesis tasks. ReadAloud accepts 5,000 characters per request, so split longer text at sentence boundaries. (source: docs.aws.amazon.com)
  • Custom voices, as Amazon Polly's pages describe them: Brand Voice: custom voice built with AWS on request via an AWS account manager or contact form; cost and timeline scoped per engagement. No self-serve cloning stated on pages read. Those voices cannot be exported to ReadAloud. (source: aws.amazon.com)

Steps on the ReadAloud side

  1. Create a key (/developers#get-started) and run the test call below.
  2. Pick ReadAloud Live (the low-latency tier, built for live calls and voice agents) or ReadAloud Studio (the higher-priced tier for read-aloud and narration, with several English voices).
  3. Keep the base URL, key and voice in a setting, and move a small share of traffic first.
A first test call to ReadAloud (works from any language that can send an HTTP request)
# ReadAloud through its OpenAI-compatible speech route (/v1/audio/speech).
# Replace YOUR_KEY with a ReadAloud API key.
curl -X POST "https://api.readaloudai.org/v1/audio/speech" \
  -H "Authorization: Bearer YOUR_KEY" -H "Content-Type: application/json" \
  -d '{"model":"tts-1","voice":"readaloud-default","input":"Paste a sentence you send to Amazon Polly today.","response_format":"mp3"}' \
  -o test.mp3

What it costs to test ReadAloud

A one-time grant of free credits worth $0.10 (about 10,000 characters of speech), shared by all keys on the account.

Amazon Polly's Standard at $4 matches ReadAloud Live at $4 per 1M characters.

Amazon Polly's Neural at $16, Generative at $30 and Long-Form at $100 are higher than ReadAloud Live.

Amazon Polly's Standard at $4 is lower than ReadAloud Studio at $10 per 1M characters, so on list price per character Amazon Polly costs less there.

Amazon Polly's Neural at $16, Generative at $30 and Long-Form at $100 are higher than ReadAloud Studio.

These are list prices as of October 10, 2026. The tiers cover different things, so price your own volume rather than reading across.

What we could not confirm about Amazon Polly

  • Total voice count (table lists voices per language; not summed)
  • SDK languages (AWS SDK list not read)
  • WebSocket support (bidirectional streaming described over HTTP/2)
  • Quantified latency
  • SOC 2 and GDPR statements
  • Whether quotas are adjustable
  • OpenAI-compatible mode (none found)

Sources