Moving from Google Cloud Text-to-Speech to ReadAloud

A step-by-step checklist for moving text-to-speech calls from Google Cloud Text-to-Speech to ReadAloud, based on Google Cloud Text-to-Speech's public pages as of October 10, 2026.

Last verified October 10, 2026.

Details about Google Cloud Text-to-Speech come from its public pages as they were on October 10, 2026. Google Cloud Text-to-Speech may have changed them since, and a page that was unclear to us may be clear to you. Corrections: support@readaloudai.org. We make ReadAloud, so read this page with that in mind and check anything important with Google Cloud Text-to-Speech directly.

Moving from Google Cloud Text-to-Speech: what kind of move this is

How Google Cloud Text-to-Speech positions itself: Google Cloud's text-to-speech API with a wide voice and language catalogue spanning Standard voices, WaveNet and Neural2, Chirp 3 HD voices, instant custom voices and token-priced Gemini-TTS models.

Google Cloud Text-to-Speech is a cloud provider speech service, so speech calls usually sit inside a larger cloud account. Expect to change the client library and the credentials, and check whether other parts of your account depend on the same project or billing setup.

Google Cloud Text-to-Speech's pages suggest it fits: teams already on Google Cloud, or teams that need broad language and voice coverage, with Standard and WaveNet voices at $4 per 1M characters and optional Chirp 3 HD voices.

Before you leave Google Cloud Text-to-Speech

  • Collect the exact text you send to Google Cloud Text-to-Speech today: the longest, the shortest, and the ones with names, numbers and abbreviations.
  • Commercial use at Google Cloud Text-to-Speech: Quotas page: audio files created with the service may be used in applications or media in compliance with the Google Cloud Platform Terms of Service and applicable law. (source: cloud.google.com)
  • List where your code uses Google Cloud client libraries (languages not enumerated in pages read) or REST (v1, v1beta1) so you can replace each call site.
  • Condition listed by Google Cloud Text-to-Speech: Chirp 3: HD voices do not support SSML input, speaking rate and pitch parameters, or A-law encoding. (source: cloud.google.com)
  • Condition listed by Google Cloud Text-to-Speech: Requests are limited to 5,000 bytes. (source: cloud.google.com)
  • Condition listed by Google Cloud Text-to-Speech: Bidirectional streaming synthesis is documented as a Preview feature under Pre-GA terms. (source: cloud.google.com)
  • Save any audio you need to keep from Google Cloud Text-to-Speech before you cancel, and note which plan covers the right to keep using it.

Steps to move off Google Cloud Text-to-Speech

These steps come from Google Cloud Text-to-Speech's public pages. Sources are listed at the end.

  1. Google uses POST /v1/text:synthesize with JSON {input:{text|ssml}, voice:{languageCode, name}, audioConfig:{audioEncoding}}; each call site builds this body. (source: cloud.google.com)
  2. Google uses OAuth access tokens or service-account credentials rather than an API key; find where your code obtains them. (source: cloud.google.com)
  3. List the voice names you use (for example en-US-Chirp3-HD-Charon or en-GB-Neural2-A) and the languageCode paired with each. (source: cloud.google.com)
  4. If input uses ssml, find every place that builds SSML (Chirp 3: HD voices do not accept SSML either). (source: cloud.google.com)
  5. audioEncoding values LINEAR16, MP3, OGG_OPUS, MULAW and ALAW are Google's names; list which ones you request. (source: cloud.google.com)
  6. Google counts SSML tags as characters (except mark), so character totals will differ if you stop sending markup. (source: cloud.google.com)

Map each Google Cloud Text-to-Speech request field and endpoint

Google Cloud Text-to-Speech fields and endpoints and what they become at ReadAloud
In Google Cloud Text-to-SpeechAt ReadAloud
POST https://texttospeech.googleapis.com/v1/text:synthesizePOST /v1/audio/speech (the OpenAI-compatible route), which returns the audio in the response. The request body is different, so each call site is rewritten; Google uses OAuth or service-account credentials where ReadAloud uses an API key.
input.textinput on the OpenAI-compatible route (text on the WebSocket), up to 5,000 characters per request; split longer text at sentence boundaries
input.ssmlNo equivalent. ReadAloud does not accept SSML; send plain text and remove markup
voice.name (for example en-US-Chirp3-HD-Charon) with voice.languageCodevoice: readaloud-default on ReadAloud Live, or a ReadAloud Studio voice listed by GET /v1/voices. Vendor voice names and ids do not exist at ReadAloud
audioConfig.audioEncoding (LINEAR16, MP3, OGG_OPUS, MULAW, ALAW)response_format on the OpenAI-compatible route: mp3, opus, wav or pcm (aac and flac return 400); the WebSocket also offers 8 kHz mu-law and A-law
audioConfig.speakingRatespeed, 0.25 to 4.0 on the OpenAI-compatible route
audioConfig.pitchNo matching field is documented; leave it out and test the result

What changes in your code and what does not carry over

  • Endpoint: Google Cloud Text-to-Speech's request shape is its own, so rewrite the call; /v1/audio/speech is one plain HTTP request.
  • Voice: Google Cloud Text-to-Speech voices do not exist at ReadAloud. Live: one American English voice. Studio: several English voices.
  • Languages, as Google Cloud Text-to-Speech's pages put it: 75+ languages and variants (product page). ReadAloud: English today. (source: cloud.google.com)
  • Formats: Google Cloud Text-to-Speech lists LINEAR16 (WAV header), MP3, OGG_OPUS, MULAW, ALAW and PCM (streaming). ReadAloud returns wav, pcm, mp3, opus, mulaw and alaw on at least one route.
  • SSML: Google Cloud Text-to-Speech lists SSML support. ReadAloud does not accept it; strip or rewrite markup before sending text.
  • Request size: Google Cloud Text-to-Speech states 5,000 bytes per request (multi-byte characters such as ja-JP use more than one byte). ReadAloud accepts 5,000 characters per request, so split longer text at sentence boundaries. (source: cloud.google.com)
  • Custom voices, as Google Cloud Text-to-Speech's pages describe them: Chirp 3: Instant Custom Voice creates a personal voice from a recording; a consent statement in the speaker's language ('I am the owner of this voice and I consent to Google using this voice...') is required. Supports streaming and several languages; priced at $60 per 1M characters. Access conditions beyond that were not stated on the page read. Those voices cannot be exported to ReadAloud. (source: cloud.google.com)

Steps on the ReadAloud side

  1. Create a key (/developers#get-started) and run the test call below.
  2. Pick ReadAloud Live (the low-latency tier, built for live calls and voice agents) or ReadAloud Studio (the higher-priced tier for read-aloud and narration, with several English voices).
  3. Keep the base URL, key and voice in a setting, and move a small share of traffic first.
A first test call to ReadAloud (works from any language that can send an HTTP request)
# ReadAloud through its OpenAI-compatible speech route (/v1/audio/speech).
# Replace YOUR_KEY with a ReadAloud API key.
curl -X POST "https://api.readaloudai.org/v1/audio/speech" \
  -H "Authorization: Bearer YOUR_KEY" -H "Content-Type: application/json" \
  -d '{"model":"tts-1","voice":"readaloud-default","input":"Paste a sentence you send to Google Cloud Text-to-Speech today.","response_format":"mp3"}' \
  -o test.mp3

What it costs to test ReadAloud

A one-time grant of free credits worth $0.10 (about 10,000 characters of speech), shared by all keys on the account.

Google Cloud Text-to-Speech's Standard voices at $4 and WaveNet voices at $4 match ReadAloud Live at $4 per 1M characters.

Google Cloud Text-to-Speech's Neural2 voices at $16, Polyglot (Preview) voices at $16, Studio voices at $160, Chirp 3: HD voices at $30 and Chirp 3: Instant custom voice at $60 are higher than ReadAloud Live.

Google Cloud Text-to-Speech's Standard voices at $4 and WaveNet voices at $4 are lower than ReadAloud Studio at $10 per 1M characters, so on list price per character Google Cloud Text-to-Speech costs less there.

Google Cloud Text-to-Speech's Neural2 voices at $16, Polyglot (Preview) voices at $16, Studio voices at $160, Chirp 3: HD voices at $30 and Chirp 3: Instant custom voice at $60 are higher than ReadAloud Studio.

These are list prices as of October 10, 2026. The tiers cover different things, so price your own volume rather than reading across.

What we could not confirm about Google Cloud Text-to-Speech

  • All Google prices are single-source (the pricing page); no second Google page repeated the numbers
  • WaveNet is listed at $4 per 1M characters with the same SKU id as Standard; this may be a pricing-page quirk, and we did not check the pricing calculator
  • Whether streaming is WebSocket or HTTP (bidirectional streaming doc is a Preview; transport not confirmed)
  • Vendor-stated latency figures (none with numbers found)
  • Access requirements for Instant Custom Voice
  • Compliance statements (HIPAA, SOC 2, GDPR) for this product
  • Client library languages
  • OpenAI-compatible mode (none found)

Sources