Last verified October 10, 2026.
Details about Google Cloud Text-to-Speech come from its public pages as they were on October 10, 2026. Google Cloud Text-to-Speech may have changed them since, and a page that was unclear to us may be clear to you. Corrections: support@readaloudai.org. We make ReadAloud, so read this page with that in mind and check anything important with Google Cloud Text-to-Speech directly.
About Google Cloud Text-to-Speech
How Google Cloud Text-to-Speech positions itself: Google Cloud's text-to-speech API with a wide voice and language catalogue spanning Standard voices, WaveNet and Neural2, Chirp 3 HD voices, instant custom voices and token-priced Gemini-TTS models.
Price per 1 million characters
USD per one million characters, by tier. Other units are shown as the vendor writes them, not converted.
| Option | Listed price | What it covers | Against ReadAloud Live |
|---|---|---|---|
| ReadAloud Live | $4 per 1M characters | the low-latency tier, built for live calls and voice agents | Reference point |
| ReadAloud Studio | $10 per 1M characters | the higher-priced tier for read-aloud and narration, with several English voices | $6 above ReadAloud Live |
| Google Cloud Text-to-Speech: Standard voices | $4 per 1M characters (source: cloud.google.com) | Free: 0 to 4 million characters per month. | Same as ReadAloud Live |
| Google Cloud Text-to-Speech: WaveNet voices | $4 per 1M characters (source: cloud.google.com) | Page lists the same SKU id as Standard and the same $4 price; free 0 to 4 million characters per month. See the list of things we could not confirm below. | Same as ReadAloud Live |
| Google Cloud Text-to-Speech: Neural2 voices | $16 per 1M characters (source: cloud.google.com) | Free: 0 to 1 million characters per month. | $12 higher than ReadAloud Live |
| Google Cloud Text-to-Speech: Polyglot (Preview) voices | $16 per 1M characters (source: cloud.google.com) | Free: 0 to 1 million characters per month. | $12 higher than ReadAloud Live |
| Google Cloud Text-to-Speech: Studio voices | $160 per 1M characters (source: cloud.google.com) | Free: 0 to 1 million characters per month. | $156 higher than ReadAloud Live |
| Google Cloud Text-to-Speech: Chirp 3: HD voices | $30 per 1M characters (source: cloud.google.com) | Free: 0 to 1 million characters per month. No SSML, speaking rate or pitch parameters. | $26 higher than ReadAloud Live |
| Google Cloud Text-to-Speech: Chirp 3: Instant custom voice | $60 per 1M characters (source: cloud.google.com) | No free allowance ('Not available'). | $56 higher than ReadAloud Live |
| Google Cloud Text-to-Speech: Gemini 3.8 Flash TTS (Preview) | Vendor's own unit: USD per 1M tokens: $0.50 text input, $9.00 audio output through 2026-12-31; $1.00 and $18.00 from 2027-01-01; not converted (source: cloud.google.com) | Audio tokens are 25 per second of audio. Token-billed, so no per-character price. | Not comparable: Google Cloud Text-to-Speech lists a different unit |
| Google Cloud Text-to-Speech: Gemini 3.8 Flash-Lite TTS (Preview) | Vendor's own unit: USD per 1M tokens: $0.50 text input, $6.00 audio output through 2026-12-31; $1.00 and $12.00 from 2027-01-01; not converted (source: cloud.google.com) | Token-billed. | Not comparable: Google Cloud Text-to-Speech lists a different unit |
| Google Cloud Text-to-Speech: Gemini 2.5 Flash TTS / Flash-Lite Preview TTS | Vendor's own unit: USD per 1M tokens: $0.50 text input, $10.00 audio output; not converted (source: cloud.google.com) | Token-billed. | Not comparable: Google Cloud Text-to-Speech lists a different unit |
| Google Cloud Text-to-Speech: Gemini 3.1 Flash TTS (Preview) | Vendor's own unit: USD per 1M tokens: $1.00 text input, $20.00 audio output; not converted (source: cloud.google.com) | Token-billed. | Not comparable: Google Cloud Text-to-Speech lists a different unit |
| Google Cloud Text-to-Speech: Gemini 2.5 Pro TTS | Vendor's own unit: USD per 1M tokens: $1.00 text input, $20.00 audio output; not converted (source: cloud.google.com) | Token-billed. | Not comparable: Google Cloud Text-to-Speech lists a different unit |
Google Cloud Text-to-Speech's Standard voices at $4 and WaveNet voices at $4 match ReadAloud Live at $4 per 1M characters.
Google Cloud Text-to-Speech's Neural2 voices at $16, Polyglot (Preview) voices at $16, Studio voices at $160, Chirp 3: HD voices at $30 and Chirp 3: Instant custom voice at $60 are higher than ReadAloud Live.
Google Cloud Text-to-Speech's Standard voices at $4 and WaveNet voices at $4 are lower than ReadAloud Studio at $10 per 1M characters, so on list price per character Google Cloud Text-to-Speech costs less there.
Google Cloud Text-to-Speech's Neural2 voices at $16, Polyglot (Preview) voices at $16, Studio voices at $160, Chirp 3: HD voices at $30 and Chirp 3: Instant custom voice at $60 are higher than ReadAloud Studio.
Google Cloud Text-to-Speech's own headline, as of October 10, 2026: Standard voices $4 per 1M characters (first 4M characters per month free); Neural2 $16; Chirp 3: HD $30; Studio $160; Instant custom voice $60.
Free tier at Google Cloud Text-to-Speech: Monthly free characters: 4 million for Standard and WaveNet voices; 1 million each for Neural2, Polyglot (Preview), Studio and Chirp 3: HD. No free allowance for Instant custom voice or Gemini-TTS. New Google Cloud customers also get $300 in free credits (stated on the bidirectional streaming quickstart). ReadAloud: A one-time grant of free credits worth $0.10 (about 10,000 characters of speech), shared by all keys on the account.
- Time-limited or scheduled pricing listed when we reviewed: Gemini 3.8 Flash TTS and Flash-Lite TTS (Preview) list a lower price through 2026-12-31 and double on 2027-01-01 (text input $0.50 to $1.00; Flash audio output $9 to $18; Flash-Lite audio output $6 to $12 per 1M tokens). (source: cloud.google.com)
Tiers cover different things, so compare on your own monthly volume, and check Google Cloud Text-to-Speech's pricing page before you decide. Current ReadAloud prices: /developers.
Side by side
"Not stated on the pages we reviewed" marks what Google Cloud Text-to-Speech's pages did not say. We print no head-to-head latency number.
| Topic | ReadAloud | Google Cloud Text-to-Speech |
|---|---|---|
| Voices | Live: one American English voice. Studio: several English voices. | 380+ voices (product page); Chirp 3: HD voices come in 30 distinct styles (source: cloud.google.com) |
| Languages | English today | 75+ languages and variants (product page) (source: cloud.google.com) |
| Custom voices or cloning | API, with a consent step and a payment method | Chirp 3: Instant Custom Voice creates a personal voice from a recording; a consent statement in the speaker's language ('I am the owner of this voice and I consent to Google using this voice...') is required. Supports streaming and several languages; priced at $60 per 1M characters. Access conditions beyond that were not stated on the page read. (source: cloud.google.com) |
| Streaming | WebSocket and HTTP streaming | Google documents bidirectional streaming synthesis as a Preview feature under Pre-GA terms; whether the transport is WebSocket or HTTP was not confirmed on the pages read. (source: cloud.google.com) |
| Time to first audio | ReadAloud Live: about 250 ms (our measurement, median to first audio byte, warm connection, San Jose; October 10, 2026; /developers#vs-elevenlabs) | Not stated on the pages we reviewed |
| Output formats | mp3, opus, wav, pcm; 8 kHz mu-law and A-law over WebSocket | LINEAR16 (WAV header), MP3, OGG_OPUS, MULAW, ALAW, PCM (streaming) |
| SSML | No SSML and no audio tags. | Yes |
| OpenAI-compatible speech endpoint | Yes, /v1/audio/speech | No. Native API is POST https://texttospeech.googleapis.com/v1/text:synthesize with input, voice (languageCode, name) and audioConfig (audioEncoding); OAuth or Google credentials rather than an OpenAI-style bearer key. No OpenAI-compatible mode found in the pages read. SSML is supported on most voice types but not Chirp 3: HD. |
| Limit per request | 5,000 characters | 5,000 bytes per request (multi-byte characters such as ja-JP use more than one byte) (source: cloud.google.com) |
| Concurrency | Up to 12 Live streams per server; more start under load | Quotas per project: 100 concurrent streaming sessions; 200 requests per minute for Chirp 3; 500 for Studio; 1,000 for Neural2, Polyglot and default voices; 30 per minute for Chirp voice cloning. Gemini-TTS 2.5: 150 (Flash) and 125 (Pro) queries per minute. Limits subject to change. (source: cloud.google.com) |
| SDKs and plugins | Python and JavaScript libraries (readaloud), Pipecat and LiveKit plugins, any OpenAI SDK | Google Cloud client libraries (languages not enumerated in pages read) and REST (v1, v1beta1) |
| HIPAA | Not claimed | Not stated on the pages we reviewed |
| SOC 2 | Not claimed | Not stated on the pages we reviewed |
| GDPR | Not claimed | Not stated on the pages we reviewed |
About Google Cloud Text-to-Speech's compliance wording: No product-specific compliance statements found on the Text-to-Speech pages read; Google Cloud publishes compliance information at the platform level that was not reviewed here.
What Google Cloud Text-to-Speech lists as strengths
- Broad catalogue: 380+ voices across 75+ languages and variants. (source: cloud.google.com)
- Entry price: Standard and WaveNet voices at $4 per 1M characters with 4M free characters per month. (source: cloud.google.com)
- Chirp 3: HD voices (30 styles), which Google describes as designed for conversational agents with low-latency streaming, plus an instant custom voice option with a consent statement. (source: cloud.google.com)
- SSML support on Studio, Neural2, WaveNet and Standard voices, and output formats including MULAW and ALAW for telephony. (source: cloud.google.com)
- Gemini-TTS models give text-prompt control over delivery. (source: cloud.google.com)
Conditions to check with Google Cloud Text-to-Speech
Limits or conditions from the pages we reviewed.
- Chirp 3: HD voices do not support SSML input, speaking rate and pitch parameters, or A-law encoding. (source: cloud.google.com)
- Requests are limited to 5,000 bytes. (source: cloud.google.com)
- Bidirectional streaming synthesis is documented as a Preview feature under Pre-GA terms. (source: cloud.google.com)
Google Cloud Text-to-Speech usage terms
From the vendor pages we reviewed; not legal advice.
- Commercial use: Quotas page: audio files created with the service may be used in applications or media in compliance with the Google Cloud Platform Terms of Service and applicable law. (source: cloud.google.com)
Where ReadAloud is different
- Voices: Google Cloud Text-to-Speech lists 380+ voices (product page); Chirp 3: HD voices come in 30 distinct styles. ReadAloud: Live: one American English voice. Studio: several English voices. (source: cloud.google.com)
- SSML: Google Cloud Text-to-Speech lists SSML support; ReadAloud does not accept it, so markup in your text would have to be removed or rewritten.
- Custom voices, as Google Cloud Text-to-Speech's pages describe them: Chirp 3: Instant Custom Voice creates a personal voice from a recording; a consent statement in the speaker's language ('I am the owner of this voice and I consent to Google using this voice...') is required. Supports streaming and several languages; priced at $60 per 1M characters. Access conditions beyond that were not stated on the page read. (source: cloud.google.com)
Which one fits
Google Cloud Text-to-Speech may suit you if this describes you: teams already on Google Cloud, or teams that need broad language and voice coverage, with Standard and WaveNet voices at $4 per 1M characters and optional Chirp 3 HD voices.
ReadAloud may suit you for streaming English speech at a low price per character.
What we could not confirm about Google Cloud Text-to-Speech
- All Google prices are single-source (the pricing page); no second Google page repeated the numbers
- WaveNet is listed at $4 per 1M characters with the same SKU id as Standard; this may be a pricing-page quirk, and we did not check the pricing calculator
- Whether streaming is WebSocket or HTTP (bidirectional streaming doc is a Preview; transport not confirmed)
- Vendor-stated latency figures (none with numbers found)
- Access requirements for Instant Custom Voice
- Compliance statements (HIPAA, SOC 2, GDPR) for this product
- Client library languages
- OpenAI-compatible mode (none found)
Sources
- Text-to-Speech pricing (retrieved October 10, 2026)
- Text-to-Speech product page (retrieved October 10, 2026)
- Quotas and limits (retrieved October 10, 2026)
- Supported voices and languages (retrieved October 10, 2026)
- Text-to-Speech basics (retrieved October 10, 2026)
- Chirp 3: Instant Custom Voice (retrieved October 10, 2026)
- Bidirectional streaming quickstart (retrieved October 10, 2026)
- AudioEncoding reference (retrieved October 10, 2026)