Last verified October 10, 2026.
Details about Cartesia come from its public pages as they were on October 10, 2026. Cartesia may have changed them since, and a page that was unclear to us may be clear to you. Corrections: support@readaloudai.org. We make ReadAloud, so read this page with that in mind and check anything important with Cartesia directly.
About Cartesia
How Cartesia positions itself: Cartesia offers the Sonic family of streaming text-to-speech models, plus speech-to-text and managed voice agents, aimed at real-time conversational use.
Price per 1 million characters
USD per one million characters, by tier. Other units are shown as the vendor writes them, not converted.
| Option | Listed price | What it covers | Against ReadAloud Live |
|---|---|---|---|
| ReadAloud Live | $4 per 1M characters | the low-latency tier, built for live calls and voice agents | Reference point |
| ReadAloud Studio | $10 per 1M characters | the higher-priced tier for read-aloud and narration, with several English voices | $6 above ReadAloud Live |
| Cartesia: Free | Vendor's own unit: 20K credits/month, $0; not converted (source: cartesia.ai) | Evaluation and personal use; no commercial license listed on the plan card. | Not comparable: Cartesia lists a different unit |
| Cartesia: Pro | $50 per 1M characters (source: cartesia.ai) | Per-1M figure derived by us: $5 per month divided by 100K credits, at about 1 credit per character (the rate Cartesia's docs state). Includes commercial use license and instant voice cloning. Overage $65 per 1M credits if enabled. Applies to Sonic-3.6; the first paid plan with a commercial license. | $46 higher than ReadAloud Live |
| Cartesia: Startup | $39.20 per 1M characters (source: cartesia.ai) | Per-1M figure derived by us: $49 per month divided by 1.25M credits, at about 1 credit per character. Adds professional voice cloning (2 slots). Overage $45 per 1M credits. | $35.20 higher than ReadAloud Live |
| Cartesia: Scale | $37.38 per 1M characters (source: cartesia.ai) | Per-1M figure derived by us: $299 per month divided by 8M credits (rounded), at about 1 credit per character. 4 PVC slots, priority support, higher concurrency. Overage $38 per 1M credits. | $33.38 higher than ReadAloud Live |
| Cartesia: Enterprise | Vendor's own unit: custom; not converted (source: cartesia.ai) | Volume pricing; BAAs/DPAs, SSO, custom concurrency. | Not comparable: Cartesia lists a different unit |
Every per-character list price we found for Cartesia, Pro at $50, Startup at $39.20 and Scale at $37.38, is higher than both ReadAloud Live ($4) and ReadAloud Studio ($10) per 1M characters.
Cartesia's own headline, as of October 10, 2026: Pro plan: $5/month for 100K credits, an effective $50 per 1M characters; larger plans lower the effective rate to about $37 per 1M.
Free tier at Cartesia: 20K credits per month (about 20K characters of Sonic TTS), 2 concurrent TTS requests. ReadAloud: A one-time grant of free credits worth $0.10 (about 10,000 characters of speech), shared by all keys on the account.
Tiers cover different things, so compare on your own monthly volume, and check Cartesia's pricing page before you decide. Current ReadAloud prices: /developers.
Side by side
"Not stated on the pages we reviewed" marks what Cartesia's pages did not say. We print no head-to-head latency number.
| Topic | ReadAloud | Cartesia |
|---|---|---|
| Voices | Live: one American English voice. Studio: several English voices. | Not stated on the pages we reviewed (source: docs.cartesia.ai) |
| Languages | English today | 44 languages for Sonic 3.6 (vendor-stated) (source: docs.cartesia.ai) |
| Custom voices or cloning | API, with a consent step and a payment method | Instant voice cloning from a short clip (Pro and above); professional voice cloning on Startup and above; clones require verified consent. (source: docs.cartesia.ai) |
| Streaming | WebSocket and HTTP streaming | WebSocket: yes; HTTP streaming: yes (source: cartesia.ai) |
| Time to first audio | ReadAloud Live: about 250 ms (our measurement, median to first audio byte, warm connection, San Jose; October 10, 2026; /developers#vs-elevenlabs) | Vendor-stated: first audio in under 90 ms (model latency), 'sub-90ms'. Not independently measured. (source: cartesia.ai) |
| Output formats | mp3, opus, wav, pcm; 8 kHz mu-law and A-law over WebSocket | wav, mp3, raw: pcm_f32le, raw: pcm_s16le, raw: pcm_mulaw, raw: pcm_alaw |
| SSML | No SSML and no audio tags. | Not stated on the pages we reviewed |
| OpenAI-compatible speech endpoint | Yes, /v1/audio/speech | Not stated on the pages we reviewed |
| Limit per request | 5,000 characters | Not stated on the pages we reviewed |
| Concurrency | Up to 12 Live streams per server; more start under load | TTS concurrent requests by plan: Free 2, Pro 3, Startup 5, Scale 15, Enterprise custom. Exceeding returns 429. (source: docs.cartesia.ai) |
| SDKs and plugins | Python and JavaScript libraries (readaloud), Pipecat and LiveKit plugins, any OpenAI SDK | JavaScript/TypeScript (@cartesia/cartesia-js) and Python (cartesia) |
| HIPAA | Not claimed | Mentioned on their pages; see the note below the table for what it covers (source: cartesia.ai) |
| SOC 2 | Not claimed | Mentioned on their pages; see the note below the table for what it covers (source: cartesia.ai) |
| GDPR | Not claimed | Mentioned on their pages; see the note below the table for what it covers (source: cartesia.ai) |
About Cartesia's compliance wording: Self-attested on the Sonic product page: SOC 2 Type 2, HIPAA-eligible with BAAs, GDPR; BAAs/DPAs on the Enterprise plan. Zero data retention is described as available for enterprise customers.
What Cartesia lists as strengths
- Vendor states Sonic delivers first audio in under 90 ms (vendor-stated, not independently verified). (source: cartesia.ai)
- Sonic 3.6 supports 44 languages from one model. (source: docs.cartesia.ai)
- Offers instant and professional voice cloning plus an SSML-like tag set for speed, volume and emotion. (source: cartesia.ai)
- Publishes SOC 2 Type 2, HIPAA-eligibility with BAAs, and GDPR statements, with on-prem/VPC deployment for enterprise. (source: cartesia.ai)
- Failed requests do not consume credits; official Python and JavaScript SDKs and LiveKit/Pipecat integrations. (source: docs.cartesia.ai)
Conditions to check with Cartesia
Limits or conditions from the pages we reviewed.
- Billing is a monthly subscription with a credit pool; Pro's effective rate is $50 per 1M characters and overage rates are $38 to $65 per 1M credits depending on plan. (source: cartesia.ai)
- Free-plan concurrency is 2 TTS requests and Pro is 3; higher concurrency requires Startup or above. (source: docs.cartesia.ai)
Cartesia usage terms
From the vendor pages we reviewed; not legal advice.
- Commercial use: Commercial use license is listed from the Pro plan upward; the Free plan is described as for evaluation and personal use. (source: cartesia.ai)
- Attribution or conditions: Voice clones need verified consent; cloning people without permission is prohibited by the Terms of Use (vendor-stated). (source: cartesia.ai)
Where ReadAloud is different
- Voices: we could not read a voice count on Cartesia's pages, so we make no comparison of choice.
- Custom voices, as Cartesia's pages describe them: Instant voice cloning from a short clip (Pro and above); professional voice cloning on Startup and above; clones require verified consent. (source: docs.cartesia.ai)
- Compliance: Cartesia's pages mention HIPAA, SOC 2 and GDPR; ReadAloud claims none. (source: cartesia.ai)
Which one fits
Cartesia may suit you if this describes you: teams building real-time voice agents or multilingual products who want a model Cartesia describes as low-latency, with 44 languages, cloning options and published compliance statements, and who are comfortable with a credit-based subscription.
ReadAloud may suit you for streaming English speech at a low price per character.
What we could not confirm about Cartesia
- Total number of voices in the voice library
- Maximum characters per TTS request
- SSML support is a Cartesia-specific tag subset (speed, volume, break/emotion), not full W3C SSML
- Whether any OpenAI-compatible speech route exists
- Credit use is 'approximately' 1 per character and may vary slightly with pre-processing
- Plan prices and credit allotments are single-source (cartesia.ai/pricing); the docs pricing page confirms just the ~1 credit per character rate. Per-1M figures are derived by dividing plan price by credits
Sources
- Cartesia pricing page (retrieved October 10, 2026)
- Cartesia docs: Pricing (retrieved October 10, 2026)
- Cartesia docs: Sonic 3.6 model (retrieved October 10, 2026)
- Cartesia Sonic product page (retrieved October 10, 2026)
- Cartesia docs: Concurrency and WebSocket limits (retrieved October 10, 2026)
- Cartesia API reference: TTS bytes (retrieved October 10, 2026)
- Cartesia docs: Client libraries (retrieved October 10, 2026)