Cartesia · Audio
Sonic 3 API pricing
cartesia/sonic-3
Low-latency text to speech. Sonic 3 by Cartesia, called through one INFRO key alongside every other model you run. We price it 33% below Cartesia's direct rate — the same model, the same weights, reached through a different integration.
Pricing
Cartesia direct
$55
per 1M chars
INFRO
$37
per 1M chars
You save
33%
on every unit
Illustrative launch rates, reconciled against published provider pricing on 2026-08-24. Live rates come from GET /v1/models once your key is active.
Calling Sonic 3
POST /v1/audio/speechAt launch — Committed to the first release. In build now, not usable yet.import os
from infro import Infro
client = Infro(api_key=os.environ["INFRO_API_KEY"])
audio = client.audio.speech.create(
model="cartesia/sonic-3",
input="Every AI model, one API.",
)What this modality supports
- Text to speech
- Music generation
- Transcription
- Diarization
Text-to-speech with streaming playback, full-track music generation, and transcription with timestamps and diarization. Billed per character generated, per track, and per hour transcribed.
Alternatives for low-latency text to speech
Same modality, nearest in price. All of them share the request shape above, so comparing them is a one-word change.
Work out what Sonic 3 would cost you
Put your current spend on this model into the calculator and see the same workload priced here — or have us read your real usage.