Gemini 3.8 Live Is 10× Cheaper Than GPT-6 Astra for Voice Agents — What Developers Need to Know
Google dropped two new voice models on September 15 that change the cost math for anyone building real-time voice agents. Gemini 3.8 Live and its reasoning-capable sibling, Gemini 3.8 Live Extended Thinking, are now generally available via the Gemini API — and the pricing gap versus GPT-6 Astra's voice layer is hard to ignore once you run the numbers. This isn't a minor incremental release. Extended Thinking adds genuine background reasoning during an active voice conversation, something that used to require breaking out of the audio loop entirely. Whether that matters for your use case depends on what you're building, but the cost story alone is worth understanding before you start a new voice project.

Photo by Olia Danilevich on Pexels
How the New Models Work
Google's new Live models are audio-to-audio — voice in, voice out — with background reasoning happening while the model continues to speak. Both models support real-time conversation, async function calling, visual context inputs, and more than 97 languages. The Extended Thinking variant runs multi-step chain-of-thought during a live conversation continuously in the background, rather than making the user wait for a separate processing pass.
That distinction is the whole point. In earlier architectures, if your agent needed to reason through a customer's account history or verify a multi-step request, it had to stop talking, spin up a separate reasoning call, and then resume. Users heard the dead air. Extended Thinking overlaps the reasoning with the conversation, so the agent can keep the interaction natural while it works through logic underneath.
The Benchmarks
The benchmarks are solid across the board: 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice, 35.1% on Sierra's τ-Voice-banking benchmark, and 97.7% on Big Bench Audio. These aren't cherry-picked synthetic tests — they measure latency, naturalness, and accuracy together, which is closer to what actually matters in a production voice agent than any single-dimension leaderboard.
The banking number (35.1%) is the one worth watching if you're in a regulated or high-stakes domain. Voice-banking tasks demand strict accuracy and correct function calls, and no current model handles them cleanly. Read that as: fine for balance checks and routing, still risky for anything involving money movement without a human in the loop.
The Numbers
The headline story is the price. Here's how Gemini 3.8 Live stacks up against GPT-6 Astra's voice layer:
| Model | Audio Input | Audio Output | Reasoning | Languages |
|---|---|---|---|---|
| Gemini 3.8 Live | $0.005/min | $0.018/min | No | 97+ |
| Gemini 3.8 Live Extended Thinking | $0.005/min | $0.018/min + reasoning tokens | Yes (background) | 97+ |
| GPT-6 Astra (voice layer) | ~$0.05/min | Bundled | Backend (billed separately) | 50+ |
At $0.005/min audio input versus ~$0.05/min for GPT-6 Astra's voice layer, Gemini 3.8 Live is roughly 10× cheaper on audio input.
Put that in concrete terms. A voice agent handling 1,000 minutes of conversation per day costs about $7.20/day with Gemini 3.8 Live (input plus output combined) versus roughly $50+/day with GPT-6 Astra. Over a month, that's about $216 versus $1,500+ — you're saving close to $1,300/month on that single workload. Scale to 10,000 minutes a day and the gap widens to well over $13,000/month. If you're running a call center assistant or a voice IVR at any meaningful volume, that difference lands in your budget the first billing cycle.
Choosing Between the Two Models
The two models serve different use cases, even within the Gemini 3.8 Live family. The base model — no extended thinking — is built for speed and throughput. Think IVR replacements, real-time customer service bots, appointment schedulers, or any scenario where the agent mostly listens, routes, and responds without reasoning through complex multi-step logic. The latency profile is tuned for snappy back-and-forth, not deep deliberation.
Extended Thinking earns its keep when the conversation itself requires judgment mid-stream: troubleshooting flows, insurance claims triage, multi-condition eligibility checks. You pay for reasoning tokens on top of the audio rate, so it's not free — but for a workflow that would otherwise need a human, the math still favors the model. Just budget for those extra tokens; on reasoning-heavy calls they can double your per-minute cost.
Where It Falls Short
Two things to weigh before you commit. First, language coverage: 97+ languages sounds comprehensive, but quality varies sharply outside the top 20 or so. Test your specific target languages before assuming parity. Second, the reasoning-token pricing on Extended Thinking is variable, which makes cost forecasting harder than the clean $0.005/min headline suggests. If your traffic is spiky or reasoning-intensive, model your worst case, not your average.
Bottom Line
If you're starting a new voice project today and cost is a real constraint — and for most teams it is — Gemini 3.8 Live is the default choice to beat. A 10× reduction on audio input isn't a rounding error; it's the difference between a voice feature being viable and being shelved. I'd start with the base model for anything conversational-but-simple, and only reach for Extended Thinking once you've confirmed a workflow genuinely needs mid-call reasoning.
The one caveat: don't migrate an existing production system on price alone. Run your own latency and accuracy tests against your real traffic, especially if you're in a regulated domain where that 35.1% banking score should give you pause. But for greenfield builds, the burden of proof has shifted — you now need a specific reason not to use Gemini here, and "we're already on Astra" isn't a good enough one.
Comments
Post a Comment