Gemini 3.8 Live: Voice Agents Get Real (and Cheaper)
Gemini 3.8 Live holds live voice conversations while doing tasks mid-call. Review per-minute pricing, features, and what it means for service calls.
Gemini 3.8 Flash brings agentic reasoning and coding at 75 cents per million input tokens. Where it runs, what the 3.8 family adds, and how to use it.
Google released Gemini 3.8 Flash on September 2, 2026, as its best reasoning and coding model at the time, designed for agentic workflows where the model plans, uses tools, and carries multi-step work forward. The pricing is the part owners should read twice: 0.75 dollars per million input tokens and 3.75 dollars per million output tokens on introductory rates, rising to 1.50 and 7.50 dollars on January 1, 2027. Low entry cost, a scheduled step up, and a full family of voice models around it.
Agentic workflow is the phrase Google attaches to this model, and it has a concrete meaning: the model is expected to break a goal into steps, call tools, read results, and continue, rather than produce a single answer and stop. Reasoning strength plus coding strength is the combination that makes that loop work, since most business agents spend their time reading structured data, transforming it, and writing outputs back somewhere.
Availability follows Google's surfaces. Gemini Enterprise carries it for organizational use, Google AI Pro and Ultra subscribers get it in the consumer tiers, and AI Mode in Google Search exposes it where customers already ask questions. That last placement matters strategically: if research and comparison work happens inside Search itself, the businesses that publish clear, structured, checkable information benefit when the model synthesizes answers. The model card on DeepMind's site is the primary reference for capabilities and limits, and it is worth a read before any build decision.
For teams comparing everyday models, our Sonnet 5.5 business guide covers Anthropic's counterpart, and the Grok 4.7 explainer covers xAI's long-horizon entry. The three were released within weeks of each other, which says more about the market than about any single model.
The introductory rates are genuinely low for this capability class: 0.75 dollars per million input tokens and 3.75 dollars per million output tokens. The scheduled rates from January 1, 2027, are exactly double: 1.50 and 7.50 dollars. Google published both sets together, which is good practice, and owners should return the favor by budgeting at the higher set from the start.
Why the caution matters: agent workloads consume tokens differently from chat. A single delegated job can read large contexts, iterate across steps, and generate long outputs, so per-task cost depends on design choices such as context size, step limits, and output length caps. A pilot that looks cheap at introductory rates with small contexts can surprise you at scale. Model the pilot at January rates, set step and length limits in the brief, and track cost per completed task alongside quality. If the task economics work at 1.50 and 7.50, the introductory period is a bonus rather than a trap.
| Item | Introductory rate | From January 1, 2027 |
|---|---|---|
| Input tokens | 0.75 dollars per million | 1.50 dollars per million |
| Output tokens | 3.75 dollars per million | 7.50 dollars per million |
| Planning advice | Pilot freely | Budget and cap at these rates |
September brought a full voice stack alongside the reasoning model, and service businesses that live on calls should understand the pieces. Gemini 3.8 Live and Live Extended Thinking, released September 15, 2026, are native speech-to-speech models: they hold a conversation and perform tasks while keeping the dialogue going, rather than transcribing, thinking, and then speaking in separate passes. Pricing is per minute, 0.005 dollars for audio input and 0.018 dollars for audio output, integrating through partners including Agora, LiveKit, Pipecat, and Vercel.
Around that core sit three supporting releases. Live with Live Avatar, from September 24, adds real-time video presence with speech for Gemini Enterprise, meaning a visual persona for live dialogue. The 3.8 Flash and Flash-Lite text-to-speech models, from September 23, provide expressive voice output for agents and content. And 3.5 Transcribe, from August 2026, handles speech-to-text at a 4.0 percent streaming word error rate across more than 85 languages. Together the family covers the full loop: hear the caller, reason about the request, speak the answer, and optionally appear on screen.
Our Gemini 3.8 Live voice guide goes deeper on the phone implications. For the operational side of answering today, the AI voice agent service describes configured call handling and the voice receptionist workflow shows the call flow.
The lowest-risk starting point is text-based agentic work inside tools your team already opens: research briefs, code-adjacent automation, document drafting from large inputs, and multi-step data cleanup. These tasks exercise the reasoning and tool-use strengths with outputs a person can review before anything customer-facing happens. Keep the human review step mandatory during the pilot, set explicit step limits so runaway loops cannot burn budget, and log inputs, actions, and outcomes.
Voice comes second, and only with measured unit economics. Per-minute audio pricing looks small until call volumes multiply it, so compute cost per resolved call, compare it against your current answering cost, and confirm escalation behavior for the calls the agent cannot resolve. The transcription accuracy figure and the 85-plus language coverage are genuine advantages for multilingual customer bases, but they do not replace testing on your actual callers, accents, and background noise. One month of shadow operation, where the agent drafts and a person sends, teaches more than any specification sheet.
A final note on lock-in: building inside Gemini Enterprise, Workspace-adjacent flows, and Search-facing content deepens dependence on one vendor's surfaces. That dependence buys genuine convenience, and for many small teams it is the right trade. Just keep your prompts, evaluation sets, and outcome logs portable, so a future model switch is a reconfiguration rather than a rebuild.
What is Gemini 3.8 Flash?
Gemini 3.8 Flash is Google's reasoning and coding model released September 2, 2026, described at launch as its best in that class and built for agentic workflows. It is available in Gemini Enterprise, to Google AI Pro and Ultra subscribers, and in AI Mode in Google Search. Introductory pricing is 0.75 dollars per million input tokens and 3.75 dollars per million output tokens.
How long does the introductory pricing last?
The 0.75 and 3.75 dollar rates are introductory and rise to 1.50 dollars per million input tokens and 7.50 dollars per million output tokens on January 1, 2027. Owners planning 2027 budgets should model the higher figures now, since agent workloads that process large contexts can scale usage quickly once automations run daily.
What else is in the Gemini 3.8 family?
The September line includes 3.8 Live and Live Extended Thinking for speech-to-speech voice agents, Live with Live Avatar for real-time video presence in Gemini Enterprise, Flash and Flash-Lite text-to-speech for voice output, and 3.5 Transcribe for speech-to-text across more than 85 languages. Together they cover hearing, thinking, speaking, and appearing.
Where should a service business start with Gemini 3.8 Flash?
Start with high-volume agentic chores inside Google's surfaces you already use: drafting and coding assistance, research synthesis, and document work in Gemini Enterprise or AI Pro. Add voice capabilities only after the text workflows prove themselves, and measure cost per resolved task rather than cost per token.
Pick one repeatable research or document task and run it on 3.8 Flash for two weeks with human review on every output. The free six-step AI automation plan structures the pilot: start your AI automation plan.
Gemini 3.8 Live holds live voice conversations while doing tasks mid-call. Review per-minute pricing, features, and what it means for service calls.
Claude Sonnet 5.5 is 30 percent faster and cheaper for everyday business work, with 1M context. Here is where small teams should deploy it first.
Grok 4.7 targets many-hour coding and knowledge tasks with self-checks and long context. Pricing, access paths, and what long-horizon means for owners.
More articles: browse the full Praktivo blog.