the wire · #ai · 2026-09-02
Google says its new Gemini 3.8 Flash model ‘works harder’ but might cost more
Cech Tech Reviews

Google just released Gemini 3.8 Flash only weeks after version 3.7, and the upgrade comes with an interesting tradeoff, according to The Verge. The new model doesn't just answer faster, it thinks harder, performing more reasoning steps and calling tools multiple times to tackle complex requests. That extra effort might deliver better results, but it also means burning through more tokens per query.
The base pricing stays identical to 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens. But Google is warning developers upfront that 3.8 could rack up higher bills because it uses more tokens when tackling difficult problems, especially at higher effort settings. It's a calculated bet: pay for quality when you need it, or stick with 3.7 if you're watching costs closely.
This reflects a broader shift in how AI companies are positioning their models. Instead of just racing toward cheaper and faster, we're seeing differentiation around reasoning depth. OpenAI's o1 model similarly prioritizes extended thinking over speed. Google is essentially offering two tools for different jobs: 3.7 for straightforward tasks where efficiency matters, and 3.8 when you need the model to work through something genuinely complex.
For developers, this creates a real decision point. If you're running high-volume, repetitive tasks like summarization or simple data extraction, the token overhead of 3.8 probably isn't worth it. But if you're building agents that need to troubleshoot, plan multi-step workflows, or handle ambiguous requests, the extra reasoning could save you from brittle outputs and manual cleanup.
The iterative tool-calling feature is particularly interesting. Rather than making one API call and hoping for the best, 3.8 can refine its approach as it goes, checking results and adjusting. That's closer to how a human would debug or research, and it's exactly what agentic workflows need. The cost is variable compute, which means variable pricing.
What this means for you: if you're prototyping or running cost-sensitive production workloads, start with Gemini 3.7 Flash and only upgrade specific endpoints to 3.8 where quality gaps appear. For a practical test, try this prompt with both models and compare token usage and output quality: "Review this customer support ticket, identify the core issue, check our knowledge base for relevant solutions, then draft a response that addresses all concerns and prevents follow-up questions." Track which model gives you fewer revision cycles, that's your real cost savings.
Reporting basis: original story
← back to The Wire







