On August 13, 2026, Google released Gemini 3.7 Flash at half the price of Gemini 3.6 Flash, which had shipped just three weeks earlier. For applications tied to a fixed model, that creates a practical problem: a cost decision can become outdated before the next billing review.
The listed input price is $0.75 per million tokens through December 31, with a rate of $1.50 per million in 2027. Google also reports capability gains. FrontierCode scores increased from 34.4% to 43.6%, while DeepSWE, a benchmark for long-horizon software engineering tasks, rose from 49.0% to 65.3%.
At 340 output tokens per second, Artificial Analysis ranks the model first for speed among 186 models tested. Google's own coding measurements put it ahead of Claude Sonnet 5 and GPT-5.6 Terra. Those are useful signals for model selection, though Google's comparisons remain vendor measurements.
The infrastructure question is how to take advantage of changes like this without rewriting application code or sending difficult work to a model that can't handle it reliably.
The price spread is getting harder to ignore
Comparisons suggesting AI API price declines of 90% to 97% since GPT-4 cover a market with very different capability tiers and billing rates. GPT-4 launched at around $60 per million input tokens. GPT-5.5, a current top-tier option, runs $30 per million. Mid-tier models that benchmark above GPT-4o cost under $2. At the low-cost end, DeepSeek V4 Flash supports workloads with a 1-million-token context at $0.28 per million output tokens.
Input and output rates aren't interchangeable, and a token price alone doesn't establish the cost of completing a task. Still, the workload comparison is striking: a task priced at $788 on GPT-5.5 is reported to cost $11 on DeepSeek V4 Flash. The comparison describes the outputs as functionally equivalent. Its usefulness for a particular application depends on whether that equivalence holds for the work being done.
Several price cuts arrived within the six weeks leading up to August 16. OpenAI cut GPT-5.6 Luna prices by 80% on July 30, three weeks after launch. Anthropic announced a 67% cut to Claude pricing. Google then halved the price of its Flash workhorse model.
A production deployment whose cost assumptions haven't been reviewed in 90 days may now have substantially cheaper options. Whether switching pays off depends on task quality as well as the advertised rate.
Cheap and premium models are moving in different directions
DeepSeek has taken a different approach at the premium end. While other providers cut rates, DeepSeek raised V4 Pro prices by up to 1,100% on premium workloads. Its V4 Flash model remains at $0.28 per million output tokens.
That looks like deliberate market segmentation. The apparent bet is that cost-sensitive work will move toward inexpensive models, while quality-sensitive work will support premium prices and higher margins. Whether that can sustain the business remains an open question.
There is a loose parallel with AWS Reserved Instances and On-Demand pricing: different buyers accept different tradeoffs to control costs. The distinction here is model capability, rather than a purchasing commitment. A single application can have requests that belong at both ends of the price range.
That makes request routing an operational decision. Choosing a provider once doesn't settle which model should handle every task.
What production routing looks like
On July 22, 2026, Cursor launched Cursor Router, a classifier trained on more than 600,000 live coding requests. It evaluates the query, surrounding code context, task complexity, and domain, then sends the request to a suitable model.
The router offers three modes:
- Intelligence: prioritize model capability.
- Balance: aim for strong quality at lower cost.
- Cost: minimize spending while maintaining reasonable quality.
In A/B testing across millions of production requests, Cursor Router reportedly reduced costs by 60% compared with sending everything to Opus 4.8. Early enterprise deployments reported reductions of 30% to 50% while maintaining comparable frontier-model performance.
Those results support task-aware routing as a production pattern. They don't establish the same savings for every application, but they show that choosing among models can materially affect costs at scale.
The networking analogy is useful. DNS-based load balancing distributes traffic through name resolution. Anycast uses network routing to reach a service location. Application-layer routing can make more detailed decisions using request content and session state. Model routing similarly needs to understand something about the request, rather than merely choose an available endpoint.
As model prices fall and capabilities remain uneven, the routing layer becomes a place to manage both. It can keep routine work inexpensive while reserving more capable models for tasks that justify their cost.
Three approaches that can waste money
Sending everything to one model. A single provider was easier to justify when capability gaps left few workable substitutes. That case is weaker for workloads where several models perform well. Google's coding results suggest Gemini 3.7 Flash may be a lower-cost alternative to Claude Sonnet 5 for some tasks at $0.75 per million input tokens. The same decision needs task-specific evidence for document work or other workloads; coding scores alone don't settle it.
Reviewing costs only once a quarter. Price changes are arriving within weeks of model launches. Automated monitoring of model rates and actual spending can expose changes sooner than a quarterly review. Infrastructure teams already track changing egress and compute economics. AI API rates need similar attention if the application can switch models when a better option appears.
Treating all requests as equally difficult. Summarizing a support ticket, generating a pull request description, and debugging a segmentation fault in unfamiliar code place different demands on a model. Routing all three to the most expensive option can waste money. Routing all three to the cheapest can sacrifice useful results. A reported $788 versus $11 spread makes classification worth examining, provided the lower-cost output is good enough for the task.
Separate the application from model selection
An application with claude-sonnet-5 or gpt-5.6-luna scattered across service calls is expensive to adjust. A price change shouldn't require edits at fifty call sites.
A practical starting point is a configuration-driven dispatch table that maps task types to model endpoints. The application identifies the kind of work it needs done, and the dispatch layer selects the model. That centralizes changes to provider, model, and routing policy.
The next requirement is a feedback loop. Logging cost per task alongside task outcomes gives the team a way to measure whether a routing change saves money without reducing useful performance. A lower token bill is a poor result if the task no longer gets completed adequately.
More sophisticated systems can use a trained classifier, as Cursor does, to make decisions for individual queries. That adds complexity, but it can distinguish requests that share a broad task label while requiring different levels of capability.
Multi-provider API layers such as OpenRouter are a reasonable starting point for provider abstraction. Access to several models doesn't, by itself, establish which model belongs on which task. Task-aware selection is another layer. The reported 30% to 60% savings come from routing deployments, not simply from placing multiple providers behind one API.
More capacity could keep prices under pressure
Big Tech has accumulated nearly $1.5 trillion in purchase commitments for chips, data centers, and energy infrastructure. As that capacity comes online, it could lower the marginal cost of serving tokens and support further API price competition.
The expectation is that this supply expansion will keep competitive pressure on prices for at least the next two to three years. Purchase commitments alone don't guarantee lower prices, particularly for premium models, but they support the case for avoiding a rigid model dependency.
The load-balancer analogy has a practical limit: an AI router must account for output quality, not just availability and traffic distribution. Its job is to send each task to a model that can complete it at an acceptable cost. Centralized configuration and measured task outcomes provide a workable foundation before a trained routing system is necessary.