Anthropic released Claude Opus 5.5 on the morning of September 22 at $4 per million input tokens and $20 per million output tokens. About 90 minutes later, OpenAI released GPT-6 Sol at $2 and $10 and GPT-6 Luna at $0.10 and $0.50. Both companies cut the price of cached input reads harder than anything else on the sheet. If you run agents against these APIs, that line item is the one to re-check today.
The sequence ran over about 24 hours. xAI shipped Grok 4.7 on September 21 at $2 and $6, the same rate as Grok 4.6. Anthropic followed the next morning. OpenAI followed Anthropic before lunch. Three frontier labs repriced in a day.
What each vendor changed
Opus 5.5 costs 20% less than Opus 5 on list price. The larger cut is underneath. Cache reads dropped from $0.50 to $0.20 per million tokens, a 60% cut. Cache writes dropped from $6.25 to $5. Anthropic says the model performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads, because it finishes the same job in fewer tokens. It has a one million token context window and a 128,000 token output limit. A fast mode runs at $8 and $40 for up to 2.5 times the speed. The model ID is claude-opus-5-5, and it landed on AWS, Google Cloud, and Azure the same day. Subscribers got something too: five hour usage allowances went up 20% and weekly allowances went up 25% on Pro, Max, and Team plans.
OpenAI's cuts are larger on paper. Sol dropped from $4 and $20 for GPT-5.6 Sol to $2 and $10. Luna dropped from $0.20 and $1.20 to $0.10 and $0.50. Cached input reads carry a 90% discount. An OpenAI spokesperson told VentureBeat the pricing is "permanent, not promotional." OpenAI also changed cache behavior. Changing the reasoning effort or the tool list no longer invalidates the cache. That change is one line in the announcement and a large number in production, because tool lists change constantly inside agent loops.
Astra stays at $10 and $50. Fable 5.1 stays at its price. The top tier at both companies did not move. The tier that most production code calls did.
The benchmarks belong to the vendors
Anthropic's numbers: Terminal-Bench 4.0 at 66.4%, against 55.8% for Fable 5.1 and 52.3% for Opus 5. GDPval-AA at 1846 Elo against 1708 for Opus 5. OpenAI's numbers: Sol at 68.8% on DeepSWE 1.1, Luna at 66.6%. On AutomationBench, a test of agents working across 47 business tools, Sol scored 33.2% at 27 cents per task, and OpenAI says Astra on low effort scored 30.3% at roughly 3.9 times the cost per task. Anthropic quotes Opus 5.5 at 40.0% on the same test.
Each lab picked its own tables. Each lab ran its own model at the highest effort setting and the rival at whatever setting made the comparison work. I have read enough of these to stop caring which one leads by four points. The pattern that survives across all of them is plainer. The mid tier now does what the top tier did in July, and the top tier at both companies has become a price anchor.
The agent behavior numbers are more useful than the leaderboard numbers. OpenAI reports Sol's coding deception rate at 1.3%, down from 10.4% for GPT-5.6 Sol, and its failure to disclose rate at 5.4%, down from 77.8%. By OpenAI's own measure, the model many of us have been shipping on failed to disclose what it had done more than three times out of four. Anthropic reports that Opus 5.5 attempted to get around boundaries about 85% less often than Opus 5. Both are self-reported, and both are admissions about the previous model as much as claims about the new one.
Cache reads are most of the bill
A coding agent that makes 40 tool calls in a session sends the system prompt, the tool definitions, the transcript so far, and the newest tool result on every call. Nearly all of that is cached input. Anthropic said outright that cache reads are the majority of the cost of agentic and coding work. So a 20% cut on list price plus a 60% cut on cache reads works out closer to a 50% cut for the workload most of us run. OpenAI's 90% cache discount on Sol puts a cached token at $0.20 per million, the same figure as Anthropic's new Opus rate, from a model that lists at half the price.
Run the numbers on your own traffic before you trust either vendor's percentage. Pull a week of usage logs and split input tokens into cached and uncached. Then apply the new rates. The answer depends on your cache hit rate, and cache hit rate depends on how stable your prompt prefix is. Teams that put a timestamp or a user ID at the top of the system prompt have been paying full price on tokens that could have been cached for two years. This is the week to fix that.
One more thing about Anthropic's model. Some cybersecurity requests to Opus 5.5 get routed to Opus 4.8 as a fallback. The response still comes back. If you do security tooling through the API, test that your prompts land on the model you paid for, and check the model field on every response.
What I'd change in a codebase this week
- Stop hardcoding model IDs. Put them in config with an environment override. Three model generations in three months means the string changes faster than your deploy cycle.
- Log cached and uncached input tokens separately. If your metrics only show total tokens, you cannot evaluate either of these price changes.
- Re-test the routing rules. If you send hard tasks to the top tier and easy ones to a cheap model, the boundary moved on Tuesday. The mid tier may cover most of what you were sending to Astra or Fable.
- Check the effort setting on every call. Sol's headline numbers came from its highest effort mode. Default effort is a different model for practical purposes, and a different bill.
Two dates to put on the calendar. Anthropic said Sonnet 5.5 and Haiku 5.5 arrive "in the coming weeks." Google's Gemini 3.8 Flash sits on introductory pricing of $0.75 and $3.75 through December 31, which is now seven times Luna's list price for the cheap tier. Google will have to move before then.
Dario Amodei has called for "pacing the frontier" to match progress on alignment. TechCrunch notes Opus 5.5 shipped 60 days after Opus 5 on July 24. Whatever pacing means, it does not mean slower releases or higher prices.
My prediction: the $10 and $50 top tier at both companies drops before the end of the year, and the next cut comes from Google within 30 days. Budgets set in January for API spend are now wrong, and they are wrong in the direction that lets you run more.