On 22 September 2026 Anthropic released Claude Opus 5.5, the first model in its Claude 5.5 family, on every platform at once. Input and output tokens cost $4 and $20 per million, 20% less than Opus 5, and cache reads, which Anthropic says make up the majority of agentic and coding costs, fall 60% to $0.20 per million. Anthropic’s framing is the release in a sentence: Opus 5.5 performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.
Two months ago Opus 5 brought the frontier to half the price of Fable 5, and three weeks ago GPT-6 Astra moved the top of the frontier again. Anthropic calls Opus 5.5 the new leading model, and on its own benchmark table it leads agentic coding, computer use and knowledge work. It does not lead everywhere, and Anthropic’s models guidance now starts most workloads on Opus 5.5 while still pointing demanding reasoning and long-horizon agentic work at Fable 5.1. What the release argues hardest is the ratio: how much frontier-level work you get for each dollar you spend on a task.
These dispatches read a release through the operator lens: what changed for a business that runs agents on real work with real money. By that lens this release is one argument, made six times over on Anthropic’s own charts, plus two changes to how you run it.
The token price is the smaller half
The sticker price is the part everyone will quote, and it is the less interesting number. A 20% cut per token is real money, but Anthropic’s headline is a 40% drop in the cost of typical workloads at default settings, and the gap between those two figures is the model using fewer tokens to finish the same work. In Anthropic’s words, the advantage that is very clear is efficiency: it costs less per token than Opus 5 and uses fewer tokens per task.
That distinction matters because intelligence per token is the wrong unit for a business. Opus 5.5 is not the cheapest model per token. Claude Sonnet 5 is half its price and Haiku 4.5 a quarter. A model that costs less per token but takes three times as many steps, or fails and retries, costs more per finished task. The unit that shows up on an invoice is the work, and every chart below plots cost per task or per attempt, never per token.
The early-tester reports Anthropic published point the same way. Optiver says Opus 5.5 matched Opus 5’s quality on its agentic coding tasks in about half the turns, time and output tokens, cutting that workload’s cost by 40 to 50%. Factory calls it the first model it would default to at medium effort, matching Opus 5 on high with 20 to 25% fewer output tokens. Box reports a third of the tokens Opus 5 used. These are customer statements quoted by Anthropic, not independent measurements, and they agree with the direction of the charts.
The curve, not the score
A release table invites you to read down the new model’s column. The more useful chart plots score against what the work cost, each model swept from low to max effort, because that curve is what an operating business actually buys. Anthropic published six of them, and on every one most of the best-value frontier, the line of results nobody beats for less money, is Opus 5.5.
On Terminal-Bench 4.0, multi-step professional work in a command line, Opus 5.5 at its default medium effort scores 57.6% at $2.94 per attempt, the unit Anthropic charts for this benchmark. That beats Opus 5 at max effort for about a fifth of the cost and matches GPT-6 Astra’s best, 57.9%, for about 40% of Astra’s cost. Its own peak, 66.4% at xhigh, is the highest score on the chart.
On FrontierCode v1.1, which asks whether an agent’s code changes would be merged, the default setting scores 54.6% at $0.80 per task, higher than every other model’s best result on the chart. GPT-6 Astra’s top score, 53.3%, costs about five times as much. On CursorBench 4.0, ambiguous multi-file tasks taken from real Cursor sessions, the default setting scores 52.5% at $2.90, above Fable 5.1 at max effort, 51.8%, which the chart plots at $17.28. Against GPT-5.6 Sol’s best it is 11 points higher for about a third of the cost.
The pattern is specific. On all three coding charts, a single default setting of Opus 5.5 already matches or clears every other model’s best result, with Terminal-Bench the closest call at 0.3 points behind GPT-6 Astra. Raising effort buys more score, but the argument for the release is made at the default.
The same curve on knowledge work
For an operating business the coding rows are only part of the picture. The rows that look like an actual operation are knowledge work, business workflows spanning several apps, and data collection at scale.
On GDPval-AA v2.1, Artificial Analysis’s evaluation of professional work across 44 occupations, Opus 5.5 at max effort scores 1846 Elo against 1735 for Fable 5.1 and 1708 for Opus 5. At its default medium setting it beats GPT-6 Astra at max effort for about a fifth of the cost per task. On Perplexity’s WANDR data-collection benchmark, it outperforms Fable 5.1 and Opus 5 at a lower cost: 71.3% at xhigh for $29.06 per attempt, against Fable 5.1’s best of 68.7% at $49.02.
AutomationBench, Zapier’s test of real business workflows across many connected apps, is where the frontier is shared. Opus 5.5 outscores Opus 5 and GPT-5.6 Sol at every effort level, and its best result, 40.0%, costs $1.37 per task. GPT-6 Astra’s best is 41.4% at $1.77. Below Astra’s top setting, Opus 5.5 is cheaper for the same score. At the very top of the workflow chart, Astra still leads. Fable 5.1 is not plotted on the cost chart; the table puts it at 31.4%, below both.
Anthropic’s internal tests describe the same trade in hours rather than percentages. Translating HAProxy from C into Rust, both Opus 5.5 and Fable 5.1 produced rewrites that passed nearly all of HAProxy’s own regression tests; Opus 5.5 finished in 9.5 hours against 12 and cost 51% less. Building a merger analysis and executive presentation, it finished in 63 minutes against Opus 5’s 93 at half the cost. On a quarterly-earnings report graded for invented figures and quotes, 16 of its 18 reports cleared the bar, where neither Fable 5.1 nor Opus 5 cleared it in any attempt.
Where this honestly stands
The discipline of a dispatch is to mark the edges, and this release has several.
The charts are Anthropic’s, and so is most of the harness. Opus 5.5 was evaluated at max effort unless noted, with its production safeguards on: when they intervened, cybersecurity tasks were completed by Opus 4.8 and biology and frontier model development tasks by Opus 5, which Anthropic says likely reduces Opus 5.5’s scores. The GPT-6 Astra and GPT-5.6 Sol figures on Terminal-Bench are as reported by OpenAI. AutomationBench was run by Zapier without fallback models, so every safeguard intervention counted as a failure. The WANDR setup differs from Perplexity’s published one, and Anthropic says the scores are not directly comparable with it. The outside measurements the release cites are Zapier’s early-access AutomationBench run, a Gray Swan prompt-injection benchmark, and pre-release testing by METR and Frontier Design, all as reported by Anthropic.
Anthropic adds a caveat of its own that belongs next to every chart above: at these levels of capability, benchmark margins have become a less reliable guide to real-world differences, and in its own use the gap between Opus 5.5 and Fable 5.1 is narrower than the scores suggest. Read that both ways. It tempers the margins Opus 5.5 posts over Fable 5.1, and the models guidance still sends the hardest reasoning and long-horizon work to Fable 5.1, or to it whenever Opus 5.5 at higher effort still falls short on your own evals.
The frontier is not Opus 5.5’s alone. GPT-6 Astra keeps the top of AutomationBench and leads the table on Terminal-Bench-Science, 64.6% against 58.7%. On CursorBench and AutomationBench, GPT-5.6 Sol at low effort is cheaper still, for a much lower score, and on WANDR Opus 5 at low undercuts Opus 5.5’s medium setting. And “per task” means per task on these benchmarks. Your workload has its own curve, and the only way to know where Opus 5.5 sits on it is to plot it.
On safety, Anthropic reports that Opus 5.5 posts the best scores of any model to date on its automated behavioral audit, and that in a new containment evaluation it tried to cross boundaries about 85% less often than Opus 5 or Mythos 5.1, every attempt low severity and self-reported. It also reports signs that the model often suspects it is being evaluated, a limit on what any pre-release audit can show.
The door, and the discipline
Opus 5.5 is the first Opus model to launch with safeguards in the same class as Fable 5.1’s, on cybersecurity, biology and distillation. Finding and fixing vulnerabilities in your own code stays allowed, but most cybersecurity tasks are re-routed to Opus 4.8, and in Claude apps and Claude Code most flagged messages move to an older model and the work continues there, behind a setting named “Switch models when a message is flagged”. On the API a declined request returns stop_reason: "refusal" with the policy area named, and server-side fallback with fallbacks: "default" is in beta. A new refusal category, reasoning_extraction, covers requests to reproduce the model’s internal reasoning in the reply. Vetted organizations can apply to a new Life Sciences Verification Program for biology work, and the Cyber Verification Program is expanding to Opus 5.5 in the coming weeks.
The API asks for a discipline too, and most of it is breaking. Thinking is always on: thinking: {"type": "disabled"} and manual thinking budgets now return a 400, and effort is the only control for thinking depth. Forcing a tool call with tool_choice set to any or a named tool returns a 400. Thinking blocks are tied to the model and the conversation, so a conversation that moves from Opus 5.5 onto most other models loses the earlier reasoning. On the Claude API and Google Cloud, computer use needs the computer_toolset_20260801 toolset. And the short notes the model writes between tool calls now arrive in thinking blocks that are empty at the default display setting, so an interface that streamed them goes quiet without an error.
The quiet change is the one that touches cost. The default effort is now medium, where Opus 5 defaulted to high, and at any given effort level the model tends to think more per turn than Opus 5 did. A request that omits effort now runs at medium, and a request that pins a setting carried over from Opus 5 may spend more than before. The 40% saving is measured at the defaults. The claude.dev guide makes the same point from the prompt side: delete the “think carefully” lines, because the model already thinks before every reply, and say what done looks like instead.
What to do with this
Three things, in order of how quickly they pay.
Reprice per task. Run your own workload at medium effort before anything higher, and compare cost per finished task, not cost per token. If a project was scoped as too expensive under Opus 5 or Fable 5.1 numbers, the arithmetic has moved again, and on Anthropic’s charts it has moved most at the default.
Recalibrate effort. Set effort explicitly, re-run your sweep instead of carrying Opus 5 settings over, leave room in max_tokens for thinking, and remove anything that disables thinking or forces a tool call before it reaches production as a 400.
Log what answers. Fallback now reaches cyber, biology and reasoning-extraction requests. Record the model that produced each response, alert on the fallback path, and decide deliberately which processes want the fallback and which should fail loudly. If the way your business runs agents is now a question of effort settings, routing and record-keeping rather than raw capability, that is the conversation to have. Start it with us as a Discovery Phase.
References
- Anthropic. Claude Opus 5.5. 22 Sep 2026. anthropic.com/claude-opus-5-5
- Anthropic. What’s new in Claude Opus 5.5. Claude Docs. accessed 25 Sep 2026. platform.claude.com/docs
- Anthropic. Models overview. Claude Docs. accessed 25 Sep 2026. platform.claude.com/docs
- Addy Osmani. Getting the most out of Opus 5.5 in Claude and Claude Code. claude.dev Blog. 22 Sep 2026. claude.dev/blog