ORF-N-2026-020·Dispatch

Claude Opus 5: the frontier at half the price

Claim

Claude Opus 5 puts Fable-5-class intelligence on the tier anyone can buy, at half Fable 5's cost per task and at Opus 4.8's unchanged sticker price. The sharper reading is where the gains landed: not on coding benchmarks but on agentic business workflows and computer use, the shape of the work an operating business actually automates. And the safety door stopped refusing. A flagged request now falls back to a weaker model instead of failing, which is a better product and a new thing an operator has to log.

July 24, 2026 · 9 min · dispatch · co-authored by Claude Opus 5

Today, 24 July 2026, Anthropic shipped Claude Opus 5: available on every platform at once, the new default on Claude Max and the strongest model on Claude Pro, and priced at $5 per million input tokens and $25 per million output, which is exactly what Opus 4.8 cost yesterday. Anthropic’s own framing is the release in a sentence: a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.

For six weeks the frontier’s story has been about doors. Fable 5 shipped as one model behind two on 9 June, a government order closed both, they reopened on 1 July, and eight days later GPT-5.6 arrived through a preview staged for the US government. Opus 5 runs the other way, and it does it on price. Fable 5’s open door costs $10 in and $50 out per million tokens. Opus 5 asks half of that for work Anthropic places close to it, on a tier whose sticker has not moved since May.

These dispatches read a release through the operator lens: what changed for a business that runs agents on real work with real money. By that lens there are three things in this box. A price move, a shift in where the capability landed, and a quiet change to what happens when the safety layer says no.

The curve, not the score

A release table invites you to read down the new model’s column. The more useful chart plots the score against what it cost to get there, because that curve is what an operating business actually buys.

FRONTIER-BENCH v0.1 · AGENTIC CODING, SCORE AGAINST COST PER ATTEMPT 0 10 20 30 40 50 $1 $2 $5 $10 $20 $30 cost per attempt (log $) GPT-5.6 Sol Opus 4.8 Fable 5 Opus 5 · peak at max effort each curve sweeps effort low to max; Opus 5 clears Fable 5's best score at roughly half the cost per attempt
Figure 1. Agentic coding score against cost per attempt on Frontier-Bench v0.1, each curve sweeping reasoning effort from low to max. Opus 5's band sits above Fable 5's across the range, peaks higher than Fable 5's best, and gets there at roughly half the cost per attempt. Opus 4.8 trails both, and GPT-5.6 Sol runs the cheap end of the chart to a peak between them.

The chart carries the argument. On Frontier-Bench, Anthropic’s claim is that Opus 5 more than doubles Opus 4.8 at a lower cost per task; read the plotted curves and it also clears Fable 5’s best score at around half the spend per attempt. The pattern repeats on a harness other people built: on CursorBench 3.2 at max effort, Opus 5 lands within 0.5% of Fable 5’s peak at half the cost per task, and beats every other model’s performance per dollar at high, xhigh, and max effort. Anthropic charts the Artificial Analysis Coding Agent Index alongside those two, where the same ordering holds. Fable 5 still sets the top of the frontier. What changed is that standing next to it no longer costs what it cost in June.

The widest gap in the release is not on coding at all.

ARC-AGI-3 · NOVEL PROBLEM SOLVING, SCORE AGAINST EVALUATION COST 0 10 20 30 $10,000 $15,000 $20,000 $25,000 total cost of the evaluation run (log $) GPT-5.6 Sol Opus 4.8 (high) Opus 5 (high) Anthropic: three times the next-best model the x axis is what the whole evaluation run cost, not the price of one task; high effort unless noted
Figure 2. ARC-AGI-3, novel problem solving, plotted against the cost of the whole evaluation run rather than the price of one task. Opus 5 at high effort sits alone near the top of the chart while every other model stays under the ten-point line. Anthropic reports its score as three times the next-best model's.

ARC-AGI-3 measures problems the model cannot have seen, and until now every frontier model has been close to the floor on it. Read the axis carefully before quoting the number: it is the cost of the whole evaluation run, in the tens of thousands of dollars, not what a task costs you. This is a capability datapoint, not a budget one. It says the new model is doing something the previous generation was not doing at any price.

The gains landed where the work is

For a lab, the headline is coding. For an operating business, the interesting rows are the ones that look like an actual operation: a workflow that spans several systems, a screen that has to be driven, a document that has to come out right.

ZAPIER AUTOMATIONBENCH · AGENTIC BUSINESS WORKFLOWS, PASS RATE AGAINST COST 0 10 20 30 $0.20 $0.50 $1.00 $2.00 $3.00 cost per task (log $) Opus 5 Fable 5 Opus 4.8 GPT-5.6 Sol Opus 5 at its lowest setting, above every other model's best the workflow benchmark, not a coding one: the whole Opus 5 band sits above the field at lower cost per task
Figure 3. Zapier's AutomationBench, pass rate against cost per task. The whole Opus 5 band sits above every other model's, and the dashed line marks its lowest effort setting, which still passes more tasks than any other model's best result at any price.

AutomationBench is Zapier’s, built from the multi-step business workflows its customers automate, and it is the closest public proxy we have to the work we are asked to build. Opus 5’s entire band sits above the field, at around 1.5 times the next-best model’s pass rate for the same money, and its cheapest setting outperforms everyone else’s most expensive one. The same shape holds on OSWorld 2.0, the computer-use evaluation, where Anthropic reports Opus 5 beating every other model at any cost and passing Fable 5’s best result at just over a third of the price. On the knowledge-work evaluations, GDPval-AA v2, Humanity’s Last Exam and DeepSearchQA, Anthropic puts it first and cheapest as well.

Two things follow. The first is that the release’s largest margins are on the tasks least like a coding benchmark and most like a business process, which is a different claim from “the frontier got smarter” and a more useful one. The second is arithmetic. Since we wrote about the floor rising on 30 June, the price of running an agent through a real workflow at a useful pass rate has fallen again, this time at the top of the line rather than the middle. Any unit economics you set before this month are stale, and this is the third month running that has been true.

The door swings more often than it stops you

The safety architecture we have been tracking all quarter shows up here in a form worth naming precisely, because it changes an operator’s day.

OSS-FUZZ · FINDING VULNERABILITIES AGAINST DEVELOPING EXPLOITS vulnerability identification (pass@1, % of challenges) 0 50 100 61.5 80.0 79.4 Opus 4.8 Mythos 5 Opus 5 exploitation success (challenges solved outright) 0 7 14 0 13 4 Opus 4.8 Mythos 5 Opus 5 Opus 5 finds what the gated Mythos 5 finds, and turns far fewer of those findings into working exploits
Figure 4. OSS-Fuzz, one of Anthropic's cybersecurity evaluations. On the left, Opus 5 identifies vulnerabilities about as well as the gated Mythos 5, 79.4% against 80.0%. On the right, it turns far fewer of those findings into working exploits, 4 challenges against 13. The gap between the panels is the shape of the safeguard.

Opus 5 ships behind cyber classifiers that let it find vulnerabilities in source code while blocking binary-based scanning, penetration testing, and exploit generation. Fallback rather than refusal is not the new part: we described that mechanism in June, when Fable 5’s classifiers fired in roughly 5% of sessions and handed the request to Opus 4.8 unbilled. What changed is the calibration and the reach. Anthropic expects Opus 5’s classifiers to intervene around 85% less often than Fable 5’s, and the fallback to Opus 4.8 now covers Claude.ai, Claude Code and Claude Cowork by default. On the API the same posture is available through the fallbacks parameter, which gains a "default" mode that applies Anthropic’s recommended fallback model per refusal category instead of a list you maintain, in beta. Around it the access tiers hold their shape: the standing Cyber Verification Program gives eligible enterprises and researchers a version of Opus 5 with fewer restrictions, and requests that Fable 5 blocks on biology now route to Opus 5 instead.

For an operator this is a genuine product improvement and a widening obligation. A refusal you could see is a model substitution you cannot, it now reaches across three surfaces and an API default rather than one model’s classifiers, and a silent substitution changes three things you are accountable for at once: what the answer cost, how good it was, and which model produced the artifact sitting in your audit trail. If you run agents in a regulated process, the response metadata that names the answering model stops being telemetry and starts being a record you keep.

The discipline it asks for

Opus 4.8 asked for a discipline when it shipped, and Opus 5 asks for a slightly different one. Thinking is now on by default, so requests that ran without it will think, and because max_tokens caps thinking plus response text together, budgets carried over from Opus 4.8 need revisiting. The effort ladder runs low through max with no beta header, and disabling thinking is only accepted at high effort or below: send thinking: {"type": "disabled"} with xhigh or max and you get a 400. The minimum cacheable prompt drops to 512 tokens, generally available. Two more arrive gated: tool lists can change mid-conversation without discarding the cache, in beta behind a header, and Fast mode, a research preview on the Claude API only, runs about 2.5 times quicker at twice the price, $10 and $50 per million.

The behavior changes matter more than the parameters. Opus 5 writes longer by default, narrates its progress more often in agentic sessions, delegates to subagents more readily, and verifies its own work without being asked. Anthropic’s guidance is to delete the verification instructions you carried from earlier models, because on this model they cause over-verification: a step you built into your harness is now built into the model, and leaving both in place costs tokens and time.

That is the harness thesis playing out on schedule. The half you own gets thinner wherever the model absorbs a piece of scaffolding, and it does not get thinner anywhere else. Permissions, review gates, the definition of done, the record of which model answered, all of that is still yours, and the last of those just got more important. A harness built to stop getting in the model’s way is one you can strip a verification loop out of in an afternoon. A harness that ossified around what last year’s model could not do is one you now pay to run twice.

Where this honestly stands

The discipline of a dispatch is to mark the edges, and this release has several.

Nearly every number above is Anthropic’s harness measuring Anthropic’s models, and the headline coding chart is an internal run in which Opus 4.8 served as the fallback whenever a safety classifier refused, for both Opus 5 and Fable 5. That is a sound way to keep a benchmark running and it means the curve describes a system, not purely a model. The independent yardsticks in the release, CursorBench and the Artificial Analysis Coding Agent Index, are charted by Anthropic rather than published by their owners; no outside run of Opus 5 had appeared at the time of writing, and on past form the independent read agrees on direction and disagrees on degree.

AUTOMATED BEHAVIORAL AUDIT · MISALIGNED BEHAVIOR, 1–10, LOWER IS BETTER Opus 4.82.85 Mythos 52.81 Sonnet 53.35 Opus 52.30 lower is better; the audit is Anthropic's own, and it re-scores the field each release
Figure 5. Anthropic's automated behavioral audit, 1 to 10, lower is better. Opus 5 scores 2.30, below Opus 4.8, Mythos 5 and Sonnet 5 on the same chart. Read within the chart and not across releases: the audit re-scores the whole field each time, and the same models scored differently a month ago.

The alignment result is real and it needs one caveat of its own. Opus 5 is the cleanest model on the chart Anthropic published today, and we quoted the same audit in June when it scored Sonnet 5 at 2.53 and Opus 4.8 at 2.10, against 3.35 and 2.85 today. Mythos 5 moved further, from 2.06 at the Fable 5 release to 2.81 now. The audit is re-run against a moving instrument. Compare bars inside one chart, not across two releases.

On capability, Anthropic is explicit that Opus 5 does not advance the frontier in risky dual-use work and stays behind Mythos 5 in both biology research and offensive cybersecurity, which the OSS-Fuzz panels above show plainly on the cyber half. Its gains in the sciences are real but bounded: 10.2 points over Opus 4.8 on organic chemistry, 7.7 on protein prediction, with the same stated limitations on long-running autonomous research. And the model that beat everything on ARC-AGI-3 did so on an evaluation costing tens of thousands of dollars to run, which is worth remembering before that number appears in a business case.

What to do with this

Three things, in order of how quickly they pay.

Reprice. The work an agent does in a real business workflow, not a coding benchmark, got roughly 1.5 times better per dollar this week, at a tier whose sticker did not change. If a project was scoped as too expensive to automate under the May or June numbers, it is worth re-running the arithmetic, and worth doing so now rather than after the next release, because that arithmetic has moved three months running.

Clean the harness. Drop the verification instructions, revisit max_tokens where thinking is now on by default, walk your effort settings down before you assume you need max, and check that nothing in your stack disables thinking above high effort, since that request will now fail outright.

Log what answers. Silent fallback to Opus 4.8 is a good default and a bad surprise. Record the model that produced each response, alert on the fallback path, and decide deliberately whether a given process wants the fallback or wants to fail loudly. If the way your business runs agents is now a question about routing, verification and record-keeping rather than raw capability, that is the conversation to have. Start it with us as a Discovery Phase.

References

  1. Anthropic. Claude Opus 5. 24 Jul 2026. anthropic.com/news/claude-opus-5
  2. Anthropic. What’s new in Claude Opus 5. Claude Docs. accessed 4 Aug 2026. platform.claude.com/docs