BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Models/Claude Opus 5: Half the Price, and a New Effort Tr…
ModelsJuly 27, 2026
Read · 5 min
claude opus 5 · anthropic

Claude Opus 5: Half the Price, and a New Effort Tradeoff

Anthropic's Claude Opus 5 tops the intelligence index and halves token pricing, yet its highest effort tier often produces messier code than the middle one.

Turn the reasoning effort up and the model gets better. That assumption held for two years of frontier releases, and Claude Opus 5 just broke it in public. Across at least three independent evaluations of Anthropic's new flagship, released July 25, 2026, the highest effort setting produced worse results than the middle one. The model thought harder, refactored more, and shipped more errors. That anomaly is the most useful thing in the launch data, and it is buried under a pile of leaderboard numbers.

The headline facts are simpler. Claude Opus 5 sits at the top of the Artificial Analysis Intelligence Index, costs half of what Claude Fable 5 costs per token, and posted a score on a novel reasoning benchmark that is roughly four times the previous record. Then the caveats begin.

Key takeaways
  • Claude Opus 5 prices at $5 per million input tokens and $25 per million output, half of Fable 5's rates, with a 1 million token context window.
  • It leads the Artificial Analysis Intelligence Index at 61 points against Fable 5 at 60.
  • It scored 30.2 percent on ARC-AGI-3 against 7.8 percent for GPT-5.6 Sol, and ARC Prize reports behavior it had not seen before.
  • It hallucinates on 50 percent of AA-Omniscience items because it answers more often when uncertain.
  • Multiple evaluators found the maximum effort tier underperforms the high or medium tier on coding.

What Anthropic actually shipped

Per The Decoder's launch coverage, Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens. Claude Fable 5, which went public on July 1, 2026, sits at $10 and $50 for the same units. The context window is 1 million tokens. The model became the default on Claude Max and the most capable option available on Claude Pro.

Halving the price of a frontier tier is the structural news. For most production workloads the deciding number is not peak capability but cost per completed task, and a 50 percent token discount changes which jobs are economical to automate at all. A workflow that costs $4,000 a month in tokens becomes a $2,000 workflow without a single line of code changing.

One caution belongs next to that number. Published token rates describe price, not spend. Earlier Opus releases consumed more tokens per task even at flat rates, so the effective cost rose while the price list stayed still. The Decoder raises exactly this point, and it is the first thing any team should measure on its own traffic rather than take from a launch post. The method for doing that, counting a request category by category instead of multiplying a rate, sits in our guide to what a real feature costs once retries and caching are included.

Where Claude Opus 5 leads

The Artificial Analysis results put Opus 5 at 61 on the Intelligence Index, one point ahead of Fable 5 at 60. A single point is inside the noise of any composite index, so the ranking matters less than the fact that the cheaper model reached the same tier.

The coding numbers are more separated. Opus 5 shares first place on the Coding Agent Index at 67 points when paired with Claude Code, and reached 89 percent on Terminal-Bench v2.1 at its top performance level. On Frontier-Bench agentic coding it scored 43.3 percent against 33.7 percent for Fable 5, a gap wide enough to survive normal benchmark variance.

Knowledge work shows the same pattern. On GDPval-AA v2 the model posted an Elo of 1,861 against Fable 5's 1,747. On AA-Briefcase, Artificial Analysis reported Opus 5 taking the lead while cutting cost per task by about 20 percent, with the high reasoning tier completing tasks at $10.41 against $22.30 for Fable 5.

Those two figures together are the real story of the release. Better output, less than half the cost per task. That combination compounds faster than a benchmark point.

The ARC-AGI-3 jump, and how much to read into it

The number that generated the loudest reaction came from ARC-AGI-3, which tests whether a system can infer the rules of an interactive environment it has never seen, plan inside it, and execute step by step. It is deliberately built to resist memorization.

Opus 5 scored 30.2 percent. GPT-5.6 Sol at maximum settings had held the record at 7.8 percent, per The Decoder's report on the ARC Prize results. On the older versions of the benchmark the model reached 90.4 percent on ARC-AGI-2 and 97.5 percent on ARC-AGI-1.

The qualitative observation is more interesting than the score. ARC Prize reported that the model translated tasks into algebraic notation and derived reflection equations on its own, behavior they had not recorded from any previous system. It solved five environments that no model had solved, four of them at human level.

"One benchmark cannot settle whether a model generalizes better."The caution attached to the ARC-AGI-3 result

ARC Prize's Greg Kamradt attached his own caveat, acknowledging that a single benchmark cannot establish broad reasoning gains even while allowing that the gains may be real. Independent testing on the Witness benchmark from Guanghan Ning showed a much narrower improvement, with Kimi K3 and Fable 5 tied at 43.4. When one evaluation shows a fourfold jump and another shows a modest one, the honest read is that ARC-AGI-3 measures something Opus 5 is unusually good at, not that general reasoning quadrupled in three weeks.

Where it trails

Opus 5 does not lead everywhere, and the exceptions are informative. On DeepSWE v1.1 it scored 68.8 percent while GPT-5.6 Sol led at 72.7 percent. On physics benchmarks it sits behind both GPT-5.6 Sol and GPT-5.5 Pro.

Epoch's Capability Index tells a similar story with different framing. Latent Space's launch roundup notes an ECI of 159 for Opus 5 against 161 for Fable 5, with the two tied at 161 on the software engineering subscore. Slightly behind overall, level on code, at half the price.

The pattern across every source is consistent. This is not a model that dominates the frontier. It is a model that reaches the frontier at a materially lower cost, with a distinctive spike on novel interactive reasoning and a soft spot in the sciences.

Four leaderboards, four different answers

This launch is a useful case study in why benchmark coverage feels contradictory. Four respected evaluations reached four different conclusions about the same model in the same week, and none of them is wrong.

The Artificial Analysis Intelligence Index put Opus 5 first at 61. It is a composite, blending many task categories into one number, so it rewards breadth and smooths over spikes. A one point lead in a composite means the two models are effectively tied, and anyone reporting it as a decisive win is overreading the instrument.

Epoch's Capability Index put Opus 5 second at 159 against 161. It weights differently, which is the entire explanation for the flipped ordering. When two composites disagree by a couple of points in opposite directions, the correct conclusion is parity, not contradiction.

ARC-AGI-3 measured one narrow thing and found a large gap. The Witness benchmark measured a related thing and found a small one. Narrow benchmarks have high variance by construction, because a single capability that a model happens to have can move the whole score. That is what makes them informative and also what makes them unreliable as summaries.

The reading that survives all four: Opus 5 is at the frontier without dominating it, is genuinely strong at agentic coding, has a real and unusual spike on interactive rule inference, and is cheaper than the model it sits beside. Every one of those claims holds across sources. Anything stronger depends on which leaderboard you picked first.

What a million token context actually changes

The 1 million token context window is easy to skim past because large context numbers have become a routine line in launch posts. Its practical effect is on architecture rather than on capability.

At that size, an entire mid sized repository, its dependency manifests and its recent commit history fit in a single request. Retrieval stops being mandatory for many code tasks, which removes a whole layer of chunking, embedding and relevance tuning that teams currently maintain to work around smaller windows. Removing infrastructure is usually worth more than adding capability, because infrastructure has ongoing failure modes.

The tradeoff is cost and attention. A million tokens of input at $5 per million is $5 for a single call, and the model still has to locate the relevant few hundred tokens inside that mass. Long context does not eliminate the retrieval problem so much as move it inside the model, where you cannot inspect or tune it. For work where precision matters more than convenience, a well built retrieval layer over a smaller window remains easier to debug.

The sensible pattern is to use the large window for exploration, when you do not yet know which files matter, and to narrow the context once you do. That also keeps the token bill closer to the advertised discount.

Availability and what changes for existing plans

Opus 5 arrives as the default model on Claude Max and as the most capable option on Claude Pro, which matters because plan level access has been a moving target across recent Anthropic releases. We tracked an earlier round of that in our piece on how Fable 5 access was structured across plans.

For API users the substantive change is the price sheet rather than the availability. A team already running Opus workloads sees the same interface at half the token rate, which makes the migration decision unusually simple: the only open question is token consumption per task, and that is measurable in an afternoon.

Measuring that is more concrete than it sounds. Replay a representative sample of last month's requests against both models, record total output tokens rather than wall clock time, and compare spend per successfully completed task rather than per call. Failed runs that get retried are where hidden cost accumulates, so a model that succeeds slightly more often can be cheaper even at an identical token rate. Teams that skip this step tend to discover the answer later, in a billing statement rather than a spreadsheet.

The hallucination number nobody should skip

Buried in the Artificial Analysis writeup is a figure that should shape how anyone deploys this model: a hallucination rate of 50 percent on AA-Omniscience. The stated cause is not that the model knows less. It is that Opus 5 answers more often when it is uncertain.

That is a calibration choice, not a knowledge gap, and it interacts badly with agentic use. A model that declines to answer produces a visible gap a human can fill. A model that answers confidently produces a plausible string that flows into the next tool call, the next commit, the next customer facing summary. In a single turn chat the user catches it. In a twenty step agent run, nobody does.

The practical response is unglamorous. Ground factual work in retrieval rather than parametric recall, require citations for claims that leave the system, and keep a verification step that does not use the same model that produced the answer. None of that is new advice. The 50 percent figure just raises the cost of ignoring it.

The distinction between calibration and knowledge is worth holding onto, because the two call for opposite fixes. A knowledge gap is closed by giving the model better sources. A calibration gap is closed by changing what the model does when its sources are thin, which is a prompting and scaffolding decision rather than a data one. Asking explicitly for an abstention when confidence is low, and treating an empty answer as a valid result rather than a failure, recovers most of what the calibration shift costs.

More effort is not always better

Here is the finding that deserves its own section, because three independent sources reported versions of it.

Vals.ai found that the highest reasoning tiers produce more complex solutions that contain errors more often, and identified the high tier rather than the maximum tier as the best balance for coding. The Decoder's coverage describes higher effort settings yielding worse outcomes because the model refactors code that did not need refactoring. Latent Space flagged a FrontierCode pattern where medium effort beat higher effort, calling it unusual precisely because more compute normally helps.

The mechanism is intuitive once stated. Extra reasoning budget has to go somewhere. On a task with a short correct answer, additional deliberation does not find a better answer, so it finds a bigger one. The model restructures working code, generalizes a function nobody asked to generalize, and introduces surface area where there was none. Effort becomes ambition.

This inverts the default assumption in most agent frameworks, which treat maximum reasoning as the safe setting for hard problems and dial down only to save money. On Opus 5 the maximum tier appears to be a specialized tool rather than a strictly better one.

Choosing an effort tier on purpose

Treating the effort dial as a real parameter rather than a budget knob leads to a straightforward policy.

  1. Bounded edits. Bug fixes, small refactors and single file changes belong at medium. The correct answer is small, so a larger budget mostly buys scope creep.
  2. Multi step agentic work. Planning across many files and tool calls is where the high tier earned its results on the coding indices.
  3. Novel reasoning. Puzzle like problems with unfamiliar rules are the case ARC-AGI-3 measured, and the one where the top tier has a real argument.
  4. Factual retrieval. Effort does not fix calibration. Given the AA-Omniscience result, spend the budget on grounding rather than on thinking.

Measure this on your own workload before trusting any of it. Benchmark tasks have a shape, and yours may not share it. The cheap experiment is to run the same evaluation set at two adjacent tiers and compare error rate rather than output quality, since the failure mode being described here is confident overreach rather than visible incompetence.

What the price move signals about the frontier

Anthropic released a model that approaches its own premium tier at half the token cost, three weeks after that premium tier shipped. Read that as a statement about where competitive pressure now sits. Capability parity at the top has become common enough that price and cost per task are the differentiators, which is the same dynamic reshaping the open weight side of the market, as we covered in our look at how open weight models compare with frontier systems.

The competitive field around this release is crowded. Fable 5, GPT-5.6 Sol, Grok 4.5 and Kimi K3 all occupy overlapping capability bands, with coding ability serving as the wedge each is pushing hardest. Anthropic's answer is that you should not have to pay frontier prices for frontier coding, and the independent numbers largely support the claim on cost per task even where they complicate it on raw capability.

For anyone deciding what to run in production, the launch reduces to three measurable questions. Does the token discount survive contact with your actual token consumption? Does the calibration change increase your review burden? And is your default effort tier the one that actually performs best on your tasks? Every one of those has a numeric answer available from a weekend of evaluation, which is a better use of time than reading another leaderboard.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building