• Sat. Aug 1st, 2026

Ravody

Where VPNs, Games, AI & Software Meet Honest Reviews

Claude Opus 5: Anthropic Delivers Near-Frontier Intelligence at Half the Price

ByRavody

Jul 26, 2026

A Fourth Model in Eight Weeks

On July 24, 2026, Anthropic shipped Claude Opus 5, the fourth model to carry the Anthropic name into daily production use since early June. Mythos 5, Fable 5, and Sonnet 5 all arrived within days of one another that month, and Opus 5 lands only two months after Opus 4.8 went live on May 28. Few labs have ever compressed a release calendar this tightly, and the pace itself is now part of the story: Anthropic is treating “ship a materially better model every few weeks” as a strategy rather than an occasional event.

What makes Opus 5 different from a routine mid-cycle refresh is where it lands relative to the rest of Anthropic’s lineup. The company’s own comparisons, published alongside the launch, show Opus 5 beating Fable 5 — Anthropic’s most capable publicly available model — on eight of thirteen internal benchmarks, while costing roughly half as much per million tokens. That is an unusual thing for a mid-tier model to do to its own flagship sibling, and it reframes how buyers should think about the Opus line: not as “the affordable option,” but as the model most teams will actually run every day, with the flagship reserved for the narrower set of tasks that genuinely need it.

  • Release date: July 24, 2026
  • Pricing: $5 per million input tokens / $25 per million output tokens (standard mode) — identical to Opus 4.8’s rate
  • Fast mode: $10/$50 per million tokens, roughly 2.5x faster generation
  • Context window: 1 million tokens, with a 128,000-token maximum output (300K via the Batch API beta)
  • Knowledge cutoff: May 2026 — the freshest of any model in Anthropic’s current lineup
  • API identifier: claude-opus-5, also available as anthropic.claude-opus-5 on Amazon Bedrock and claude-opus-5 on Google Vertex AI

What the Benchmarks Actually Show

Anthropic’s headline claim rests on two evaluation suites: Frontier-Bench v0.1, a 74-task battery spanning physics, chemistry, and cryptography designed to test terminal-based agentic reasoning, and GDPval-AA v2, a knowledge-work benchmark meant to approximate the kind of open-ended analytical tasks that show up in real professional settings.

On Frontier-Bench v0.1, Opus 5 scored 43.3%, compared with 33.7% for Fable 5, 21.1% for the outgoing Opus 4.8, and 34.4% for OpenAI’s GPT-5.6 Sol. That is more than double Opus 4.8’s prior result and a clear lead over every other model Anthropic tested against, including its own flagship. Independent aggregator Artificial Analysis has since folded Opus 5 into its capability boards and now ranks it first on its overall Intelligence Index, first on its Agentic Index, and tied for first on its Coding Index — all achieved, notably, at half the per-token price of the model it just displaced from the top spot.

The margin that matters is not the one-point gap between Opus 5 and GPT-5.6 Sol on raw coding benchmarks — that is close enough to be a tie. It is the much wider gap on reasoning and agentic knowledge work, where Opus 5 pulls meaningfully ahead of everything else at its price point.

The model ships with a low/medium/high “effort” toggle, a per-request setting that lets developers decide how much reasoning the model spends on a given call. Low effort is cheaper and faster for routine requests; high effort lets Opus 5 think longer on harder problems, trading latency and token spend for accuracy. This is Anthropic’s version of a lever that has become standard across the frontier-model market this year — OpenAI’s GPT-5.6 family ships a similar reasoning slider, and xAI’s Grok 4.5 offers configurable reasoning effort as well — but Anthropic’s implementation is notable for applying the toggle to the Opus tier specifically, rather than reserving it for a separate “mini” or “ultra” branded variant.

Where Opus 5 Still Falls Short

Anthropic has been unusually direct about the model’s limits. The Claude Opus 5 system card states plainly that the model’s cybersecurity-relevant capabilities exceed those of Opus 4.8 but remain behind Mythos 5, the internal-only model Anthropic has not released publicly. Evaluators found that while Opus 5 has become noticeably better at identifying software vulnerabilities, it remains substantially weaker than Mythos 5 at actually exploiting them — a distinction Anthropic’s safety team treats as meaningful enough to spell out rather than bury in a footnote.

The system card also describes testing conducted in partnership with the UK AI Security Institute, which ran external cyber-range evaluations alongside Anthropic’s own five-benchmark suite — ExploitBench, OSS-Fuzz, Firefox 147, and two newly introduced evaluations, CyScenarioBench and ExploitGym. The conclusion Anthropic draws from that combined testing is that Opus 5’s safeguards are calibrated to match those already in place for Fable 5, rather than being loosened to reflect the model’s lower raw capability ceiling.

That caution is a useful data point for anyone tracking how AI safety evaluation has matured this year. It has become routine for major labs to publish not just what a new model can do, but where independent or semi-independent evaluators found it wanting — a pattern that stands in some contrast to earlier release cycles, when capability claims arrived with far less accompanying nuance.

How Opus 5 Fits Into the July Release Wave

Opus 5 did not launch into a quiet market. July 2026 has been one of the most crowded months in frontier AI history: OpenAI’s GPT-5.6 family (Sol, Terra, and Luna) reached general availability on July 9 after a government-gated preview; xAI released Grok 4.5 on July 8 as its first model co-trained with Cursor; Moonshot AI shipped Kimi K3, a 2.8-trillion-parameter open-weight model, on July 16, with weights following on July 26; and Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, released its first open-weights model, Inkling, on July 15. Anthropic itself spent the first days of July restoring access to Claude Fable 5 and Mythos 5 after a US Department of Commerce export-control directive had forced both models offline between June 12 and July 1.

Against that backdrop, Opus 5’s positioning is deliberate. Rather than trying to reclaim the outright capability crown — a title that currently sits with Fable 5 on the benchmarks Anthropic itself publishes — Opus 5 is built to be the model teams reach for by default. It is now the default model on Claude Max and the strongest model available on Claude Pro, while Free-tier users continue to be served by Claude Sonnet 5, which became Anthropic’s new default for that tier on June 30.

Pricing Context Across the Lineup

Understanding where Opus 5 sits requires seeing the full Anthropic price ladder as it stood at the end of July. Sonnet 5 carries an introductory rate in effect through August 31, 2026, after which standard pricing of $3 per million input tokens and $15 per million output tokens applies. Opus 5, at $5/$25, sits one tier above Sonnet and one tier below Fable 5’s flagship pricing. Anthropic bills cached input separately, at one-tenth of the base input rate, and offers a 50% discount on asynchronous batch processing — details that matter more than the headline number for any team running Opus 5 at meaningful volume.

One detail that gets less attention than the pricing table but arguably matters just as much for coding and research workloads: Opus 5 carries a May 2026 knowledge cutoff, a full four months fresher than the January 2026 cutoff shared by both Fable 5 and Opus 4.8. For teams working with recent software releases, current events, or newly published research, that gap can be as consequential as any single benchmark point.

What to Watch Next

Two open questions will shape how Opus 5 is judged over the coming weeks. First, whether independent evaluators — Artificial Analysis, Epoch AI, and others who have become a standard part of how frontier releases get verified in 2026 — reproduce Anthropic’s benchmark claims at scale, rather than the vendor-reported figures that accompanied launch day. Second, whether Anthropic’s stated safeguard parity with Fable 5 holds up as more red-teaming groups get access to the model outside Anthropic’s own testing environment.

Why the Release Cadence Itself Is Becoming the Story

It is worth pausing on the calendar math, because it says something about how the frontier-model market has changed shape over the past year. A single lab shipping four distinct, named models — Mythos 5, Fable 5, Sonnet 5, and now Opus 5 — inside a roughly eight-week window would have been an extraordinary pace even by 2025 standards. In 2026, it has become one of several comparable examples: OpenAI moved from a restricted GPT-5.6 Sol preview to a full three-model general-availability launch inside two weeks, xAI took its V9 architecture from private beta to public release in a little over a week, and Moonshot compressed the gap between its Kimi K3 API launch and open-weight release to eleven days, beating its own stated timeline in the process.

What that compressed cadence means in practice is that the useful unit of comparison for buyers is no longer “which lab has the best model,” a question with an answer that now has a shelf life measured in weeks rather than quarters. It is closer to “which model is the best fit for this specific workload, this month” — a genuinely different kind of decision, and one that rewards infrastructure built to swap models in and out cheaply rather than infrastructure built around a single vendor relationship.

How Enterprises Are Likely to Route Requests

The practical question most engineering teams are asking after Opus 5’s launch is not “should we switch,” but “how should we split traffic.” The emerging pattern among early adopters, based on the benchmark gaps Anthropic itself published, looks less like a wholesale migration and more like a routing decision made per task type. Routine coding tasks, first-pass drafting, and standard agentic workflows are strong candidates for Opus 5’s low- or medium-effort settings, where the cost savings relative to Fable 5 are largest and the capability gap, per Anthropic’s own numbers, is smallest. Harder, more open-ended reasoning tasks — the kind Frontier-Bench v0.1 was specifically built to probe — remain better suited to Fable 5 or to Opus 5’s high-effort setting, where the token cost rises but so does reliability on genuinely difficult problems.

That kind of task-level routing, rather than a single blanket model choice, has become the default recommendation across most of the independent evaluation guides published this month — not just for Opus 5, but for every frontier release in the current cycle. The reasoning-effort toggle Anthropic shipped with Opus 5 is a direct acknowledgment of that shift: rather than forcing developers to choose between two separately priced models, it lets a single model flex across a wider range of cost-capability tradeoffs on a per-request basis, which is likely to become the standard shape of frontier-model pricing going forward rather than an unusual exception.

For now, Opus 5 represents something genuinely useful for the market rather than a marginal update: a model priced like last year’s mid-tier option that performs, on several measures, like this year’s frontier. In a month defined by government-gated previews, export-control disruptions, and a scramble over open-weight licensing, that combination of accessibility and capability may end up being the most quietly significant release of the bunch.

By Ravody

Ravody

Leave a Reply

Your email address will not be published. Required fields are marked *