Introducing Claude Opus 5.5 Anthropic


We’re introducing Claude Opus 5.5, the primary mannequin in our new Claude 5.5 household. It performs on the degree of Claude Fable 5.1 on most work and prices 40% much less to run than Opus 5.

Claude Opus 5.5 is our first launch since we known as for pacing the frontier. It was examined earlier than launch by exterior evaluators, together with Frontier Design and METR. On our automated behavioral audit, probably the most complete alignment check we run, Opus 5.5 is the strongest-performing mannequin we’ve examined thus far. It additionally comes with the safeguards we’ve developed for our most succesful fashions.

Here are a few of the enhancements you may count on from Opus 5.5:

Performance. Opus 5.5 is a significant step up from Opus 5. It’s the brand new main mannequin, and early testers noticed giant jumps in efficiency on their most complicated work. One tester accomplished a 680,000-line code migration in lower than a day—work that might have taken an engineering staff weeks. It’s good at discovering and fixing inefficiencies in software program: after we requested it to chop load instances throughout each web page of an internet app, Opus 5.5 succeeded 39 of 40 instances, whereas Opus 5 made smaller enhancements that additionally altered the app’s habits. A unique tester had a number of Claude fashions construct a recreation from a single immediate; Opus 5.5 scored larger than every other mannequin on the power of its graphics and polish.

Safety. Opus 5.5 achieves one of the best scores of any mannequin thus far on our automated behavioral audit, our alignment suite that exams Claude throughout hundreds of simulated situations. It is far much less seemingly than current fashions to take hard-to-reverse actions or act outdoors the boundaries it’s been given, and it’s extra resistant than Opus 5 to immediate injection. We’ve additionally broadened our alignment testing to cowl longer duties, inconceivable duties, and situations modeled on actual incidents, although it nonetheless has limits. Full particulars of our analysis can be found within the Opus 5.5 System Card.

Because Opus 5.5 is similar to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards just like these on Claude Fable 5.1. Vetted organizations can apply in the present day to our Life Sciences Verification Program to make use of Opus 5.5 for biology analysis. In the approaching weeks we may also be increasing entry to our Cyber Verification Program, and verified cybersecurity practitioners will be capable of use Opus 5.5 for his or her work.

Cost and pace. Opus 5.5 requires much less compute to serve than Opus 5, and its pricing displays that. Our exams present that at default settings it should value 40% lower than Opus 5 on typical workloads. Input and output tokens are $4 and $20 per million, 20% lower than Opus 5. Cache reads (which make up nearly all of agentic and coding work prices) are $0.20 per million tokens, 60% lower than Opus 5. Opus 5.5 additionally generates output greater than 30% sooner than Opus 5.

In addition to the worth drop, we’re growing five-hour utilization limits on Pro, Max, and Team plans. We’re additionally offering subscription customers a charge restrict reset, which now you can save and use everytime you select.

Communication. Opus 5.5 communicates extra naturally than prior fashions. Early testers discovered its writing clearer and simpler to comply with, which addresses a few of the frequent suggestions we heard about Opus 5. It places an important info up entrance, and its type makes it a greater work accomplice over lengthy classes. As one early tester put it, “it writes the way in which I do.” In our personal use, this has made Opus 5.5’s work simpler to comply with and verify—which is a security profit in addition to a sensible one.

Claude Sonnet 5.5 and Claude Haiku 5.5 will comply with within the coming weeks, with most of the identical enhancements to efficiency, effectivity, and security.

Performance and cost-effectiveness

On our benchmarks, Claude Opus 5.5 leads in agentic coding, pc use, and information work. That mentioned, at these ranges of functionality we’ve discovered that benchmark margins have turn out to be a much less dependable information to real-world variations. In our personal use, the hole between Opus 5.5 and Claude Fable 5.1 is narrower than these scores counsel.

Opus 5.5 Fable 5.1 Opus 5 GPT-6 Astra GPT-5.6 Sol
Agentic codingTerminal-Bench 4.0¹ 66.4% 55.8% 52.3% 57.9% 37.3%
Agentic codingFrontierCode v1.1 (Main) 54.4% 50.3% 48.0% 53.3% 47.5%
Agentic codingCursorBench 4.0 57.8% 51.8% 46.6% 41.7%
Knowledge workGDPval-AA v2.1 1846 1735 1708 1542 1588
Business workflowsAutomationBench² 40.0% 31.4% 26.9% 41.4% 28.8%
Multidisciplinary reasoningHumanity’s Last Exam 67.7%with instruments 65.6%with instruments 63.6%with instruments 57.2%with instruments
Agentic scientific analysisTerminal-Bench-Science 0.1³ 58.7% 52.6% 29.0% 64.6% 22.4%
Computer useOSWorld 2.0 81.8%partial 80.7%partial 74.0%partial
Visual chart recognitionChartography 89.0%with instruments 88.4%with instruments 83.4%with instruments

Unless in any other case famous, all Claude Opus 5.5 outcomes use adaptive pondering at max effort. Terminal-Bench 4.0 outcomes are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at excessive effort, as reported by OpenAI; these symbolize every mannequin’s highest rating. Claude Opus 5.5 was evaluated with its manufacturing safeguards enabled. When they intervened, cybersecurity duties had been accomplished by Claude Opus 4.8, and biology and frontier LLM growth duties had been accomplished by Claude Opus 5. This seemingly reduces Claude Opus 5.5’s efficiency on these benchmarks.

1Terminal-Bench 4.0: The normal error is ±2.6 pts for Claude Opus 5.5 and ±1.6–2 pts for the opposite Claude fashions. The public leaderboard (5 trials/process, Claude Code harness) experiences Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, inside noise. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.

2AutomationBench: AutomationBench outcomes had been run and reported by Zapier. These runs had been carried out with out fallback fashions, so safeguard interventions had been thought-about failures—this resulted in a decrease rating than Claude Opus 5.5 would obtain in observe. Claude Opus 5.5 outcomes come from Zapier’s personal analysis throughout early entry. Results for Opus 5, GPT-5.6 Sol, and GPT-6 Astra come from Zapier’s public leaderboard.

3Terminal-Bench-Science 0.1: The normal error is ±3.5–5 pts per mannequin. The public leaderboard (3 trials/process, Claude Code harness) experiences Claude Opus 5 at 30.0%; our setup reproduces it at 29.0%, inside noise. The GPT-6 Astra determine is as reported by OpenAI.

Where Opus 5.5’s benefit could be very clear is effectivity. It prices much less per token than Opus 5 and makes use of fewer tokens per process, which nets out to a 40% drop in prices.

Pricing

Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25

Fast mode for Opus 5.5 can be out there in Claude Code and the Claude Platform with as much as 2.5x pace. It prices $8 per million enter tokens and $40 per million output tokens.

Coding

Opus 5.5 is especially good at lengthy and sprawling jobs like codebase-wide migrations and audits. An early tester used it to audit and repair a 200,000-line codebase in below three hours, the place Opus 5 took over 20 hours and used 2.5x as many tokens. In an inside check, we requested Opus 5.5 and Fable 5.1 to translate HAProxy, broadly used software program that balances net visitors masses throughout servers, from C into Rust. Both rewrites handed practically all of HAProxy’s personal regression exams, however Opus 5.5 completed in 9.5 hours in comparison with 12 for Fable 5.1, and value 51% much less.

Opus 5.5 delivers frontier outcomes on agentic coding at a fraction of the price. At its default effort degree on FrontierCode, it beats GPT-6 Astra at roughly 20% of the price per process. On Terminal Bench 4.0, it matches Astra for about 40% of the price, whereas on CursorBench it beats GPT-5.6 Sol by 11 factors for a couple of third of the price.

Agentic terminal codingAgentic coding: FrontierCodeAgentic coding: CursorBench

Terminal-Bench 4.0Accuracy vs Cost
010203040506070Score (%)251020Cost per try (USD, log scale)lowmedexcessivexhighmax

Our early testers reported related effectivity and intelligence good points:

GitHubClioLovableQuantiumSpotifyOptiverColumnKiro

Quote

“Developers need brokers that may tackle actual software program work and end it. In our testing throughout GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the many fewest tokens and steps we measured. In VS Code, it solved extra terminal duties than Opus 5 in lower than half the steps. More than making particular person duties extra environment friendly, it’s making builders’ larger tasks extra achievable.”

CompanyGitHub

AuthorMario Rodriguez, Chief Product Officer

The most safe coding agent

Enterprises that use brokers inside their methods have to know that these brokers are working as meant, notably once they run autonomously for a lot of hours. Opus 5.5 has a classifier that screens each motion earlier than it runs, an open-source sandbox that safety groups can audit, and code assessment that catches vulnerabilities earlier than they merge.

The mannequin itself additionally has stronger defenses. On immediate injection assaults, it matches or beats Opus 5 in each setting we examined, together with coding, software use, pc use, and net looking. On a benchmark run by the AI safety agency Gray Swan, Opus 5.5 ties Fable 5.1 for the bottom immediate injection success charge of any mannequin examined.

Knowledge work

Opus 5.5 is a dependable and adept researcher. In one inside check, we requested Opus 5.5, Fable 5.1, and Opus 5 to write down a report on an organization’s quarterly efficiency utilizing solely the data it might discover on a duplicate of the online the place the earnings launch was arduous to find. An automated grader checked each determine and quote towards sources. Across totally different effort settings, 16 out of 18 of Opus 5.5’s experiences cleared our high quality bar, the place any invented determine or quote would have failed. Neither Fable 5.1 nor Opus 5 cleared that bar in any try.

It’s additionally robust in monetary evaluation and trade work. Walleye Capital, an funding agency and early tester, reported that Opus 5.5 largely solved their analysis suite on its lowest setting; on larger settings, it carried out even higher, noticing an error of their analysis directions and correcting for it. No different mannequin had caught this error earlier than.

In one other check, we tasked each Opus 5.5 and Opus 5 with analyzing a proposed merger between two fictional HR software program firms. Each constructed a monetary mannequin in Excel, then turned it into an government presentation on whether or not the deal made sense at its value. Both fashions reached the identical conclusions concerning the deal, however Opus 5.5’s mannequin was extra thorough and its presentation simpler to learn, whereas Opus 5’s had minor errors. Opus 5.5 completed in 63 minutes in comparison with 93 for Opus 5, and value 50% much less to supply.

On information work evaluations, Opus 5.5 outperforms different fashions whereas additionally utilizing fewer tokens. On GDPval-AA v2.1, a check of real-world work throughout 44 occupations, Opus 5.5 scores 1846 Elo, forward of Fable 5.1 and Opus 5. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for a couple of fifth of the price per process. It likewise outperformed different fashions on benchmarks measuring trade workflows and large-scale knowledge assortment.

GDPval-AA v2.1AutomationBenchWANDR

GDPval-AA v2.1Elo vs Cost
12001300140015001600170018000Elo0.200.5012510Estimated value per process (USD, log scale)lowmedexcessivexhighmax

Our prospects have reported related outcomes. Here’s what they advised us about working with the mannequin:

Deloitte Consulting LLPRogoLexisNexis Legal & ProfessionalWalleye CapitalHexThomson Reuters LabsHebbiaViktor

Quote

“Even at its lowest effort setting, Claude Opus 5.5 caught 72% of recognized bugs in our code critiques to Opus 5’s 56% at excessive effort, with fewer false alarms and a fraction of the output. On US consulting evaluation, low pondering effort matched its larger pondering settings on half the output and handed our high quality checks. When extra decrease pondering efforts are deployed in manufacturing, that’s client-ready work delivered effectively.”

CompanyDeloitte Consulting LLP

AuthorCarl Bennett, CIO

Communication

We’ve made main enhancements to the way in which Opus 5.5 writes and communicates, one of the frequent areas of suggestions we heard about Opus 5. Its messages are a lot simpler to know at a look, which testers mentioned helped throughout lengthy working classes. It places an important info up entrance, is much less seemingly to make use of jargon or idiosyncratic phrases, and follows the writing guidelines you give it. We discover that this makes Opus 5.5 a noticeably higher collaborator. Here’s a side-by-side comparability of the 2 fashions:

Our prospects’ suggestions helps these findings:

RampStripeBoxChicago Trading CompanyFactory

Quote

“Verbose, hard-to-follow output has been my largest frustration with frontier fashions, and Claude Opus 5.5 fixes it. It writes like colleague, and follows our writing guidelines. A design spec got here out usable with very minimal edits, and when it rewrote certainly one of our prompts I most popular its model to my very own. When it optimized our check suite, I might comply with its reasoning simply and shipped the change with confidence.”

CompanyRamp

AuthorJohn Ruelas, Staff Software Engineer

Safety

Pacing the frontier

Last week, our CEO, Dario Amodei, argued that AI progress should be paced in order that security practices keep forward of mannequin capabilities. Pacing is an method to maintaining AI protected, remaining aggressive with China, and realizing AI’s advantages, notably in areas like biology and drugs.

We largely perceive the dangers in the present day’s fashions current and are nicely outfitted to handle them. However, extra critical dangers might emerge rapidly as capabilities enhance, and we have to put together for them now. For that purpose, our security work takes place on two time horizons without delay:

Safety practices for present fashions. The present era of fashions depends on a longtime set of practices: in depth alignment testing, pre-release analysis by outdoors organizations akin to METR and Frontier Design, and safeguards matched to every mannequin’s capabilities in high-risk areas like cybersecurity and biology. We refine these practices with every launch. We imagine they’re applicable to the worst dangers in the present day’s fashions current, and imagine they offer us a broad, although not excellent image of the vary of great dangers.

Additionally, we observe our capacity to coach and consider aligned fashions, and report on each our public and inside fashions within the danger experiences we publish below our Responsible Scaling Policy, our voluntary framework for managing catastrophic dangers from superior AI methods.

Preparing for future fashions. We’re making ready our coaching and analysis processes in anticipation of extra superior fashions. We’re tightening how we filter the environments utilized in reinforcement studying, since flawed environments are a significant supply of misaligned habits. Additionally, we’re enhancing our alignment rewards and creating automated processes for producing new, numerous situations for security coaching. And we’re strengthening our security and monitoring, together with a targeted effort to enhance interpretability-based monitoring and analysis. We hope such methods will assist scale back our reliance on auditing a mannequin’s chain-of-thought, or the reasoning it writes out whereas it really works.

Models with higher capabilities—akin to these that may absolutely automate the work of AI analysis itself—require the next security normal nonetheless. Our requires pacing had been primarily based largely on our expectation that such fashions might be skilled quickly. For these fashions, we don’t assume the measures described above will meet that security normal on their very own. As AI turns into extra succesful, public coverage ought to play a bigger function in ensuring the methods individuals depend on are protected. That capability takes time to construct, and we’ve began to place the transport systems in place to help it, as described in “We Must Pace the Frontier” and our current announcement with Accenture; we count on to share extra particulars on these efforts quickly. We may also proceed to contribute to coverage discussions with authorities and trade, together with on approaches to regulation and overseas coordination.

Alignment

On our major analysis suite, an automatic behavioral audit that assesses Claude throughout practically 2,000 situations, Opus 5.5 scored higher than any current Claude mannequin on practically each measure of misaligned habits. It’s additionally our strongest mannequin on most measures of honesty.

In explicit, Opus 5.5 improves over earlier fashions on a number of of the behaviors that contributed to current cybersecurity incidents, together with biased or motivated reasoning, trying to flee a sandbox, and taking dangerous actions after concluding it was in a simulated setting. In a brand new analysis designed to check a mannequin’s propensity to cross containment boundaries, Opus 5.5 tried to avoid boundaries round 85% much less usually than Opus 5 or Claude Mythos 5.1, and each try it made was low severity and self-reported. For groups working Claude unattended throughout their codebases and methods, that is simply as essential as uncooked functionality.

However, as we described in our current alignment assessment, constructing evaluations that reliably catch each failure previous to deployment stays an unsolved drawback. We see indicators that Opus 5.5 usually suspects it’s being evaluated, which challenges our capacity to evaluate the way it will act within the huge number of real-world settings it’s deployed in. As these settings develop and mannequin capabilities improve, we count on this problem to develop, until we make progress on interpretability. Although we’re assured that Opus 5.5 reveals broad enhancements within the areas we’re capable of measure, we pair our personal alignment work with the safeguards described beneath.

Safeguards

As our fashions develop extra highly effective, stricter safeguards are a method we stop new capabilities from changing into instruments for misuse. Opus 5.5 is the primary Opus mannequin to launch with the same class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall again to a different mannequin transparently.

Cybersecurity. Because Opus 5.5 has extraordinarily robust cyber capabilities, we’re making use of cybersecurity safeguards to Opus 5.5 which can be just like Fable 5.1’s. Users will be capable of determine and repair bugs of their code as a part of the routine software program growth lifecycle, however most cybersecurity duties will likely be re-routed to Opus 4.8.

For cyberdefenders, we’ll quickly be increasing our Cyber Verification Program to incorporate Opus 5.5. The new program will embrace three tiers for more and more permissive trusted entry, together with entry to Claude Mythos fashions. Claude Security is already out there with entry to Claude Mythos 5.1.

Biology. Opus 5.5 is extremely succesful in biology, exceeding Opus 5 and matching or beating Claude Mythos 5.1 throughout many areas of labor. For instance, Opus 5.5 achieved enhancements on a long-horizon molecular prediction and design analysis performed in collaboration with Dyno Therapeutics, and professional red-teamers rated its scientific novelty as similar to one of the best mannequin that they had examined.

For this purpose, Opus 5.5 makes use of the identical biology safeguards as Fable 5.1. To use Opus 5.5 for analysis and growth work impeded by these safeguards, customers can apply to our new Life Sciences Verification Program, which provides vetted organizations like educational labs, startups, and pharmaceutical firms entry to safeguards designed for the total breadth of biology-related work. Interested organizations can apply here.

Distillation

Distillation assaults, during which attackers use hundreds of pretend accounts to extract a mannequin’s capabilities at industrial scale, create security and nationwide safety dangers. Distillation permits unhealthy actors to create extremely succesful fashions with out the safeguards we construct into Claude. Our September 2026 threat intelligence report particulars the illicit distillation exercise we’ve detected and disrupted up to now.

Opus 5.5 is launching with preserved pondering, the anti-distillation safeguard we launched with Fable 5.1. It stops API customers from modifying Claude’s prior context in an try to extract Claude’s reasoning. It applies to Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026. Our Help Center article explains the change, and our preserved thinking docs present tips on how to check and replace your integrations.

Data retention and compliance

Like earlier Opus fashions, Opus 5.5 is accessible with zero knowledge retention.

As with Fable 5.1, Opus 5.5 comes with our watermarking measures to adjust to the EU AI Act, discussed here. It can be not out there with “pondering” mode switched off, as we describe here.

Availability

Claude Opus 5.5 is now out there on all platforms, together with Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, builders can get started with claude-opus-5-5.



Source link