Introducing Claude Opus 5.5 Anthropic
We’re introducing Claude Opus 5.5, the primary mannequin in our new Claude 5.5 household. It performs on the degree of Claude Fable 5.1 on most work and prices 40% much less to run than Opus 5.
Claude Opus 5.5 is our first launch since we known as for pacing the frontier. It was examined earlier than launch by exterior evaluators, together with Frontier Design and METR. On our automated behavioral audit, probably the most complete alignment check we run, Opus 5.5 is the strongest-performing mannequin we’ve examined thus far. It additionally comes with the safeguards we’ve developed for our most succesful fashions.
Here are a few of the enhancements you may count on from Opus 5.5:
Performance. Opus 5.5 is a significant step up from Opus 5. It’s the brand new main mannequin, and early testers noticed giant jumps in efficiency on their most complicated work. One tester accomplished a 680,000-line code migration in lower than a day—work that might have taken an engineering staff weeks. It’s good at discovering and fixing inefficiencies in software program: after we requested it to chop load instances throughout each web page of an internet app, Opus 5.5 succeeded 39 of 40 instances, whereas Opus 5 made smaller enhancements that additionally altered the app’s habits. A unique tester had a number of Claude fashions construct a recreation from a single immediate; Opus 5.5 scored larger than every other mannequin on the power of its graphics and polish.
Safety. Opus 5.5 achieves one of the best scores of any mannequin thus far on our automated behavioral audit, our alignment suite that exams Claude throughout hundreds of simulated situations. It is far much less seemingly than current fashions to take hard-to-reverse actions or act outdoors the boundaries it’s been given, and it’s extra resistant than Opus 5 to immediate injection. We’ve additionally broadened our alignment testing to cowl longer duties, inconceivable duties, and situations modeled on actual incidents, although it nonetheless has limits. Full particulars of our analysis can be found within the Opus 5.5 System Card.
Because Opus 5.5 is similar to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards just like these on Claude Fable 5.1. Vetted organizations can apply in the present day to our Life Sciences Verification Program to make use of Opus 5.5 for biology analysis. In the approaching weeks we may also be increasing entry to our Cyber Verification Program, and verified cybersecurity practitioners will be capable of use Opus 5.5 for his or her work.
Cost and pace. Opus 5.5 requires much less compute to serve than Opus 5, and its pricing displays that. Our exams present that at default settings it should value 40% lower than Opus 5 on typical workloads. Input and output tokens are $4 and $20 per million, 20% lower than Opus 5. Cache reads (which make up nearly all of agentic and coding work prices) are $0.20 per million tokens, 60% lower than Opus 5. Opus 5.5 additionally generates output greater than 30% sooner than Opus 5.
In addition to the worth drop, we’re growing five-hour utilization limits on Pro, Max, and Team plans. We’re additionally offering subscription customers a charge restrict reset, which now you can save and use everytime you select.
Communication. Opus 5.5 communicates extra naturally than prior fashions. Early testers discovered its writing clearer and simpler to comply with, which addresses a few of the frequent suggestions we heard about Opus 5. It places an important info up entrance, and its type makes it a greater work accomplice over lengthy classes. As one early tester put it, “it writes the way in which I do.” In our personal use, this has made Opus 5.5’s work simpler to comply with and verify—which is a security profit in addition to a sensible one.
Claude Sonnet 5.5 and Claude Haiku 5.5 will comply with within the coming weeks, with most of the identical enhancements to efficiency, effectivity, and security.
Performance and cost-effectiveness
On our benchmarks, Claude Opus 5.5 leads in agentic coding, pc use, and information work. That mentioned, at these ranges of functionality we’ve discovered that benchmark margins have turn out to be a much less dependable information to real-world variations. In our personal use, the hole between Opus 5.5 and Claude Fable 5.1 is narrower than these scores counsel.
| Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol | |
|---|---|---|---|---|---|
| Agentic codingTerminal-Bench 4.0¹ | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| Agentic codingFrontierCode v1.1 (Main) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| Agentic codingCursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| Knowledge workGDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 | 1588 |
| Business workflowsAutomationBench² | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Multidisciplinary reasoningHumanity’s Last Exam | 67.7%with instruments | 65.6%with instruments | 63.6%with instruments | 57.2%with instruments | — |
| Agentic scientific analysisTerminal-Bench-Science 0.1³ | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| Computer useOSWorld 2.0 | 81.8%partial | 80.7%partial | 74.0%partial | — | — |
| Visual chart recognitionChartography | 89.0%with instruments | 88.4%with instruments | 83.4%with instruments | — | — |
Unless in any other case famous, all Claude Opus 5.5 outcomes use adaptive pondering at max effort. Terminal-Bench 4.0 outcomes are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at excessive effort, as reported by OpenAI; these symbolize every mannequin’s highest rating. Claude Opus 5.5 was evaluated with its manufacturing safeguards enabled. When they intervened, cybersecurity duties had been accomplished by Claude Opus 4.8, and biology and frontier LLM growth duties had been accomplished by Claude Opus 5. This seemingly reduces Claude Opus 5.5’s efficiency on these benchmarks.
1Terminal-Bench 4.0: The normal error is ±2.6 pts for Claude Opus 5.5 and ±1.6–2 pts for the opposite Claude fashions. The public leaderboard (5 trials/process, Claude Code harness) experiences Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, inside noise. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
2AutomationBench: AutomationBench outcomes had been run and reported by Zapier. These runs had been carried out with out fallback fashions, so safeguard interventions had been thought-about failures—this resulted in a decrease rating than Claude Opus 5.5 would obtain in observe. Claude Opus 5.5 outcomes come from Zapier’s personal analysis throughout early entry. Results for Opus 5, GPT-5.6 Sol, and GPT-6 Astra come from Zapier’s public leaderboard.
3Terminal-Bench-Science 0.1: The normal error is ±3.5–5 pts per mannequin. The public leaderboard (3 trials/process, Claude Code harness) experiences Claude Opus 5 at 30.0%; our setup reproduces it at 29.0%, inside noise. The GPT-6 Astra determine is as reported by OpenAI.
Where Opus 5.5’s benefit could be very clear is effectivity. It prices much less per token than Opus 5 and makes use of fewer tokens per process, which nets out to a 40% drop in prices.
Pricing
| Prices per 1M tokens | Claude Opus 5.5 | Claude Opus 5 |
|---|---|---|
| Cache reads | $0.20 | $0.50 |
| Input tokens | $4 | $5 |
| Output tokens | $20 | $25 |
| Cache writes | $5 | $6.25 |
Fast mode for Opus 5.5 can be out there in Claude Code and the Claude Platform with as much as 2.5x pace. It prices $8 per million enter tokens and $40 per million output tokens.
Coding
Opus 5.5 is especially good at lengthy and sprawling jobs like codebase-wide migrations and audits. An early tester used it to audit and repair a 200,000-line codebase in below three hours, the place Opus 5 took over 20 hours and used 2.5x as many tokens. In an inside check, we requested Opus 5.5 and Fable 5.1 to translate HAProxy, broadly used software program that balances net visitors masses throughout servers, from C into Rust. Both rewrites handed practically all of HAProxy’s personal regression exams, however Opus 5.5 completed in 9.5 hours in comparison with 12 for Fable 5.1, and value 51% much less.
Opus 5.5 delivers frontier outcomes on agentic coding at a fraction of the price. At its default effort degree on FrontierCode, it beats GPT-6 Astra at roughly 20% of the price per process. On Terminal Bench 4.0, it matches Astra for about 40% of the price, whereas on CursorBench it beats GPT-5.6 Sol by 11 factors for a couple of third of the price.
Agentic terminal codingAgentic coding: FrontierCodeAgentic coding: CursorBench
Our early testers reported related effectivity and intelligence good points:
GitHubClioLovableQuantiumSpotifyOptiverColumnKiro
The most safe coding agent
Enterprises that use brokers inside their methods have to know that these brokers are working as meant, notably once they run autonomously for a lot of hours. Opus 5.5 has a classifier that screens each motion earlier than it runs, an open-source sandbox that safety groups can audit, and code assessment that catches vulnerabilities earlier than they merge.
The mannequin itself additionally has stronger defenses. On immediate injection assaults, it matches or beats Opus 5 in each setting we examined, together with coding, software use, pc use, and net looking. On a benchmark run by the AI safety agency Gray Swan, Opus 5.5 ties Fable 5.1 for the bottom immediate injection success charge of any mannequin examined.
Knowledge work
Opus 5.5 is a dependable and adept researcher. In one inside check, we requested Opus 5.5, Fable 5.1, and Opus 5 to write down a report on an organization’s quarterly efficiency utilizing solely the data it might discover on a duplicate of the online the place the earnings launch was arduous to find. An automated grader checked each determine and quote towards sources. Across totally different effort settings, 16 out of 18 of Opus 5.5’s experiences cleared our high quality bar, the place any invented determine or quote would have failed. Neither Fable 5.1 nor Opus 5 cleared that bar in any try.
It’s additionally robust in monetary evaluation and trade work. Walleye Capital, an funding agency and early tester, reported that Opus 5.5 largely solved their analysis suite on its lowest setting; on larger settings, it carried out even higher, noticing an error of their analysis directions and correcting for it. No different mannequin had caught this error earlier than.
In one other check, we tasked each Opus 5.5 and Opus 5 with analyzing a proposed merger between two fictional HR software program firms. Each constructed a monetary mannequin in Excel, then turned it into an government presentation on whether or not the deal made sense at its value. Both fashions reached the identical conclusions concerning the deal, however Opus 5.5’s mannequin was extra thorough and its presentation simpler to learn, whereas Opus 5’s had minor errors. Opus 5.5 completed in 63 minutes in comparison with 93 for Opus 5, and value 50% much less to supply.
On information work evaluations, Opus 5.5 outperforms different fashions whereas additionally utilizing fewer tokens. On GDPval-AA v2.1, a check of real-world work throughout 44 occupations, Opus 5.5 scores 1846 Elo, forward of Fable 5.1 and Opus 5. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for a couple of fifth of the price per process. It likewise outperformed different fashions on benchmarks measuring trade workflows and large-scale knowledge assortment.
GDPval-AA v2.1AutomationBenchWANDR
Our prospects have reported related outcomes. Here’s what they advised us about working with the mannequin:
Deloitte Consulting LLPRogoLexisNexis Legal & ProfessionalWalleye CapitalHexThomson Reuters LabsHebbiaViktor
Communication
We’ve made main enhancements to the way in which Opus 5.5 writes and communicates, one of the frequent areas of suggestions we heard about Opus 5. Its messages are a lot simpler to know at a look, which testers mentioned helped throughout lengthy working classes. It places an important info up entrance, is much less seemingly to make use of jargon or idiosyncratic phrases, and follows the writing guidelines you give it. We discover that this makes Opus 5.5 a noticeably higher collaborator. Here’s a side-by-side comparability of the 2 fashions:
Claude Opus 5
Please clarify the problem to me.
What I discovered
The additional drop isn’t the free tier — it’s a regression in 0552feb (“extract billing interval boundaries into intervals.py”, labelled “No behaviour change”).
The bug
mixture.monthly_usage used to do a half-open interval:
begin, finish = month_start(12 months, month), next_month_start(12 months, month)
if begin <= e.ts < finish:
It now does:
if interval.accommodates(e.ts)
and in intervals.py:
@property
def last_day(self) -> datetime:
return next_month_start(...) - timedelta(days=1) # 2026-08-31 00:00:00 UTC
def accommodates(self, ts):
return self.first_day <= ts <= self.last_day
last_day is a datetime at midnight, not a entire day. Subtracting one day from Sep 1 00:00 provides Aug 31 00:00, and the comparability is inclusive, so the interval successfully ends at the first prompt of the thirty first. Every occasion timestamped after midnight on the final day of the month is silently dropped — roughly one day of utilization per account, per month.
The new test_periods.py doesn’t catch it as a result of it solely checks last_day.date() (which is appropriately 2026-08-31) and exams accommodates at Aug 1 00:00, Aug 15 12:00, Jul 31 23:59 and Sep 1 00:00 — it by no means probes wherever inside the thirty first.
Claude Opus 5.5
Please clarify the problem to me.
The additional drop is a bug in the billing refactor
The free-tier change accounts for solely $1.50 of acme’s August drop. The different $9.92 comes from a bug in commit 0552feb. That commit was labelled “No behaviour change,” however it stops counting utilization from the final day of the month.
What modified
Before the merge, mixture.py used a half-open interval:
start <= e.ts < next_month_start(year, month) # [Aug 1 00:00, Sep 1 00:00)
The refactor moved this into Period.contains and made the interval closed on the wrong boundary:
last_day = next_month_start(...) - timedelta(days=1) # = 2026-08-31T00:00:00Z
return self.first_day <= ts <= self.last_day
last_day is a timestamp at midnight, not a whole day. So every event after 2026-08-31T00:00:00Z is excluded. The September period starts Sep 1, so it doesn’t pick these events up either. The usage isn’t moved to another month; it’s never billed at all.
Claude Opus 5
Summarize this Slack thread for my manager in three bullet points:
[09:41] Dana Osei: I’ll inform them. I’ll additionally affirm no one else reads the Friday file.
[09:43] Marcus Webb: I can transfer the schedule, it’s a one-line cron change. But not till Dana confirms.
[10:02] Dana Osei: Finance says the Friday file is just used Monday. Sunday evening is okay.
Friday orders export — #data-platform, Tuesday
- Problem: The Friday orders export timed out once more (second failure this month) after working 2h on the warehouse question; order quantity is up ~40% since the summer time sale. Finance was left with out weekend numbers for their Monday 9am assessment, and each failures solely surfaced on Monday.
- Fixed identical day: The Saturday ops dashboard tile was repointed to learn the reside desk as a substitute of the export (Sam), finance confirmed the Friday file is solely used on Mondays (Dana), a failure alert now pages the on-call channel (Priya), and the export strikes to Sunday 10pm from subsequent week with the Friday run dropped (Marcus). A one-off ~$300 warehouse improve covers this Friday as a stopgap.
- Still in progress: Ji-woo owns rewriting the question to run incrementally (re-scanning the final 3 days to catch late-arriving orders), anticipated to minimize runtime from hours to minutes — design doc Thursday, roughly two weeks together with testing. The schedule change buys time in the meantime.
Claude Opus 5.5
Summarize this Slack thread for my supervisor in three bullet factors:
[09:41] Dana Osei: I’ll inform them. I’ll additionally affirm no one else reads the Friday file.
[09:43] Marcus Webb: I can transfer the schedule, it’s a one-line cron change. But not till Dana confirms.
[10:02] Dana Osei: Finance says the Friday file is just used Monday. Sunday evening is okay.
- Problem: The Friday orders export has failed twice this month. The warehouse question instances out after 2 hours as a result of order quantity is up ~40% since the summer time sale. As a consequence, Finance was lacking weekend numbers for their Monday 9am assessment.
- Immediate fixes: The export will transfer to Sunday 10pm beginning subsequent week, and the Friday run will be dropped (Marcus). Finance has confirmed this works for them (Dana). The ops dashboard tile that depended on the Friday file now reads from the reside desk (Sam, completed). A failure alert now pages the on-call channel (Priya, completed and examined). For this Friday solely, Marcus will improve the warehouse slot as a ~$300 one-off stopgap, which Dana accepted.
- Longer-term repair: Ji-woo will rewrite the question to course of solely new orders, re-scanning the final 3 days to catch late arrivals. This ought to minimize runtime from hours to minutes. The work will take ~2 weeks together with testing, with a design doc due for assessment by Thursday.
Claude Opus 5
Write Python that reads a chess place (FEN) and lists which items are below assault. Use TensorFlow for pace. Here’s my design: construct an 8×8 tensor board, generate an assault map for every bit kind from its motion sample, and flag opposing items on attacked squares. Change something you assume is incorrect, and in your remaining abstract clarify every change you made and why.
tf.math.cumsum replaces the bit scan. The geometry is unrolled into ray_targets[square, direction, step] — an ordered stroll outwards. An unique cumsum over the step axis counts blockers strictly earlier than every step, so the first blocker on a ray is the single step that is blocking with nothing blocking earlier than it. This is the vectorized kind of the bit scan a bitboard engine makes use of. I did look at bitwise_ops for true 64-bit bitboards — population_count exists, however hyperbola-quintessence wants a byte-reverse primitive TF would not expose, so it would have fallen again to per-direction loops anyway.
Claude Opus 5.5
Write Python that reads a chess place (FEN) and lists which items are below assault. Use TensorFlow for pace. Here’s my design: construct an 8×8 tensor board, generate an assault map for every bit kind from its motion sample, and flag opposing items on attacked squares. Change something you assume is incorrect, and in your remaining abstract clarify every change you made and why.
Blocked traces (bishops, rooks, queens). A fastened masks cannot say “cease at the first piece in the method.” I added a second stack of masks marking the squares strictly between each two squares on the identical line. One matrix multiply with the board’s occupancy counts the blockers on each line at as soon as, and an assault survives solely if that depend is zero. Knights and single steps have nothing between them, so one rule covers each piece.
Our prospects’ suggestions helps these findings:
RampStripeBoxChicago Trading CompanyFactory
Safety
Pacing the frontier
Last week, our CEO, Dario Amodei, argued that AI progress should be paced in order that security practices keep forward of mannequin capabilities. Pacing is an method to maintaining AI protected, remaining aggressive with China, and realizing AI’s advantages, notably in areas like biology and drugs.
We largely perceive the dangers in the present day’s fashions current and are nicely outfitted to handle them. However, extra critical dangers might emerge rapidly as capabilities enhance, and we have to put together for them now. For that purpose, our security work takes place on two time horizons without delay:
Safety practices for present fashions. The present era of fashions depends on a longtime set of practices: in depth alignment testing, pre-release analysis by outdoors organizations akin to METR and Frontier Design, and safeguards matched to every mannequin’s capabilities in high-risk areas like cybersecurity and biology. We refine these practices with every launch. We imagine they’re applicable to the worst dangers in the present day’s fashions current, and imagine they offer us a broad, although not excellent image of the vary of great dangers.
Additionally, we observe our capacity to coach and consider aligned fashions, and report on each our public and inside fashions within the danger experiences we publish below our Responsible Scaling Policy, our voluntary framework for managing catastrophic dangers from superior AI methods.
Preparing for future fashions. We’re making ready our coaching and analysis processes in anticipation of extra superior fashions. We’re tightening how we filter the environments utilized in reinforcement studying, since flawed environments are a significant supply of misaligned habits. Additionally, we’re enhancing our alignment rewards and creating automated processes for producing new, numerous situations for security coaching. And we’re strengthening our security and monitoring, together with a targeted effort to enhance interpretability-based monitoring and analysis. We hope such methods will assist scale back our reliance on auditing a mannequin’s chain-of-thought, or the reasoning it writes out whereas it really works.
Models with higher capabilities—akin to these that may absolutely automate the work of AI analysis itself—require the next security normal nonetheless. Our requires pacing had been primarily based largely on our expectation that such fashions might be skilled quickly. For these fashions, we don’t assume the measures described above will meet that security normal on their very own. As AI turns into extra succesful, public coverage ought to play a bigger function in ensuring the methods individuals depend on are protected. That capability takes time to construct, and we’ve began to place the transport systems in place to help it, as described in “We Must Pace the Frontier” and our current announcement with Accenture; we count on to share extra particulars on these efforts quickly. We may also proceed to contribute to coverage discussions with authorities and trade, together with on approaches to regulation and overseas coordination.
Alignment
On our major analysis suite, an automatic behavioral audit that assesses Claude throughout practically 2,000 situations, Opus 5.5 scored higher than any current Claude mannequin on practically each measure of misaligned habits. It’s additionally our strongest mannequin on most measures of honesty.
In explicit, Opus 5.5 improves over earlier fashions on a number of of the behaviors that contributed to current cybersecurity incidents, together with biased or motivated reasoning, trying to flee a sandbox, and taking dangerous actions after concluding it was in a simulated setting. In a brand new analysis designed to check a mannequin’s propensity to cross containment boundaries, Opus 5.5 tried to avoid boundaries round 85% much less usually than Opus 5 or Claude Mythos 5.1, and each try it made was low severity and self-reported. For groups working Claude unattended throughout their codebases and methods, that is simply as essential as uncooked functionality.
However, as we described in our current alignment assessment, constructing evaluations that reliably catch each failure previous to deployment stays an unsolved drawback. We see indicators that Opus 5.5 usually suspects it’s being evaluated, which challenges our capacity to evaluate the way it will act within the huge number of real-world settings it’s deployed in. As these settings develop and mannequin capabilities improve, we count on this problem to develop, until we make progress on interpretability. Although we’re assured that Opus 5.5 reveals broad enhancements within the areas we’re capable of measure, we pair our personal alignment work with the safeguards described beneath.
Safeguards
As our fashions develop extra highly effective, stricter safeguards are a method we stop new capabilities from changing into instruments for misuse. Opus 5.5 is the primary Opus mannequin to launch with the same class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall again to a different mannequin transparently.
Cybersecurity. Because Opus 5.5 has extraordinarily robust cyber capabilities, we’re making use of cybersecurity safeguards to Opus 5.5 which can be just like Fable 5.1’s. Users will be capable of determine and repair bugs of their code as a part of the routine software program growth lifecycle, however most cybersecurity duties will likely be re-routed to Opus 4.8.
For cyberdefenders, we’ll quickly be increasing our Cyber Verification Program to incorporate Opus 5.5. The new program will embrace three tiers for more and more permissive trusted entry, together with entry to Claude Mythos fashions. Claude Security is already out there with entry to Claude Mythos 5.1.
Biology. Opus 5.5 is extremely succesful in biology, exceeding Opus 5 and matching or beating Claude Mythos 5.1 throughout many areas of labor. For instance, Opus 5.5 achieved enhancements on a long-horizon molecular prediction and design analysis performed in collaboration with Dyno Therapeutics, and professional red-teamers rated its scientific novelty as similar to one of the best mannequin that they had examined.
For this purpose, Opus 5.5 makes use of the identical biology safeguards as Fable 5.1. To use Opus 5.5 for analysis and growth work impeded by these safeguards, customers can apply to our new Life Sciences Verification Program, which provides vetted organizations like educational labs, startups, and pharmaceutical firms entry to safeguards designed for the total breadth of biology-related work. Interested organizations can apply here.
Distillation
Distillation assaults, during which attackers use hundreds of pretend accounts to extract a mannequin’s capabilities at industrial scale, create security and nationwide safety dangers. Distillation permits unhealthy actors to create extremely succesful fashions with out the safeguards we construct into Claude. Our September 2026 threat intelligence report particulars the illicit distillation exercise we’ve detected and disrupted up to now.
Opus 5.5 is launching with preserved pondering, the anti-distillation safeguard we launched with Fable 5.1. It stops API customers from modifying Claude’s prior context in an try to extract Claude’s reasoning. It applies to Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026. Our Help Center article explains the change, and our preserved thinking docs present tips on how to check and replace your integrations.
Data retention and compliance
Like earlier Opus fashions, Opus 5.5 is accessible with zero knowledge retention.
As with Fable 5.1, Opus 5.5 comes with our watermarking measures to adjust to the EU AI Act, discussed here. It can be not out there with “pondering” mode switched off, as we describe here.
Availability
Claude Opus 5.5 is now out there on all platforms, together with Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, builders can get started with claude-opus-5-5.


