All About AI covers Claude Opus 4.7 in this video.
Introducing Claude Opus 4.7
Our latest model, Claude Opus 4.7, is now generally available.
Opus 4.7 is a notable improvement on Opus 4.6 in advanced software engineering, with particular gains on the most difficult tasks. Users report being able to hand off their hardest coding work, the kind that previously needed close supervision, to Opus 4.7 with confidence. Opus 4.7 handles complex, long-running tasks with rigor and consistency, pays precise attention to instructions, and devises ways to verify its own outputs before reporting back.
The model also has substantially better vision: it can see images in greater resolution. It’s more tasteful and creative when completing professional tasks, producing higher-quality interfaces, slides, and docs. And, although it is less broadly capable than our most powerful model, Claude Mythos Preview, it shows better results than Opus 4.6 across a range of benchmarks.

Comparison across Opus 4.7, Opus 4.6, GPT-5.4, Gemini 3.1 Pro, and Mythos Preview.
Last week we announced Project Glasswing, highlighting the risks, and benefits, of AI models for cybersecurity. We stated that we would keep Claude Mythos Preview’s release limited and test new cyber safeguards on less capable models first.
Opus 4.7 is the first such model: its cyber capabilities are not as advanced as those of Mythos Preview. We are releasing Opus 4.7 with safeguards that automatically detect and block requests that indicate prohibited or high-risk cybersecurity uses.
Opus 4.7 is available today across all Claude products and our API, Amazon Bedrock, Google Cloud’s Vertex AI, and Microsoft Foundry. Pricing remains the same as Opus 4.6: $5 per million input tokens and $25 per million output tokens.
The Benchmark Picture
Anthropic's own framing of Opus 4.7 is that it is a meaningful step on hard software engineering rather than a clean sweep, and the published numbers bear that out. Below is how the model lands against Opus 4.6, the frontier competition, and Anthropic's own unreleased Mythos Preview, drawn from Anthropic's launch materials, AWS's Bedrock announcement and third-party benchmark aggregation.
| Benchmark | Opus 4.6 | Opus 4.7 | Best rival |
|---|---|---|---|
| SWE-bench Verified | 80.8% | 87.6% | Mythos Preview 93.9% |
| SWE-bench Pro | 53.4% | 64.3% | GPT-5.4 57.7% |
| Terminal-Bench 2.0 | 65.4% | 69.4% | GPT-5.4 75.1% |
| MCP-Atlas (tool use) | 75.8% | 77.3% | Gemini 3.1 Pro 73.9% |
| Finance Agent v1.1 | 60.1% | 64.4% | GPT-5.4 Pro 61.5% |
| OSWorld-Verified (computer use) | 72.7% | 78.0% | Mythos Preview 79.6% |
| GPQA Diamond | 91.3% | 94.2% | Mythos Preview 94.6% |
| CharXiv Reasoning (no tools) | 69.1% | 82.1% | Mythos Preview 86.1% |
| BrowseComp (web research) | 83.7% | 79.3% | GPT-5.4 Pro 89.3% |
Three things stand out. The first is the size of the coding jump: nearly seven points on SWE-bench Verified and close to eleven on the harder, multi-language SWE-bench Pro. That is a larger single-release gain than the 4.5-to-4.6 step produced, and it is concentrated exactly where the marketing says it is, on the difficult end of the distribution rather than on tasks the previous model already handled.
The second is CharXiv. Reasoning over scientific charts and figures rises from 69.1% to 82.1% without tools, which is by some distance the biggest proportional move on the board. That is the vision upgrade showing up in a measurable place rather than as a marketing claim, and it is the number to point at if you are trying to justify the migration to anyone who works with diagrams, dashboards or dense PDFs.
The third is the one Anthropic does not lead with. BrowseComp went backwards, from 83.7% on Opus 4.6 to 79.3% on 4.7, and it now trails both GPT-5.4 Pro and Gemini 3.1 Pro on open-web research. If your workload is primarily deep web research rather than engineering, 4.7 is not automatically an upgrade and you should measure before switching. Terminal-Bench 2.0 tells a similar if less dramatic story: 4.7 improves on its predecessor but still sits behind GPT-5.4 on raw command-line work.
Read the ceiling, not just the score
Mythos Preview leads or ties on most of this table, and Anthropic has been explicit that it is deliberately holding that model back on cyber-capability grounds. Opus 4.7 is therefore best understood as the most capable model Anthropic is currently willing to put in general circulation, which is a different claim to being the most capable model it has.
On availability, Opus 4.7 ships with a 1 million token context window and is exposed through Amazon Bedrock as anthropic.claude-opus-4-7, initially in US East (N. Virginia), Asia Pacific (Tokyo), Europe (Ireland) and Europe (Stockholm). For how this sits against the rest of the line-up and what each tier costs, see our Claude API pricing breakdown and the Opus 4.6 review.
Testing Claude Opus 4.7
Claude Opus 4.7 has garnered strong feedback from our early-access testers. Below are some highlights and notes from our early testing of Opus 4.7:
Instruction following
Opus 4.7 is substantially better at following instructions. Interestingly, this means that prompts written for earlier models can sometimes now produce unexpected results: where previous models interpreted instructions loosely or skipped parts entirely, Opus 4.7 takes the instructions literally. Users should re-tune their prompts and harnesses accordingly.
Improved multimodal support
Opus 4.7 has better vision for high-resolution images: it can accept images up to 2,576 pixels on the long edge (~3.75 megapixels), more than three times as many as prior Claude models. This opens up a wealth of multimodal uses that depend on fine visual detail: computer-use agents reading dense screenshots, data extractions from complex diagrams, and work that needs pixel-perfect references.
Real-world work
Our internal testing showed Opus 4.7 to be a more effective finance analyst than Opus 4.6, producing rigorous analyses and models, more professional presentations, and tighter integration across tasks. Opus 4.7 is also state-of-the-art on GDPval-AA, a third-party evaluation of economically valuable knowledge work across finance, legal, and other domains.

GDPVal-AA Elo scores: Opus 4.7 leads the field in economically valuable knowledge work.
Memory
Opus 4.7 is better at using file system-based memory. It remembers important notes across long, multi-session work, and uses them to move on to new tasks that, as a result, need less up-front context.
Safety and Alignment
Overall, Opus 4.7 shows a similar safety profile to Opus 4.6: our evaluations show low rates of concerning behavior such as deception, sycophancy, and cooperation with misuse.
On some measures, such as honesty and resistance to malicious “prompt injection” attacks, Opus 4.7 is an improvement on Opus 4.6. Our alignment assessment concluded that the model is “largely well-aligned and trustworthy, though not fully ideal in its behavior”.

Automated behavioral audit: Opus 4.7 shows a modest improvement over 4.6, with Mythos Preview remaining the benchmark leader.
Also Launching Today
More effort control
Opus 4.7 introduces a new xhigh ("extra high") effort level between high and max, giving finer control over reasoning vs latency.
Task budgets (API)
Now in public beta, allowing developers to guide token spend so Claude can prioritize work across longer runs.
/ultrareview
New Claude Code command flags bugs and design issues that a careful human reviewer would catch.
In Claude Code, we’ve raised the default effort level to xhigh for all plans. We’ve also extended auto mode to Max users, where Claude makes decisions on your behalf for longer tasks with fewer interruptions.
A few details worth pinning down. /ultrareview is not an inline linting pass; it opens a dedicated review session, and Pro and Max subscribers get three of them free before usage-based billing applies. Task budgets, in public beta, let you tell the model roughly how much token spend a job is worth so it can prioritise across a long run rather than exhausting itself on the first subtask it encounters. And raising the Claude Code default to xhigh rather than leaving it at high is a deliberate statement about where Anthropic thinks the cost-quality trade-off now sits for engineering work: slower and more expensive per turn, but with fewer turns needed overall.
The practical upshot of auto mode reaching Max is that the supervision model changes shape. Rather than approving each step, you are approving a direction and reviewing a result, which puts considerably more weight on the review tooling that shipped alongside it. That is not a coincidence; the two features only make sense together.
Migrating from Opus 4.6 to Opus 4.7
Opus 4.7 is a direct upgrade to Opus 4.6, but two changes are worth planning for:
- Updated Tokenizer: Improved text processing may shift token mapping. Roughly 1.0–1.35× increase depending on the content type.
- Deeper Thinking: At higher effort levels, the model produces more output tokens to achieve its improved reliability on hard problems.
The net effect is favourable, token usage across all effort levels is improved on internal coding evaluations, but we recommend measuring the difference on real traffic using our migration guide.
The tokeniser change is the one that bites
A 1.0 to 1.35x shift in how the same text maps to tokens is not a rounding error at scale. At the top of that range, an unchanged workload costs a third more on input alone before the deeper reasoning at higher effort levels is accounted for. If you operate under a fixed monthly budget or a per-customer margin, re-run your cost model on a representative sample of real prompts before flipping the default, and do it on your actual traffic rather than on a synthetic benchmark, because the multiplier varies with content type.
Who Should Actually Upgrade
Stripping out the launch-day enthusiasm, the case for moving to Opus 4.7 is strong for some workloads and genuinely marginal for others. Here is how we would frame the decision.
Upgrade now
- Hard agentic coding. The SWE-bench Pro gain is the headline result, and it lands on exactly the class of task, long-running and multi-file, that people previously could not leave unsupervised.
- Anything visual. If your pipeline reads screenshots, scanned documents, charts or scientific figures, the resolution increase to roughly 3.75 megapixels plus the CharXiv jump make this a step change rather than an increment.
- Computer-use agents. OSWorld-Verified moves more than five points, and dense UI screenshots are precisely the case the vision upgrade was built for.
- Financial analysis. Finance Agent v1.1 puts 4.7 ahead of GPT-5.4 Pro, and Anthropic's internal testing pointed the same way.
Wait, or test carefully first
- Open-web research. BrowseComp regressed against 4.6. Benchmark your own retrieval workload rather than assuming the newer model wins.
- Cost-sensitive high-volume work. The tokeniser shift and the deeper reasoning at higher effort both push spend up. Per-token pricing has not changed, but tokens per task may have.
- Heavily-tuned legacy prompts. Because 4.7 follows instructions literally, prompt scaffolding written to compensate for older models, redundant checks, restated constraints, belt-and-braces verification steps, can now produce double passes and unnecessary verbosity. Simplifying usually helps more than tweaking.
- Security research workflows. The automatic cyber safeguards will block some legitimate red-team and vulnerability-research requests. Anthropic's Cyber Verification Programme is the route around this, and it takes time to clear.
A sensible migration sequence
Move in stages rather than flipping a global default. Start by pointing a read-only or low-stakes workload at 4.7 and comparing outputs side by side with 4.6 on the same inputs, not on fresh ones. Next, strip the defensive padding out of your three or four highest-volume prompts and re-measure; in our experience this is where most of the unexpected verbosity disappears. Only then adjust effort levels, and do it per-endpoint rather than globally, because the right setting for a classification call and the right setting for a refactor are not the same. Finally, re-baseline your cost per task on real traffic before you turn off the old model.
One last piece of context: Opus 4.7 was not the end of the line. Anthropic has since shipped Claude Opus 4.8, so if you are only now planning a migration it is worth reading both before committing engineering time to a version that has already been superseded.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.






