AI Tools Review
Claude 3.5 Sonnet System Card Deep Dive: Analyzing Safety, Computer Use, and Autonomous Engineering

Insights

Claude 3.5 Sonnet System Card Deep Dive: Analyzing Safety, Computer Use, and Autonomous Engineering

AI Tools Review Editorial TeamApril 11, 2026

    1. Introduction to the 3.5 Sonnet Paradigm

    When Anthropic published the system card for the upgraded Claude 3.5 Sonnet, it wasn't just a routine technical update, it fundamentally rewrote the boundaries of autonomous software engineering and "agentic" capabilities in consumer models.

    The Claude 3.5 Sonnet System Card offers a rare, granular look into how Anthropic's alignment teams evaluated a model that possessed, for the first time, the ability to control standard computer operating systems via the "Computer Use" API. Unlike previous models that were effectively "brains in jars," Sonnet 3.5 was a model given hands.

    This deep dive pulls out the critical insights from that dense technical document, looking closely at how Anthropic ensured the model couldn't be manipulated into executing catastrophic cybersecurity attacks while maintaining industry-leading coding benchmarks.

    A point of housekeeping first, because it trips up a lot of readers searching for this document. There is no standalone file called the "Claude 3.5 Sonnet System Card". What Anthropic actually published on 22/10/2024 is a 14-page PDF titled Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnet, which sits on top of the original Claude 3 Model Card from March 2024. The addendum is the authoritative source for every figure quoted in this review, and it covers both models released that day, which is why the upgraded Sonnet numbers are always presented alongside a Haiku column. If you want the sibling analysis, we have covered the smaller model separately in our Claude 3.5 Haiku system card review.

    The addendum is unusually candid by the standards of the period. It records the model's knowledge cutoff as April 2024, the same as the original 3.5 Sonnet, and it explicitly excludes OpenAI's o1 family from the comparison tables on the grounds that reasoning models "depend on extensive pre-response computation time, unlike the typical models", making like-for-like comparison unfair. That kind of methodological footnote is rare in marketing-adjacent documents, and it is a useful signal of how the safety team wanted the numbers read.

    2. The "Computer Use" Controversy

    The most widely discussed element of the 3.5 Sonnet card was the inclusion of Computer Use beta capabilities. Anthropic explicitly noted that allowing an AI to autonomously drag a cursor, type keystrokes, and read screen states introduced profound new threat vectors.

    The capability itself is narrower than the headlines suggested. As the addendum describes it, the model interprets screenshots of a graphical user interface and generates appropriate tool calls: moving the cursor, clicking, typing and reading what comes back. Anthropic deliberately fed the model screenshots only during evaluation, even though the OSWorld benchmark permits richer inputs such as an accessibility tree rendered as text. That is a harder setting, and the scores reflect it.

    OSWorld Category15 Steps50 StepsHuman
    OS tasks54.2%41.7%75.00%
    Office7.7%17.9%71.79%
    Daily16.7%24.4%70.51%
    Professional24.5%40.8%73.47%
    Workflow7.9%10.9%73.27%
    Overall14.9%22.0%72.36%

    Fourteen point nine per cent was a state-of-the-art result at the time, and Anthropic showed that simply raising the interaction budget from the standard fifteen steps to fifty lifted the overall score to 22 per cent. The interesting wrinkle is that more steps did not help uniformly: the OS category actually got worse, falling from 54.2 per cent to 41.7 per cent, which suggests the model could talk itself out of a correct answer when given room to keep fiddling. Office and Professional tasks, by contrast, roughly doubled.

    "While this represents a significant advancement over previous results, it remains well below human performance of 72.36%, indicating substantial room for future improvement in this domain."

    Anthropic, Model Card Addendum, Section 2.1

    The mitigations are more prosaic than the threat narrative implies. Anthropic ran dedicated Trust & Safety red-teaming for computer use and named four abuse vectors it considered plausible but not imminent: scaled account creation, scaled content distribution, age assurance bypass and abusive form filling. None of those are cinematic. All of them are the sort of thing a spam operation would try first. In response the company built new classifiers specifically to identify and evaluate computer-use activity for Usage Policy violations, and shipped the capability as a public beta rather than a general-availability feature.

    Crucially, the reference implementation runs in the customer's own environment. Anthropic's published guidance for computer use is to run it inside a dedicated virtual machine, limit access to sensitive data, restrict internet access to the domains the task genuinely needs, and keep a human in the loop for anything consequential. The logic lives on Anthropic's clusters; the mouse pointer lives on your machine, which is precisely where the liability sits.

    3. ASL-2 and CBRN Guardrails

    Anthropic operates under a rigorous Responsible Scaling Policy (RSP) categorising models into AI Safety Levels (ASL). According to the system card, 3.5 Sonnet was classified as an ASL-2 model.

    Crucially, rigorous red-teaming was executed across three frontier-risk domains. Anthropic does not publish per-domain pass rates for these evaluations in the addendum, and it is worth being blunt about that: any table circulating online that claims precise "pre-intervention" and "post-intervention" percentages for CBRN or exploit execution on this model is not drawn from the published document. What Anthropic did publish is the scope of the testing and the conclusion it reached.

    DomainWhat Anthropic TestedPublished Finding
    CBRNAutomated knowledge tests, new baselines, manual red-teaming, and whether the model uplifts non-expert performance on CBRN tasksImproved on both knowledge retrieval and skills assessments, but did not reach the ASL-3 threshold
    CybersecurityVulnerability discovery and exploit capability across pwn, reverse engineering, cryptography, web and network capture-the-flag challengesImproved at solving certain classes of CTF challenge; no ASL-3 safeguards required
    AutonomySoftware-engineering tasks such as submitting a pull request that satisfies test requirements, treated as a precursor to autonomous capabilityIncreased proficiency, but below the autonomy threshold that would trigger ASL-3

    The summary judgement in Section 3.2.1 is a single sentence: the model "did not demonstrate capabilities requiring ASL-3 safeguards and security in any domain", though Anthropic noted stronger capabilities across all three. Anthropic explicitly attributes part of that increase to better elicitation and testing technique rather than raw capability alone, which is an important caveat when comparing scores across model generations.

    Anthropic also asked a narrower question that turns out to be the more interesting one: does computer use itself change the frontier-risk picture? The conclusion was no, for three separate reasons. On CBRN, the ability to drive a GUI does not help without the underlying knowledge and skills, which the model did not have at a dangerous level. On cybersecurity, computer use "likely does not enable significant new capabilities beyond what can already be achieved with existing tools" — it may lower the barrier for a novice to run a script through a graphical interface, but Anthropic judged that such an actor would lack the surrounding knowledge to be an extreme threat. On autonomy, visual computer use is "not on the critical path" to the sort of software-engineering ability that autonomy evaluations actually measure.

    Two further details are worth flagging because they rarely make it into secondhand summaries. First, this was not a purely internal exercise: the upgraded 3.5 Sonnet went through joint pre-deployment testing with both the US AI Safety Institute and the UK AI Safety Institute, and Anthropic separately commissioned METR to run an independent assessment. Second, the Trust & Safety red-team covered fourteen policy areas across six languages — English, Arabic, Spanish, Hindi, Tagalog and Chinese — with particular attention to elections integrity, child safety, cyber attacks, hate and discrimination, and violent extremism.

    4. Autonomous Engineering Benchmarks

    While the safety metrics are important, the addendum also formally documented 3.5 Sonnet's blistering pace on agentic coding and tool-use benchmarks. This section is where the document's real significance lies, because it was the first time Anthropic put SWE-bench Verified and TAU-bench into its own official evaluation suite.

    Upon release, the upgraded Claude 3.5 Sonnet solved 49.0% of the SWE-bench Verified dataset, up from 33.4% for the original 3.5 Sonnet, 22.2% for Claude 3 Opus and 7.2% for Claude 3 Haiku. Anthropic footnotes the comparison point honestly: the published leaderboard state of the art on 22/10/2024 was 45.2%. The headline claim was therefore a genuine but modest lead, not the order-of-magnitude jump the coverage implied. The company also noted that "dedicated scaffolding and prompting can further improve the results", which is a quiet acknowledgement that harness quality, not just model weights, drives these figures.

    What makes SWE-bench Verified meaningful is the loop it forces. The model is handed a real GitHub issue from a popular open-source Python repository. It writes a script to reproduce the bug, searches and reads the repository, edits source files, runs its script, and keeps iterating until it decides it is finished. Only then is the result graded against the actual unit tests from the pull request that resolved the issue — tests the model never sees. Half of that, unassisted, was a genuine threshold moment.

    Evaluation3.5 Sonnet (New)3.5 Haiku3.5 Sonnet (Original)Claude 3 Opus
    SWE-bench Verified49.0%40.6%33.4%22.2%
    TAU-bench (retail)69.2%51.0%62.6%45.1%
    TAU-bench (airline)46.0%22.8%36.0%34.5%
    Internal agentic coding eval78%74%64%38%

    TAU-bench deserves more attention than it usually gets. It simulates customer-service scenarios in which the model must interact with both a simulated user and a set of APIs whilst obeying a policy document. Its scoring metric is unusual: rather than pass@k, where one of several attempts merely has to succeed, it reports pass^k, the fraction of problems where all k samples are correct. That is a reliability measure rather than a capability measure, and it is far less forgiving. The upgraded Sonnet led across k values from one to eight in both domains, meaning its advantage was consistency, not luck.

    Note the gap between the retail and airline domains: 69.2% against 46.0%. Airline tasks involve more constrained policies and more chances to violate one. That roughly twenty-three-point spread is a reasonable proxy for how much harder regulated, rule-dense customer workflows are than general commerce, and it is a number worth keeping in mind before promising an agent will handle your compliance-heavy queue.

    These numbers marked the point at which AI models stopped being advanced auto-correct and started behaving like junior engineers capable of parsing vast repositories, establishing context, and shipping verifiable patches. The direct line from this addendum runs through Claude Code and on to collaborative agent workflows like Claude Cowork.

    5. Refusals, Red-Teaming and Prompt Injection

    One of the addendum's more honest sections concerns refusals, and it contains a genuine regression that Anthropic did not bury. Measured on the WildChat toxic prompt set, the upgraded 3.5 Sonnet correctly refused 89.2% of harmful prompts, down from 96.4% for the original 3.5 Sonnet. That is a meaningful drop in the "correct refusal" column.

    The trade-off appears in the other direction. Incorrect refusals — the model declining a perfectly harmless request — fell from 11.0% to 5.3% on the non-toxic WildChat set. On XSTest, a suite designed specifically to catch exaggerated safety behaviour, the upgraded model scored 4.3% incorrect refusals against 8.3% for Claude 3 Opus and a startling 36.6% for Claude 3 Sonnet. Anyone who used the early Claude 3 models and found them maddeningly prim now has a number that explains the feeling.

    Read together, those two movements describe a deliberate recalibration: Anthropic traded some strictness on genuinely harmful prompts for a large reduction in false alarms on benign ones. Whether that is the right trade depends entirely on your deployment. A consumer assistant benefits enormously; a moderation pipeline might not.

    The Trust & Safety team's own summary is that overall harm rates for the upgraded model were "similar to, but slightly improved over" the original, with no increased risk of real-world harm identified. It did, however, name a persistent weakness: both models "struggled with nuanced requests or those framed as fiction, roleplaying, or artistic content". Fictional framing as a jailbreak surface was, and remains, an unsolved problem.

    Prompt injection gets its own short section. Anthropic built internal test sets of injection attacks and trained specifically on adversarial interactions so the model would better recognise a user trying to override its system prompt. The document stops well short of claiming the problem is solved, which is why the computer-use guidance immediately follows it. If you are building on this capability, the addendum's own advice — dedicated VM, restricted data, domain allowlist, human in the loop — is still the correct baseline in 2026.

    6. Reasoning, Maths and Vision Results

    Outside the agentic headline figures, the upgraded model posted solid but incremental gains on the standard academic suite. GPQA Diamond, the graduate-level science question set, rose from 59.4% to 65.0%. MMLU Pro climbed from 75.1% to 78.0%. The MATH benchmark moved from 71.1% to 78.3%, and HumanEval from 92.0% to 93.7% — a benchmark already so saturated that the movement is close to noise.

    The addendum introduced two benchmarks to Anthropic's suite for the first time. IFEval, which measures precise instruction-following, came in at 90.2%. AIME 2024, the American Invitational Mathematics Examination, produced the most sobering number in the entire document: 16.0% zero-shot with chain of thought, rising to 27.6% with majority voting over 64 samples. Set against the 78.3% on MATH, that gap is the clearest evidence in the addendum that competition-level mathematics was still genuinely out of reach for a non-reasoning model in late 2024.

    On vision, the upgraded Sonnet claimed state-of-the-art results on MathVista (70.7%, up from 67.7%), AI2D science diagrams (95.3%) and MMMU validation (70.4%). ChartQA held flat at 90.8%, and DocVQA actually slipped slightly, from 95.2% to 94.2%. Claude 3.5 Haiku launched as a text-only model, so no multimodal figures were reported for it.

    Anthropic also ran human preference evaluations, using the original 3.5 Sonnet as a fixed 50% baseline. The upgraded model won 61% of head-to-heads on document analysis, 58% on creative writing, 57% on vision instruction-following, 55% on honesty, 52% on coding and 51% on precise instruction-following. Two categories are more telling than the wins: multilingual came in at 48% and harmlessness at 48%, both marginally below the baseline. The model that got better at agentic work did not get better at everything, and the addendum says so in a chart rather than hiding it.

    7. Why the Addendum Still Matters in 2026

    Read from here, the document's numbers look quaint. Frontier models now score in the high seventies and beyond on SWE-bench Verified, and computer use has moved from a screenshot-driven beta to a mainstream production capability. Yet the addendum remains one of the most useful artefacts in the recent history of AI safety documentation, for three reasons.

    First, it established the template. The structure — capability benchmarks, then Trust & Safety red-teaming, then Responsible Scaling Policy frontier-risk evaluations, then an explicit ASL determination — is the shape almost every subsequent Anthropic system card has followed. If you want to learn how to read one of these documents critically, this is the cleanest place to start, because it is short enough to read end to end in half an hour.

    Second, it captured a genuine capability discontinuity in the specific domain that mattered most. Anthropic's autonomy evaluations treat software-engineering ability as a precursor to autonomous capability, and this was the model where that ability first became commercially interesting. Everything from agentic IDEs to the industry-wide vulnerability-hunting work of Project Glasswing descends from the assumption, first made defensible here, that a model can be trusted to iterate against a test suite without supervision.

    Third, it is a useful corrective against benchmark inflation. The addendum reports a regression in correct refusals, a below-baseline harmlessness win rate, a 16% AIME score and a computer-use figure less than a quarter of human performance, all in the same document as a state-of-the-art claim. Frontier labs do not always publish numbers that make them look worse. When they do, it is worth noticing — and worth being sceptical of any summary that quotes only the flattering half.

    Review Methodology

    Every figure in this review is taken from Anthropic's official Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnet (October 2024), supplemented by Anthropic's launch announcement for the same models. Where the addendum does not publish a number, we have described the finding qualitatively rather than estimating it.

    1. Anthropic, Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnet (PDF, October 2024)
    2. Anthropic, Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku (22/10/2024)

    Frequently Asked Questions

    What is the Claude 3.5 Sonnet System Card?
    Anthropic's official technical document outlining the safety evaluations, alignment scoring, and capability benchmarks of the Claude 3.5 Sonnet model.
    Did Claude 3.5 Sonnet introduce Computer Use?
    Yes, an updated version of the Claude 3.5 Sonnet model introduced the controversial 'Computer Use' API, allowing the model to autonomously move the cursor and click around a virtual desktop.
    Is Claude 3.5 Sonnet dangerous?
    According to the System Card, Claude 3.5 Sonnet complies with ASL-2 (AI Safety Level 2) standards, meaning it does not present a catastrophic CBRN threat.
    How well did Claude 3.5 Sonnet perform on real-world coding benchmarks?
    The system card reports that the upgraded Claude 3.5 Sonnet solved 49% of the SWE-bench Verified dataset - meaning it autonomously resolved nearly half of real-world pull requests from complex open-source repositories with zero human intervention.
    How effective were Claude 3.5 Sonnet's CBRN safety mitigations?
    Anthropic does not publish per-domain pass rates for these evaluations in the Claude 3 Model Card October Addendum. Any table circulating online quoting precise 'pre-intervention' and 'post-intervention' percentages for CBRN or exploit execution on this model is not drawn from the published document. What Anthropic did publish is the scope of the testing - automated knowledge tests, new baselines, manual red-teaming, and whether the model uplifts non-expert performance on CBRN tasks - and its conclusion that the model remained within the ASL-2 standard. Anthropic separately assessed whether computer use itself changed the frontier-risk picture and concluded it did not, on the grounds that driving a GUI does not supply the underlying knowledge a dangerous actor would still need.

    Explore more AI tool comparisons

    In-depth reviews, benchmarks and guides to help you choose the right AI tools.

    Browse all reviews
    AI Tools Review Editorial Team

    AI Tools Review Editorial Team Expert verified

    Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.