Anthropic’s Transparency Hub

A look at Anthropic's key processes, programs, and practices for responsible AI development.

Model Report

Last updated July 23, 2026

Select a model to see a summary that provides quick access to essential information about Claude models, condensing key details about the models' capabilities, safety evaluations, and deployment safeguards. We've distilled comprehensive technical assessments into accessible highlights to provide clear understanding of how the models function, what they can do, and how we're addressing potential risks.

Claude Sonnet 5 Summary Table

Model descriptionClaude Sonnet 5 can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.
Benchmarked CapabilitiesSee our Claude Sonnet 5 system card’s Section 8 on capabilities.
Acceptable UsesSee our Usage Policy
Release dateJune 2026
Access SurfacesClaude Sonnet 5 can be accessed through:
  • Claude.ai
  • Claude Code
  • The Anthropic API
  • Amazon Bedrock
  • Google Vertex AI
  • Microsoft Azure AI Foundry
Software Integration GuidanceSee our Developer Documentation
ModalitiesClaude Sonnet 5 can understand both text (including voice dictation) and image inputs, engaging in conversation, analysis, coding, and creative tasks. Claude can output text, including text-based artifacts, diagrams, and audio via text-to-speech.
Knowledge Cutoff DateClaude Sonnet 5 has a knowledge cutoff date of January 2026. This means the models’ knowledge base is most extensive and reliable on information and events up to January 2026.
Software and Hardware Used in DevelopmentCloud computing resources from Amazon Web Services and Google Cloud Platform, supported by development frameworks including PyTorch, JAX, and Triton.
Model architecture and training methodologyClaude Sonnet 5 was pretrained on large, diverse datasets to acquire language capabilities. After the pretraining process, Sonnet 5 underwent substantial post-training, with the goal of making it an effective assistant whose behavior aligns with the values described in Claude’s constitution.
Training DataClaude Sonnet 5 was trained on a proprietary mix of publicly available information from the Internet, public and private datasets, and synthetic data generated by other models. Throughout the training process we used several data cleaning and filtering methods, including deduplication and classification.
Testing Methods and ResultsBased on our assessments, we have decided to deploy Claude Sonnet 5 under CB-1 capabilities and autonomy threat model 1. See below for select safety evaluation summaries.

The following are summaries of key safety evaluations from our Claude 4.8 system card. Additional evaluations were conducted as part of our safety process; for our complete publicly reported evaluation results, please refer to the full system card.

User Wellbeing Summary

We run a suite of evaluations to understand how Claude responds in scenarios related to child safety and mental health. Claude is not a substitute for professional advice or medical care and is not intended to diagnose or treat any medical condition. We use these evaluations to understand how Claude performs in sensitive contexts and where we can make improvements. Claude Sonnet 5 is our most capable Sonnet model, but it is less capable compared to Opus- or Mythos-class models and its wellbeing-relevant results reflect that. For more in depth descriptions of the evaluations and their results, please see Claude Sonnet 5 system card.

  • Child Safety. Sonnet 5’s child safety behavior is comparable to or better than Claude Sonnet 4.6. Internal policy experts noted that Sonnet 5 tended to refuse clearly harmful requests more definitively than Sonnet 4.6. (Section 4.2)
  • Mental Health - Suicide and self harm. Internal policy experts found Sonnet 5’s handling of potential suicide and self-harm conversations conversations to be qualitatively comparable to Claude Sonnet 4.6. One of the clearest improvements was in Sonnet 5’s crisis response posture, providing resources sooner in a conversation compared to Sonnet 4.6. (Section 4.3.1)
    • On multi-turn suicide and self-harm testing on claude.ai, its 90% appropriate response rate exceeds Claude Opus 4.8 (85%) and Claude Sonnet 4.6 (82%), second only to Claude Fable 5 (96%).
  • Mental Health - Disordered eating. Claude Sonnet 5 performed similarly to Sonnet 4.6, with a slight improvement on responses to harmful requests on the claude.ai surface. (Section 4.3.2)
  • Misleading the user. Claude Sonnet 5 is broadly stronger than Sonnet 4.6 on measures related to deception and dishonesty, including active deception, sycophancy, sycophancy with users who appear dangerously delusional, hallucinating missing inputs, omitting important context, omitting reports of the model’s own bad actions, and falsely claiming to have completed tasks. (Section 6.4.3)
    • It is the strongest tested Claude model on the MASK measure of sycophantic dishonesty, with "the lowest lying rate of the models compared at 3.1%" — below Claude Mythos Preview (4.4%), Claude Opus 4.8 (6.1%), Claude Mythos 5 (8.6%), and Claude Sonnet 4.6 (13.3%) (Section 6.5.2).

Election Integrity

We evaluated Claude Sonnet 5 on an election integrity benchmark, which tests adherence to our Usage Policy across 300 policy-violating and 300 benign election-related requests grounded in patterns observed in real use. Results are reported for both the model on our API, without a system prompt (the standing instructions added to shape how the model behaves in a product), and with our claude.ai system prompt.

Sonnet 5 performed perfectly on the single-turn election integrity benchmark, reliably declining simple policy-violating requests without mistakenly refusing legitimate election-related requests.

We also evaluated Sonnet 5 qualitatively on our single-turn ambiguous context evaluation, as well as on a set of multi-turn test cases that are still being developed internally. In those evaluations, Sonnet 5 performed comparably to Sonnet 4.6 in identifying harmful requests and it was generally more nuanced than Sonnet 4.6 in how it handled ambiguous contexts. For example, Sonnet 5 appeared to be more adept at separating the harmful component of a request from the parts it could safely complete, resulting in offering more alternatives rather than declining outright. In one case, the model produced a requested voter-registration message but rewrote a line whose original phrasing implied it might already be too late to register, explaining that the wording would discourage the voter participation that the user was trying to encourage. Sonnet 5 was also at times more receptive than Sonnet 4.6 to a sympathetic framing (e.g., authorized red-teaming) of a request whose output could potentially be harmful regardless of intent, though in the election integrity cases we reviewed, the resulting outputs did not meaningfully increase a bad actor's ability to cause harm.

Alignment

Claude Sonnet 5 improves over Sonnet 4.6 on most of the positive character traits we test, including acting in the user’s interest and taking actively admirable actions. However, we see no improvement in creative mastery or warmth. Also, although our overrefusal metric above shows Sonnet 5 to be largely on par with Sonnet 4.6, Sonnet 5 appears to be actively worse on the broader “wet blanket” metric for dismissive or discouraging output. This is potentially linked to its improvement on sycophancy.

[Figure 6.4.6.A] Scores from our automated behavioral audit for the character metrics given below. Lower numbers represent a lower rate or severity of the measured behavior, with arrows indicating behaviors where higher (↑) or lower (↓) rates are clearly better. The y-axis is truncated below the maximum score of 10 in many cases. Reported scores are averaged across all approximately 2,900 investigations per target model (approximately 1,450 seed instructions sampled twice), with each investigation generally containing many individual conversations. Shown with 95% CI.

RSP Evaluations

Our Responsible Scaling Policy (RSP) evaluation process is designed to systematically assess our models' capabilities in areas where they could pose catastrophic risks before we release them. Sonnet 5 is less capable than our most capable model, Claude Mythos 5. Sonnet 5’s alignment risk, which is the risk that a model behaves in ways Anthropic did not intend, is very low. On automated AI research and development, Sonnet 5 performs below Claude Mythos 5 on every automated evaluation and therefore (like Mythos 5) does not cross the RSP capability threshold. On chemical and biological weapons, we conservatively treat Sonnet 5 in the same way we treated previous models such as Sonnet 4.6. We think it is capable of significantly helping individuals with basic technical backgrounds to produce (non-novel) weapons, and we deploy commensurate safeguards, including real-time classifiers to prevent harm. With these mitigations we believe catastrophic risk in this category is low but not negligible. For novel weapons development, the uplift it provides to threat actors who lack the expertise to develop such weapons is limited, with uncertainty about how much it may accelerate actors who already have that expertise.

Related content

RSP Updates

Overview of past capability and safeguard assessments, future plans, and other program updates.

Read more

Privacy Center

A central hub for information related to data privacy at Anthropic.

Read more

Trust center

This page acts as an overview to demonstrate our commitment to compliance and security.

Read more

Developer Documentation

Learn how to get started with the Anthropic API and Claude with our user guides, release notes, and system prompts.

Read more