The second risk report from Anthropic's Responsible Scaling Policy (RSP) landed like a polished stone in a still pond. Silence followed. The industry nodded, then moved on. But beneath the yield of procedural elegance lies the rot of unverifiable claims.
I have spent 21 years observing the intersection of code and trust. From the ICO gold rush to DeFi summer's structural flaws, I have learned one immutable truth: the most dangerous systems are those that look the most beautiful. Anthropic's RSP is a masterpiece of institutional design. But beauty is the mask; geometry is the bone. This report is a signal, not a data point. And as a cold dissector, I measure its depth, not its shine.
Context
Anthropic, the AI company behind the Claude model series, released its first RSP framework in May 2023. It was a pioneering document—a self-imposed set of safety thresholds modeled after biosafety levels (BSL). The framework introduced ASL levels (1 to 4) to map model capabilities to required safety measures. The second risk report, published in 2024 or early 2025, is the first substantive update. It claims to demonstrate that the RSP is a living, operational mechanism, not a static manifesto.
The report focuses on frontier risks: CBRN (chemical, biological, radiological, nuclear), cyberattack capabilities, autonomous replication, and self-improvement. These are the core dimensions for ASL-3 classification. The report does not disclose specific test results, but its existence implies that Anthropic's internal evaluation of Claude 3/3.5 models has reached or approached the ASL-3 threshold in at least some dimensions.
Hype is noise; structure is signal. The structure here is a self-assessment, self-published, and self-supervised loop. That is the first red flag.
Core: Systematic Teardown
Let us dissect the RSP v2.0 report as if it were a smart contract audit. I will apply the same forensic skepticism I used to uncover oracle manipulation vulnerabilities in DeFi protocols during the summer of 2020.
The Governance Token Illusion.
The RSP is governed entirely by Anthropic. The company decides what constitutes an ASL-3 capability, how to test it, and when to trigger restrictions. There is no external validator. This is functionally equivalent to a DAO where the team holds 100% of the voting power and the token is non-dividend equity. The community—be it researchers, regulators, or the public—has no means to verify the claims. The report may state that "Claude 3.5 Sonnet has been evaluated for CBRN knowledge diffusion," but without access to the test sets, the methodology, or the raw scores, the statement is a claim, not evidence.
Based on my experience auditing 45 whitepapers during the 2017 ICO mania, I learned that elegant language often masks logical fallacies. The RSP framework suffers from a similar fallacy: it conflates procedural rigor with substantive safety. A beautiful escalation process does not guarantee that the thresholds are set correctly. The gap between "we have a process" and "the process prevents harm" is where catastrophic failures hide.
The Oracle Problem.
In DeFi, the oracle is the weakest link. Latency, manipulation, and centralization of price feeds have caused billions in losses. In the RSP framework, the oracle is the internal safety assessment. Who feeds the data? Anthropic's red teams. Who validates the data? Anthropic's compliance team. Who audits the validator? No one. The joke that Chainlink solves decentralization with centralized nodes applies here: Anthropic solves safety with self-assessment.
The second report does not mention any independent audit. The RSP policy text from the first version stated an intention to invite third-party auditors, but the second report's silence on this point is telling. Silence is the loudest indicator of risk. I have seen this pattern before—a protocol promises transparency, but the audit trail evaporates when the market turns. In 2022, I compiled a dataset of on-chain withdrawals from three collapsed lending platforms. The pattern was consistent: beautiful dashboards, zero verifiable solvency proofs. The RSP report is a dashboard without a proof.
The Coverage Blind Spot.
The RSP focuses exclusively on catastrophic risks: CBRN, cyber, autonomous replication. It ignores the everyday social harms that are far more likely to affect billions of users: bias, discrimination, privacy violations, psychological manipulation. This is a deliberate choice. It is easier to build a framework for existential threats than to address the messy, costly, and politically charged issues of fairness and accountability.
This selective focus mirrors the behavior of certain DeFi protocols that market themselves as "decentralized" while maintaining admin keys that can drain user funds. The beautiful mask of "safety-first AI" conceals a gaping ethical void. Aesthetic perfection often hides ethical voids. The RSP report is a work of art, but it is incomplete. If Claude generates a racially biased loan denial or a privacy-violating conversation, the RSP framework offers no recourse. The company can point to the report and say, "We are safe on CBRN," while the real-world damage accumulates.
The Threshold Discretion Problem.
The ASL-3 threshold is inherently subjective. How much CBRN knowledge is too much? What constitutes a "significant" cyberattack capability? Anthropic holds the discretion, and the second report does not disclose the exact criteria. This is a classic principal-agent problem. The company has a commercial incentive to set the threshold high enough to avoid triggering restrictions, yet low enough to maintain credibility. The report's conclusion—that Claude 3.5 Sonnet has not triggered ASL-3—is precisely the result that would align with its commercial interests. Without independent verification, the conclusion is meaningless.
I recall a similar situation in 2021 when I analyzed an NFT collection's minting script. The royalty enforcement mechanism was opt-in, allowing wash trading to inflate volume. The team claimed the mechanism was "secure by default." My audit revealed the opposite. The RSP report's claim of "no ASL-3 trigger" may be similarly hollow. The code does not lie, but the contract can. The RSP is a contract between Anthropic and the public. The fine print is written by the company.
Contrarian: What the Bulls Got Right
Despite my skepticism, I must acknowledge what the RSP report gets right. It is a genuine institutional innovation. The ASL framework transforms abstract safety concerns into operational thresholds. This is more than most AI companies have done. The second report demonstrates that the RSP is not a one-time press release but a continuous process. That is valuable.
Moreover, the report's existence creates a precedent for regular safety disclosures. In a bear market for AI safety—where hype cycles dominate and accountability is scarce—Anthropic has built a structure. The bulls argue that this structure, even if imperfect, is a net positive because it forces the company to allocate resources to safety, to hire red teams, and to think about risk. They are right. The RSP is better than nothing.
But the contrarian in me must ask: is it good enough? The answer is no. The RSP is a beautiful bridge built on a single pillar. It will hold until the first real stress test. That stress test will come when a Claude model triggers ASL-3 and the company must choose between commercial deployment and safety restrictions. At that moment, the mask will slip. Until then, the report is a signal of intent, not a measure of safety.
Takeaway
The RSP second risk report is a milestone, but it is a milestone on a road that Anthropic has paved itself. The company has built a governance token for its own safety framework, and the token holders are anonymous. The public is asked to trust, not to verify. In a world where AI risks are global, this is not enough.
I do not follow the wave; I measure its depth. The depth of the RSP v2.0 report is shallow. It provides a framework, but not the data to validate it. The real test will come when the next Claude model crosses the ASL-3 threshold. Will Anthropic restrict deployment? Will it open the weights? Or will it find a way to reinterpret the threshold? The code does not lie, but the contract can.
Silence is the loudest indicator of risk. The second report's silence on independent audits, on specific test results, and on social risk coverage is deafening. The industry needs more than a beautiful mask. It needs bone-deep accountability.