Growth Cab Apply to GC
Blog/AI STRATEGY
AI STRATEGY · September 17, 2026 · 6 MIN READ

How to Evaluate AI Safety Claims From Frontier Labs

A practical six-step method for checking frontier AI safety claims against policies, evaluations, independent evidence, governance and explicit unknowns.

Federico DonatoneBy Federico Donatone · Founder, Growth Cab
How to Evaluate AI Safety Claims From Frontier Labs

To evaluate AI safety claims, copy the exact claim, identify who made it, inspect the public evidence, compare it with independent findings, check who can stop deployment, and write down what remains unknown. A resignation can be important evidence. It cannot prove a probability by itself. A company policy can be useful evidence. It cannot prove compliance by itself.

That is the practical answer behind my LinkedIn post about Jacob Coxon leaving Anthropic. The post reached 33,088 impressions, 463 reactions and 179 comments. It summarized his warning that AI labs were racing toward self-improving systems without adequate safeguards. His experience deserves attention. The size of the risk still requires evidence beyond one person or one viral thread.

Federico Donatonein
Federico Donatone
Founder, Growth Cab · This article started as a LinkedIn post

“An Anthropic researcher just quit. His reasons are wild. Here are the 7 things he said on his way out:”

33,088IMPRESSIONS
463REACTIONS
179COMMENTS
Read the original post →

How to Evaluate AI Safety Claims in Six Steps

  1. Copy the exact sentence and separate facts from forecasts.
  2. Identify the speaker's access, incentives and limits.
  3. Find the lab's dated policy, model card and risk report.
  4. Compare those documents with independent evaluations and incidents.
  5. Check thresholds, decision rights, audits and stop conditions.
  6. Record what the evidence supports, contradicts and leaves unknown.

The method takes minutes for a headline and longer for a serious buying or policy decision. The point is to keep three things separate: what happened, what somebody believes may happen and what controls exist today. Most arguments collapse those layers into one emotional answer.

1. Copy the Claim Before You Judge It

Start with the exact words, date and source. Coxon's resignation is a fact reported across several outlets and tied to his own public statement. His warning about catastrophic outcomes is a forecast. His description of private concern inside labs is testimony from someone with relevant access. Those are three different evidence types.

Write each sentence on a separate line. Label it observed event, company statement, personal estimate, technical result or policy proposal. Then ask what evidence could support that specific line. This prevents a verified employment change from silently turning into proof that every forecast attached to it is correct.

Axios reported the resignation, Coxon's stated reasoning and the equity he left before vesting. That supports the departure and his account of his motivation. It does not measure the probability of a future loss-of-control event.

2. Check Access and Incentives

An insider may know details outsiders cannot see. That access increases the value of testimony about internal priorities, conversations and working conditions. It does not make every technical forecast automatically correct. The person may also lack visibility into governance, security or evaluations handled by other teams.

Incentives matter in both directions. Someone who leaves compensation behind weakens the simple claim that the warning was designed to raise a personal stake. A company defending its safety process has commercial and reputational incentives too. Incentives help you frame evidence. They do not replace evidence.

3. Read the Lab's Current Safety Contract

A frontier lab should publish more than principles. Look for capability thresholds, required safeguards, evaluation methods, incident reporting, responsible owners and a process for delaying deployment. Dates and version history matter because a polished framework can change while a model is being trained.

Anthropic publishes a Responsible Scaling Policy, a Frontier Compliance Framework, model documentation and periodic risk reports. Its Transparency Hub says safeguards rise with identified risks and lists internal channels for escalating safety concerns. These documents show the company's stated process. They still need evidence of execution.

Anthropic's Transparency Hub describes its risk categories, reporting channels, model documentation and external collaboration.

4. Compare the Claim With Independent Evidence

Independent evidence should test the mechanism behind the warning. For loss of control, ask whether current systems can evade oversight, sustain long plans, acquire resources, resist intervention and combine those abilities reliably outside a laboratory. A dramatic demo of one behavior cannot establish the full chain.

The 2026 International AI Safety Report says current systems show early signs of relevant capabilities but do not yet reach the level required for loss of control. It also says experts disagree widely about likelihood and severity. That is a useful boundary: the risk deserves preparation, while confidence about timing remains limited.

The independent 2026 International AI Safety Report executive summary separates current capability evidence from uncertain future loss-of-control scenarios.

5. Inspect Governance Alongside Evaluations

A benchmark can tell you whether a model crossed a capability threshold. Governance tells you what happens next. Who sees the result? Who can block release? Can a commercial leader override the safety recommendation? Does an outside reviewer see enough confidential evidence to challenge the lab's conclusion?

The strongest setup links every threshold to a required action and a named decision owner. It preserves test artifacts, documents exceptions and publishes enough information for outsiders to understand the conclusion. A framework that leaves every hard choice to private discretion gives readers little reason to trust the final label.

A 2026 paper on frontier AI auditing argues that public transparency alone cannot verify every safety claim and proposes deeper third-party assessment with secure access to non-public evidence.

6. Write the Unknowns Beside the Conclusion

Finish with three columns: supported, contradicted and unknown. Supported might include the resignation, the lab's published thresholds and the current report's description of early warning signs. Contradicted might include any claim that today's systems already possess the complete capability chain required for loss of control.

AI FRONTIER
Get one useful AI play every Thursday
The AI changes that matter, Federico's direct read and one practical play, plus the free 10-page Operator Pack.
Free · under five minutes · unsubscribe anytime

Unknowns are where honest decisions live. We do not know every internal evaluation, every disagreement or how quickly several capabilities may combine. We also do not know whether a proposed global pause could be verified across countries and private labs. Naming those gaps prevents confidence from outrunning the record.

My Frontier AI Safety Claim Scorecard

  1. Claim: exact sentence, date and source.
  2. Access: what the speaker could directly observe.
  3. Evidence: policy, evaluation, incident or independent test.
  4. Mechanism: the steps required for the predicted harm.
  5. Governance: thresholds, owner, audit and stop authority.
  6. Unknowns: missing data, disputed assumptions and timing.
  7. Decision: the reversible action justified today.

Score the evidence instead of the speaker. A firsthand account can be strong on internal culture and weak on global probability. A lab policy can be strong on declared controls and weak on implementation. An independent report can be strong on the public evidence base and still miss confidential results.

What This Means for a Business Using AI

Most companies do not decide whether frontier research should stop. They decide where a model may operate. Use the same method on your own systems. Name the capability, require a representative test, define the harm, set permissions, preserve logs and give one person authority to stop the workflow when the evidence changes.

High uncertainty should narrow permissions. Let the system research, summarize or draft while a person approves external messages, payments, production changes and sensitive-data transfers. Increase autonomy only when repeated evidence shows the controls work under real conditions. Reversibility is a better default than confidence theatre.

For capability headlines and benchmark claims, use my five-question AGI claim test. It covers the model, tools, retries, benchmark and limits behind a performance announcement.

The Useful Conclusion

Coxon's resignation is meaningful because relevant insiders rarely abandon a frontier lab and publicly challenge the race they helped build. It should raise the standard of proof demanded from the labs. It should not force a precise extinction forecast from evidence that cannot support one.

The disciplined response is simple. Take the warning seriously. Read the lab's current contract. Compare it with independent evidence. Check who can stop deployment. Keep the unknowns visible. Then make the smallest reversible decision the evidence supports today.

Want a GTM engine that runs like this?

Growth Cab is the #1 GTM & sales advisory in the US & Europe. We build the outbound, LinkedIn, and closing systems behind these playbooks for founders selling high-ACV deals.

Apply to GC ← All articles
AI FRONTIER

Turn this week's AI noise
into one useful move

Every Thursday: the signal, Federico's direct view and one practical play. Join free and get the 10-page AI Frontier Operator Pack.

Free · under five minutes · unsubscribe anytime