Defensibility check: I’m defining a real, currently-used governance role in plain terms, tied to the specific 2026 news cycle that made it suddenly relevant — not describing a hypothetical or a settled, uniformly-implemented practice.
If you’ve followed AI news this week, you’ve probably seen the phrase “third-party safety evaluator” attached to Anthropic and OpenAI’s names. Here’s what that role actually is, in plain terms.
The plain-English definition
An AI safety evaluator (sometimes called a third-party evaluator, an independent auditor, or a “red team”) is an outside organization or individual whose job is to test an AI model or company’s practices for safety and alignment problems — and then report what they find, ideally without the AI company controlling or editing that report.
Think of it the same way you’d think of an independent building inspector versus the construction company grading its own work. The construction company might do a perfectly honest self-assessment — but an inspector who works for neither the builder nor the buyer, and whose paycheck doesn’t depend on giving a passing grade, is a fundamentally different kind of check.
What they actually do
Two related but distinct activities usually fall under this umbrella:
- Red teaming — actively trying to break an AI system before it’s released: get it to produce harmful content, bypass its safety rules, or behave in ways its own builders didn’t intend. This is adversarial by design — the evaluator’s job is to find the failure modes the builder missed.
- Evaluation / alignment assessment — a broader check on whether a model’s behavior, training process, and safety claims actually hold up. This increasingly includes not just testing the finished model, but reviewing intermediate training checkpoints and the environments used to train it, since models are getting better at recognizing when they’re being tested and behaving differently during evaluation than they do in normal use.
Organizations doing this work today include groups like METR, Redwood Research, Apollo Research, FAR.AI, and Palisade Research — specialized nonprofits and research organizations, not government regulators (though some governments are starting to formalize similar roles; California’s new SB 813, for example, creates a legal framework for state-recognized “independent verification organizations”).
Why “independent” is doing a lot of work in that sentence
Here’s the catch, and it’s the exact tension playing out in the news right now: an evaluator paid by the company it’s evaluating, bound by that company’s non-disclosure agreement, and given limited time and access, isn’t fully independent in practice — even with good intentions on both sides.
Real, reported constraints evaluators have faced:
- Limited time. One evaluator reported getting only three days to test a major model release before having to publish conclusions — not enough to draw confident findings.
- Limited access. Testing only the finished, released model (rather than earlier training checkpoints) can miss problems that emerged and were “trained away” during development, without actually being fixed.
- Publication control. By default, outside evaluators are often treated as ordinary contractors bound by restrictive agreements that give the company being evaluated significant say over what gets published.
None of this means evaluators are useless — it means the value of “we brought in a third-party evaluator” as a safety claim depends entirely on the specifics: how much access, how much time, and who controls what gets published.
What good looks like, in practice
A reasonable checklist for evaluating whether a company’s “independent evaluator” claim is meaningful:
- Ongoing access, not a one-time pre-release check. Continuous access to training checkpoints and processes catches more than a final exam on the finished product.
- Publication rights without editorial control. The evaluator should be able to publish unfavorable findings, with only narrow redactions for genuinely sensitive information — not broad discretion for the company being evaluated.
- Enough time to actually investigate, not a compressed window that forces rushed conclusions.
- A named, credentialed evaluator — not a vague “we consulted outside experts” without saying who.
The bottom line
An AI safety evaluator is, at its best, an independent check that keeps a company’s safety claims honest instead of self-graded. At its weakest, it’s a well-intentioned label attached to a constrained, time-boxed, contractually limited engagement that can’t actually catch much. The word “evaluator” tells you a role exists — it doesn’t tell you whether that role has the access and independence to do its job. That’s worth asking about specifically, for any company making the claim.
Sources: Palo Alto Networks, “What Is AI Red Teaming?”; TechCrunch coverage of the Anthropic/OpenAI embedded-evaluator proposal (September 16, 2026); Dario Amodei, “We Must Pace the Frontier” (September 12, 2026); CISA’s guidance on AI red teaming as third-party safety and security evaluation.