Anthropic taps Accenture's Faculty as its first embedded safety evaluator — $1B each over five years
Anthropic has named Accenture’s Faculty as the first evaluator team to be embedded inside its labs, with each side pledging at least $1 billion over five years. The experiment tests whether AI safety evaluation can survive being paid for by the company it grades.

Anthropic has named the first team that will work inside its own walls to check its models — and it is not one of the safety nonprofits the industry expected. On Friday, the company said evaluators from Faculty, the specialist AI business owned by consulting giant Accenture, will be embedded within Anthropic with access "comparable to an employee's." Each company says it expects to put at least $1 billion into the effort over the next five years.
The announcement is the first concrete deliverable from CEO Dario Amodei's September 12 essay calling for AI companies to slow the pace of capability gains. Embedded evaluators were his first proposed remedy — and this deal makes that idea real, weeks later. What makes it worth watching is the tension at the center of the design: the evaluators are being paid by the company they are meant to scrutinize.
What embedded evaluation actually means#
Most AI evaluation happens at arm's length: a lab hands a finished model to an outside tester, who probes it and reports back. The embedded model goes further — evaluators sit inside the development process, watching models take shape during training, following internal decisions about how they are built and released, and speaking directly to staff.
| Outside testing | Embedded evaluation | |
|---|---|---|
| Access | Finished model via API | Employee-comparable: training runs, internal decisions, direct staff contact |
| Timing | After (or near) release | During training, development, and deployment |
| Scope | Capability and safety probes | Red-teaming, alignment assessments, safeguard testing, incident reporting, verifying safety commitments |
The arrangement is non-exclusive in both directions: Anthropic says more evaluators will be named in the coming weeks, and Accenture will do similar work for other AI developers. The company also says it is in dialogue with METR and other nonprofit evaluators about own-funded pilots — a dialogue, not an agreement.
The money — and who pays#
The two companies each say they expect to invest at least $1 billion over five years building this capacity. Those are floors and expectations rather than a disclosed contract — no headcount, team size, or deliverable schedule has been published. Even so, the figure treats evaluation as a function to staff and fund, not a side project: a meaningful shift for a discipline that barely existed as a paid line item recently.
The awkward part is who writes the checks. Anthropic says it will fund Accenture's work directly in the near term, citing the "importance and urgency of this work," while arguing that long-term funding should come from pooled or government sources. An honest admission of a structural problem — independent scrutiny funded by the scrutinized is always one contract term away from looking compromised.
The independence tension#
The research community moved fast on this point. On the same day the deal was announced, more than 100 AI researchers — including Geoffrey Hinton — signed a public letter demanding that evaluators embedded in AI companies be "meaningfully independent." Their minimum bar: such organizations should not be owned or governed by frontier AI companies, should not have other significant commercial business with them, and should not accept payment contingent on their findings.
Accenture clears the first test and visibly strains the other two. It is not owned by Anthropic — but it is a consultancy with deep commercial relationships across the industry, and it is being paid directly by the lab it will grade. Rivals are reading the room: OpenAI CEO Sam Altman has said his company would give outside evaluators similar access, while Microsoft CEO Satya Nadella welcomed embedded evaluators but cautioned that oversight should not be controlled by a handful of entities. The caution is pointed — if the only firms capable of this work are the giant consultancies, "independent oversight" could quietly become a service the industry buys from a short list of vendors.
Safety watchers had expected Amodei's evaluator idea to be filled by research-focused nonprofits — METR, Redwood Research, Apollo Research — not a management consultancy. Faculty, which Accenture acquired earlier this year, brings serious safety credentials from public-sector work, but no track record of independent frontier-safety research. The bet is that enterprise-deployment experience transfers to safety evaluation. Plausible, not proven.
What to watch#
This is the first time anyone has tried to institutionalize evaluation inside a frontier lab at this scale, and the experiment's value is that it forces the independence debate out of the abstract. Three things to track: whether employee-level access surfaces problems that outside testing misses, whether the evaluators stay independent while their invoices are paid by the lab, and who else gets named in the coming weeks — a nonprofit piloting this on its own dime would change the conversation. If it works, every frontier lab will face pressure to open its doors. If it doesn't, the industry will have spent $2 billion to learn exactly where the model breaks — which is, fittingly, the point of evaluation.