METAL

Anthropic partners with Accenture on embedded evaluation

Anthropic will work with Accenture on independent evaluation of frontier AI. Each company expects to spend at least $1 billion over five years, and the evaluators will work inside the company with access comparable to an employee's.

Anthropic partners with Accenture on embedded evaluation

Image: METAL

Summary

  • Anthropic announced an embedded evaluation partnership with Accenture on September 18, to be led by Faculty, Accenture's AI business.
  • Each company expects to spend at least $1 billion over five years, and Anthropic is funding this engagement directly.
  • Embedded evaluators get access comparable to an employee's, covering models in training, deployment decisions, and conversations with staff.

The promise to bring evaluators inside the company rather than keep them outside has become a contract. Anthropic said on September 18 that it will work with Accenture on independent evaluation of frontier AI. The company wrote that the partnership is an important step toward the commitment made in CEO Dario Amodei's essay We Must Pace the Frontier, namely to embed evaluators within Anthropic.

The work will be led by Faculty, Accenture's specialist AI business. The scope is set out in three parts: evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards. Anthropic explained that Accenture helps businesses and governments deploy AI across many industries, and that its understanding of how enterprises actually use AI informs its safety approach.

The size of the money describes the nature of the partnership. The two companies said they each expect to invest at least $1 billion in building capacity in this area over the next five years. That is the scale of standing evaluation up as a business line rather than a research project.

The core difference from today's external evaluators is access. The announcement says embedded evaluators will work inside AI companies with access comparable to an employee's. That access lets them watch models take shape in training, follow the decisions that govern how those models are built and deployed, and speak directly to employees. It is a position from which they can assess how a company operates, verify that it is keeping its safety commitments, and identify blind spots, and it also covers reporting incidents and giving the public a more informed account of benefits and risks.

Where responsibility sits was stated plainly. Anthropic wrote that "independent embedded evaluators do not reduce our accountability, but help to make it more verifiable," and added that the safety of its models remains its own responsibility.

The same post says a great deal is still missing. There are no standards yet for what information embedded evaluators should have access to or how they should report what they find, and no settled system for funding independent evaluation. The company said that long-term it believes funding should come from pooled or government sources, but since neither exists today it plans to work with different evaluators under different funding arrangements.

So for this engagement Anthropic funds Accenture's work directly. That puts the party being evaluated in the position of paying the party doing the evaluating, and separately the company wrote that it is in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using their own funding. METAL has reported on Anthropic reclassifying four Claude incidents as misalignment while signaling an independent review by METR.

The 19-page Advanced AI Framework, whose page count METAL counted directly after downloading the original PDF, already sets out the risk in this structure. Published in June, the document writes that "self-assessment is not enough" and gives a separate heading to the risk of what it calls evaluator shopping, where companies seek out whichever evaluators will ask the least of them and be most generous in their assessments. Its prescription is for government agencies to rate evaluators against predefined criteria such as the rigor of public reasoning, and to randomly assign highly rated evaluators, particularly in high-stakes cases.

The same document sets out what evaluators should receive. Within six months of the enactment of the regulation, developers would have to regularly engage at least one qualified independent evaluator, and that evaluator would get an unredacted version of the most recent risk report and system cards along with access to the most capable models. Evaluators would be bound by obligations not to copy, retain, or disclose confidential information, but beyond that should generally not be restricted in what they can publish, including concerns about the risk report or the developer's conduct.

The shape of the contract follows that document too. The partnership is non-exclusive, and Anthropic wrote that it will announce other evaluators in the coming weeks. Accenture will also work with other AI developers in similar capacities. It is the passage where the company says it expects frontier labs to work with several organizations at once.

What embedded evaluation actually catches will turn on what the company opens up. METAL has reported that Anthropic shipped one model without pre-release checks by the UK AISI. This commitment, that resident evaluators hold employee-level access, means such a judgment would no longer be the company's alone, and whether the commitment holds has to be confirmed in a place that has neither standards nor funding yet.

Comments