Anthropic is putting independent AI safety evaluators inside the company, giving them access comparable to employees as frontier models are trained, tested and deployed.
The first partner is Accenture, through its specialist AI business Faculty, and the arrangement could become a template for how AI labs try to make safety claims more verifiable.
The move is more consequential than another red-team contract. It changes where evaluation happens, who gets to observe the development process and how much independence can exist inside a company that still controls the models, funding and access.
Anthropic Wants Safety Evaluation To Happen From Inside
In its official announcement, Anthropic says Faculty will evaluate and red-team models, conduct alignment assessments and test model safeguards. Anthropic and Accenture each expect to invest at least $1 billion in building capacity over the next five years.
The important detail is access.
Embedded evaluators will work inside Anthropic with access comparable to an employee’s, allowing them to watch models develop during training, follow decisions about how systems are built and deployed, and speak directly with staff. That gives evaluators a view that an outside audit performed after release cannot easily provide.
Anthropic says the goal is not to transfer accountability.
The lab remains responsible for the safety of its models. The argument is that independent evaluators can make that responsibility more verifiable by identifying blind spots, reporting incidents and checking whether the company is following its own safety commitments.
That distinction matters because frontier AI evaluation is moving beyond a final pre-release test. Models are increasingly used as agents, connected to tools and exposed to real systems. A safety review that only examines a finished model may miss decisions made during training or deployment that shape how the system behaves in practice.
The Hard Part Is Independence, Not Access
Anthropic openly acknowledges that embedded evaluation is still being invented. There are no settled standards for what evaluators should be allowed to see, how they should communicate findings or who should fund the work over the long term. Anthropic says the partnership is non-exclusive and that it is also talking with METR and other nonprofit evaluators.
That creates the central tension. More access can produce better evidence, but working inside a lab can also make an evaluator dependent on the lab’s cooperation, funding and internal processes. The model needs enough proximity to observe the real work without becoming another layer of internal compliance.
The choice of Accenture is revealing. Faculty brings experience evaluating complex AI systems, while Accenture brings exposure to how businesses and governments actually deploy AI. That gives the partnership a practical, enterprise perspective rather than limiting safety evaluation to laboratory benchmarks.
It also signals that AI safety is becoming an operating discipline, not only a research specialty. The question is no longer just whether a model can pass a test. It is whether an organization can show how the model was tested, what evaluators could access, what they found and whether anyone outside the core product team can challenge the result.
Google DeepMind’s move to create a public forum for the AGI debate points to the same broader shift: AI companies are increasingly being judged not only by what their systems can do, but by how visibly they explain and govern the decisions around them.
Anthropic’s partnership is therefore an early experiment in institutional trust. If embedded evaluators can retain meaningful independence while working close to the models, the approach could become part of the standard infrastructure around frontier AI. If they cannot, the arrangement may look like internal assurance with a more reassuring name.