AI systems are increasingly being asked not only to generate outputs, but to use tools, change persistent state, interact with external systems, and evaluate whether their own work succeeded. That creates a verification problem that is easy to underestimate.
A common response is to separate generation from evaluation. One model performs the task, another reviews the result, and the second model’s verdict becomes an additional trust signal. That can improve evaluation, expose disagreements, and provide additional evidence. It does not by itself establish independent verification.
This is not an argument against AI-based evaluation. It is an argument against making that evaluation sufficient authority for trusting the system being evaluated.
The harder question is whether the system being evaluated can influence the evidence, policy, or authority that determines what counts as an acceptable result.
That distinction became concrete during recent work on ODIN, Buckner Intelligence Group’s control and assurance system. A candidate passed more than 100 focused tests and still failed a separately run internal review designed to be independent of the candidate’s execution and trust authority. Here, independent refers to architectural separation, not third-party validation.
The problem was not that there was no verifier. The problem was that the candidate retained too much influence over the mechanism that made the verifier authoritative.
A different verifier identity existed. That was not enough.
Evaluation and verification are different problems
OpenAI’s May 2026 guidance on third-party evaluations describes how agent evaluations increasingly depend on the surrounding environment and harness. Models can now use tools, maintain state, recover from mistakes, and act across longer workflows, which means the evaluation setup itself can materially change the result. OpenAI also identifies reward hacking as a validity problem when a system receives credit by exploiting the task, scorer, prompt, or harness rather than demonstrating the behavior being measured. [1]
That raises one question: did the evaluation measure what we thought it measured?
Verification adds another: who or what is authorized to decide that the evidence is sufficient?
A test harness may accurately score an outcome while failing to constrain the path that produced it. A verifier may correctly inspect a trajectory while operating inside a trust structure that the subject can influence. A signed verdict may be authentic while remaining weak evidence if the system being evaluated controls which signer counts as trusted.
Outcome success can conceal execution failure
Microsoft Research’s 2026 work on computer-use-agent verification separates process from outcome because those dimensions can diverge. An agent can execute a task appropriately and still fail because of an environmental condition, or reach the desired outcome through a path that should not receive the same credit. [2]
The same distinction matters in operational systems.
An agent can create the correct file while writing somewhere it was not authorized to write. It can produce the desired database state after bypassing an approval boundary. It can complete the requested task while also creating an unsolicited side effect.
Outcome success is not execution correctness.
If verification asks only whether the desired final state exists, unauthorized or ambiguous execution can disappear behind the result. The system may know that the task worked without establishing whether it worked through an admissible path.
That is not merely an observability problem. It becomes a control problem when the system is allowed to act on the verdict.
A separate verifier can still share the same authority
Suppose Agent A performs the work and Agent B evaluates it.
The second evaluator can be useful, but its existence does not answer the more important architectural questions.
Who determines which verifier is trusted? Who controls the policy it evaluates against? Who determines which evidence it receives? Can the system being evaluated alter the trust configuration that gives the verifier’s decision force?
If those powers remain inside the same mutable authority domain, the architecture has separated roles without necessarily separating control.
That is what the recent ODIN case exposed.
The candidate under review had separate generator and verifier identities, and its focused test suite was green. Separate internal review still found that the candidate retained influence over the verifier-trust mechanism. Identity separation therefore did not establish authority separation.
The candidate was rejected.
The repair was not to add another evaluator or make the verifier more capable. The relevant change was to move verifier authority outside the candidate’s control. A later bounded successor passed more than 100 focused tests and a new review of that mechanism under the separated authority structure.
Even then, the broader control system was not declared complete. Whether consequential effect paths are forced through the relevant control boundary remains a separate assurance problem.
Independent verification and complete mediation are different requirements
For verification itself to be independent, the verifier must be able to reach a conclusion that differs from the system doing the work, and the subject should not control the mechanism that determines who is trusted to verify it.
A separate question is whether consequential actions can bypass the surrounding control path entirely.
If they can, the verifier may still be genuinely independent while the system lacks a complete enforcement boundary.
Computer security has dealt with that second problem for decades. NIST defines a reference monitor as a mechanism that is always invoked, tamperproof, and verifiable. The requirement that it always be invoked is complete mediation: a correct control is not a complete control boundary if relevant operations can route around it. [3]
AWS documents the same problem in AgentCore. A gateway can apply authorization policy, guardrails, and interceptors outside the agent environment, but AWS states that those controls protect the system only when traffic actually passes through the gateway. Direct runtime access bypasses them. [4]
Software-supply-chain verification illustrates the authority problem from another direction. SLSA’s Verification Summary Attestation identifies the verifier, subject, policy, and verification result. Consumers are expected to verify the attestation using preconfigured roots of trust and accept only appropriate signer-verifier relationships. [5]
Capability and authority do not have to live in the same place
None of this requires assuming that advanced AI systems will always behave adversarially. Ordinary bugs, stale state, ambiguous external effects, permission errors, incorrect assumptions, and evaluation failures are enough to create the problem. Greater capability can increase the range and consequence of actions available when those failures occur.
Google DeepMind’s AI Control Roadmap takes a similar system-level approach, describing trusted supervisors, permission expansion based on verified behavior, and stronger intervention as potential consequences increase. [6]
Capability and authority are different properties.
Verification can inform control without becoming control
Verification does not always need to block execution. Post-hoc evaluation can still be valuable for observability, audit, model improvement, and failure analysis.
But if verification is intended to function as part of a control boundary in a higher-consequence workflow, its result must be able to affect authority.
proposed action → evidence → verification → authority decision → effect → readbackThe exact implementation will differ across systems. The useful questions are whether the boundaries are mechanical, whether the verifier’s authority can be altered by the subject, whether relevant execution paths can bypass the control layer, and whether the system confirms what actually happened after an external effect.
AI can participate throughout this process. Models can be evaluators, critics, red-teamers, anomaly detectors, or evidence synthesizers.
The narrower requirement is that the system being evaluated should not also control the conditions under which its own work becomes trusted.
What the ODIN case establishes
The ODIN example is first-party engineering evidence from one system and should be interpreted at that scope. It shows that a large passing test suite and nominal generator-verifier separation were insufficient to establish the independence property being tested. The correction had to change the authority structure, not merely the evaluator’s behavior.
It does not establish that ODIN has solved independent verification generally, that complete mediation exists across consequential actions, or that every agent system requires the same architecture.
Verification claims should be evaluated at the level of authority and system topology, not only at the level of model identity or test performance.
A stronger standard for agent verification
As agentic systems move deeper into operational work, “another model checked it” will become an increasingly incomplete description of assurance.
A stronger verification claim should explain what system was tested, what evidence the verifier received, what policy it applied, what made the verifier trusted, whether the subject could influence that authority, and what broader control boundary surrounds the verified action.
Those questions are more demanding than ordinary output evaluation. They are also closer to the problems that determine whether an AI system can be trusted with consequential operational authority.
AI systems can participate in evaluating their own work.
They should not be able to define, by themselves, the trust structure that makes that evaluation sufficient.
Sources
- [1] OpenAI, A shared playbook for trustworthy third party evaluations, May 29, 2026. Covers evaluation validity, agent harnesses, reward hacking, contamination, refusals, sandbagging, and limits on the claims individual evaluations support.
https://openai.com/index/trustworthy-third-party-evaluations-foundations/ - [2] Microsoft Research, The Art of Building Verifiers for Computer Use Agents, April 21, 2026. Describes the Universal Verifier and the separation of process, outcome, and controllable versus uncontrollable failures.
https://www.microsoft.com/en-us/research/articles/the-art-of-building-verifiers-for-computer-use-agents/ - [3] NIST Computer Security Resource Center, Reference Monitor. Defines the reference-monitor requirements of complete mediation, tamper resistance, and verifiability.
https://csrc.nist.gov/glossary/term/reference_monitor - [4] Amazon Web Services, Security best practices for Amazon Bedrock AgentCore Runtime. Documents gateway policy, guardrails, interceptors, and the requirement that traffic flow through the governed gateway for those controls to apply.
https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/runtime-security-best-practices.html - [5] SLSA, Verification Summary Attestation. Defines verifier identity, subject, policy, verification result, and the trust relationships consumers must validate.
https://slsa.dev/spec/v1.2/verification_summary - [6] Google DeepMind, Securing the future of AI agents, June 18, 2026. Introduces the AI Control Roadmap and system-level monitoring and control mechanisms around increasingly capable agents.
https://deepmind.google/blog/securing-the-future-of-ai-agents/