AI systems are going to fail. The important question is what happens after they do. Most systems can already record failure through logs, traces, tickets, evals, postmortems, and audit trails. Those matter, but recording a failure is not the same as learning from it. If the same system can encounter the same situation tomorrow with no structural change that makes recurrence less likely or more detectable, the lesson never became part of the architecture. That is the distinction. Failure should be a first-class object.

Recording failure is only the beginning

Reliability engineering already understands this principle. Google SRE uses postmortems to understand incidents and drive preventive actions intended to reduce recurrence. [1] NASA's lessons-learned lifecycle ends with Apply, where lessons are integrated into processes, checklists, handbooks, and formal policy. [2] NIST's AI Risk Management Framework calls for post-deployment monitoring, incident response, recovery, change management, and measurable continual improvement. [3]

The systems are different, but the principle is the same: a lesson becomes institutional when it changes future behavior.

AI agents make this harder because the final result can be correct even when the execution path was not. An email can send, a file can be created, or a database can update while the system also performs actions it was never authorized to perform. Outcome success is not execution correctness, and a successful final result should not erase a bad path.

That distinction matters more as agents gain access to production systems. If success is measured only by the final output, a system can accidentally learn that expanding beyond its authorized scope is acceptable as long as the task eventually works.

The mechanism

A weak failure loop looks like:

failure → log → warning → hope the next model remembers

A stronger loop looks like:

failure → evidence → classification → correction → verification → control change → future execution → recurrence tracking

In the stronger version, the failure becomes something the system can identify, inspect, connect to evidence, and use to constrain future work. The exact schema matters less than the consequence: what is structurally different because the failure happened?

The goal is not a system that never fails. It is a system that becomes structurally harder to fail the same way twice.

A first-class failure object might preserve the expected state, what actually happened, the evidence, the failure class, the affected boundary, the accepted correction, the verification state, any resulting control change, and whether the failure later recurred. The point is not to create more documentation. The point is to give the architecture something durable that future execution can encounter.

Memory is not enforcement

Suppose an agent performs a deployment during a transaction that should only send a message. You can store a warning telling future models not to deploy, or you can change the system so deployment is unavailable during that transaction class.

Those are fundamentally different responses. The first depends on the model remembering and correctly interpreting the rule. The second changes what the model is actually allowed to do.

The stronger system does not merely remember the lesson. It enforces its consequence.

This is where observability and governance separate. Logs and traces can explain what happened. Evals can test behavior. Durable execution can preserve state across retries. Independent verification can judge whether a correction actually worked. None of those mechanisms alone guarantees that a known failure changes the next execution path.

The stronger architecture connects them.

Agent evaluation and observability systems are moving toward recurrence-aware loops

Anthropic argues that systematic agent evaluations help teams avoid reactive production loops where fixing one failure creates another, and that evaluation becomes more valuable when maintained across an agent's lifecycle. [4]

Microsoft recommends continuous agent evaluation and specifically identifies production incidents, model changes, major knowledge changes, and new tools or connectors as reasons to rerun evaluation suites. [5]

LangSmith Engine pushes the loop further by turning recurring trace failures into diagnosed issues, proposed fixes, offline evaluation examples, continued recurrence tracking, and automatic reopening when the same problem returns. [6]

That is meaningful progress. The next question is how far the loop should go.

My view is that verified failure knowledge should be able to trigger a governed change in future execution authority when warranted by the evidence. The consequence might be a regression test, a mandatory source check, a narrower permission, an approval boundary, or a capability becoming unavailable inside a specific transaction.

Not every failure deserves another control. But material failures need a path from what went wrong to what is now structurally different because we learned from it.

Recurrence resistance is the better standard

AI reliability should not depend on models eventually becoming incapable of mistakes. A more useful standard is whether the system can reconstruct a failure, classify it, connect the correction to evidence, independently verify that correction, and make the same failure harder to repeat.

That also makes the learning more durable than any individual model. Models, providers, and tools will continue to evolve, but the institution should not have to relearn the same operational lesson every time the underlying technology changes.

Failure is inevitable. Relearning the same failure indefinitely should not be.

Sources

  1. [1] Google Site Reliability Engineering, Postmortem Culture: Learning from Failure. Postmortems document incidents, contributing causes, corrective actions, and follow-up intended to reduce recurrence.
    https://sre.google/sre-book/postmortem-culture/
  2. [2] NASA APPEL Knowledge Services, Lessons Learned. NASA's lifecycle is Collect → Record → Disseminate → Apply, with lessons incorporated into future processes, checklists, handbooks, and policy.
    https://www.nasa.gov/learning-resources/for-professionals/appel-lessons-learned/
  3. [3] NIST, AI Risk Management Framework, MANAGE 4.1–4.3. Covers post-deployment monitoring, incident response and recovery, change management, incident and error tracking, and continual improvement.
    https://airc.nist.gov/airmf-resources/airmf/5-sec-core/
  4. [4] Anthropic, Demystifying evals for AI agents. Covers agent evaluation, trajectories, outcomes, regression testing, and maintaining evaluation systems across the agent lifecycle.
    https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
  5. [5] Microsoft, Review the agent evaluation checklist. Recommends continuous evaluation and rerunning evaluations after production incidents and significant system changes.
    https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/evaluation-checklist
  6. [6] LangChain, LangSmith Engine. Documents recurring-issue diagnosis, proposed fixes, offline evaluation examples, continued recurrence tracking, and automatic reopening when an issue returns.
    https://docs.langchain.com/langsmith/engine