Here’s a question I have been asking technology leaders lately: “If nobody changed the agent, why did its behavior change?” Sit with that for a second.
No new prompt was deployed. No developer changed the code. No one intentionally swapped the model, modified the guardrails, or updated the agent’s tools. The change-control system is clean. The release pipeline shows nothing. And yet, the agent that performed acceptably last month is now making more errors, taking different paths through a workflow, escalating decisions it used to handle, or worse, confidently taking actions it should have escalated.
The most common response I hear is: “That should not happen unless something changed.” Something did change.
The problem is that we are still defining “change” the way we did for traditional deterministic software: code, configuration, or infrastructure modified through an explicit deployment. AI Agents do not live inside such a clean boundary. They operate inside a changing system of models, harnesses, tools, data, other software, other Agents, and humans. Any one of those elements can change the agent’s behavior, even when no release was made to the agent itself. That is Agent drift.
Agent drift is bigger than model drift
When most teams discuss drift, they mean model drift: the statistical relationship between inputs and outputs changes, or the model’s production behavior moves away from the baseline measured during testing. That matters. But for an AI Agent, it is only one part of the problem.
An Agent is not just a model wrapped in an API. It is a software system with Agency and Autonomy. It reasons across context, invokes tools, exchanges data, interacts with other Agents, and takes action in systems that have their own release cycles and operational states. Its behavior is therefore the product of the whole system, not merely the weights and capabilities of the underlying LLM.
So when the output changes, asking “did the model drift?” is too narrow a question. The right question is: “What changed anywhere in the system of systems that produced this outcome?”
This distinction matters because model monitoring alone will not catch all Agent drift. You may have the same model, the same prompt, and the same agent code, and still get meaningfully different business outcomes. NIST’s guidance recognizes the need for post-deployment monitoring of generative AI systems rather than treating evaluation as a design-time activity alone. The operational implication is straightforward: if behavior can change continuously, evaluation must also operate continuously.
The five dimensions of Agent drift
I see Agent drift occurring across five dimensions. These dimensions are interconnected, and more than one may be changing at the same time, which is what makes diagnosing drift significantly harder than reviewing a code diff.
Model changes
This is the dimension most teams already recognize. The underlying model changes, and the Agent behaves differently as a result.
The change may be explicit: your engineering team moves from one model version to another, replaces one provider, modifies inference parameters, or introduces a fine-tuned model. But it may also happen outside your deployment process. A model provider can update, deprecate, reroute, or retire a model. A platform can change safety controls, tool-calling behavior, context handling, or content filters. Microsoft’s lifecycle guidance treats foundation models as versioned dependencies whose updates can introduce untested behavior changes, while Apple explicitly advises developers to retest prompts when its underlying on-device model changes. Even an improvement can create drift.
A newer model may reason better overall but respond differently to a prompt pattern your harness relies on. It may call tools more aggressively, interpret an escalation threshold differently, or refuse requests the previous version completed. “Better model” does not automatically mean “better behavior for this Agent in this business process.” The Agent did not change. Its decision engine did.
Harness changes
The Agent harness is everything around the model that makes the Agent operational: orchestration logic, prompts, memory, context assembly, retrieval, tool definitions, routing, guardrails, retry logic, and the mechanisms used to invoke other Agents.
A harness update can alter behavior without changing the Agent’s apparent business logic. A revised system prompt may subtly reorder priorities. A memory implementation may retrieve different history. A new summarization method may remove context the Agent previously used. A routing optimization may send certain tasks to a smaller or faster model. A changed retry policy may turn a recoverable tool error into a different decision path.
And here is the part that catches organizations off guard: the harness may be supplied by a platform or vendor with its own release cadence. Your team may not have deployed anything, but the environment coordinating the Agent may have changed underneath you.
This is why evals must cover the assembled Agent, not only the LLM. OpenAI’s own agent-evaluation guidance connects workflow traces and graders to repeatable evaluation runs that can be used to refine prompts, tools, routing logic, and guardrails. Those are harness elements, and every one of them can move the Agent away from its validated baseline.
System changes
Agents do not operate in isolation. They interact with software systems, APIs, data stores, enterprise platforms, and other Agents. Those systems change constantly.
An API changes the structure of a response. A CRM team modifies a field definition. A data pipeline begins delivering information later than before. A security team changes permissions. A knowledge base is reorganized. A SaaS vendor adds a new status value. Another Agent updates the format or semantics of the messages it sends. None of these is a deployment to your Agent. All of them can change its behavior.
Traditional integration testing catches some structural failures: a missing field, an invalid schema, a broken endpoint. Agent drift can be subtler. The API still responds. The schema is still valid. The data is simply different enough in meaning, completeness, timing, or distribution that the Agent begins reaching different conclusions. The system is technically up. The Agent is technically working. The business outcome is quietly degrading.
Everything looks healthy until someone actually measures whether the Agent is still making the right decisions.
This is the drift most likely to hide inside green dashboards, because availability and latency remain within bounds while efficacy declines. Everything looks healthy until someone actually measures whether the Agent is still making the right decisions.
Agent learning
Some Agents are designed to learn. They retain memory, incorporate feedback, evaluate their own performance, adapt plans, or modify future behavior based on prior actions and outcomes. That sounds desirable, and often is. But a self-learning or self-evaluating Agent is, by definition, changing over time.
What did it learn? From which interactions? Was the feedback accurate? Did it improve one class of task while degrading another? Did it learn from exceptional behavior and turn it into the new normal? Did a short-term optimization weaken a guardrail or create a long-term bias?
A recent survey defines self-evolving agents as systems that modify parameters, context, tools, or architecture based on their trajectories or feedback, and it identifies continuous behavior monitoring as necessary to detect longer-horizon drift. That is the central governance problem: an Agent that learns after deployment has a production change mechanism operating outside the conventional release pipeline.
We would never let a human employee rewrite a regulated operating procedure every evening based solely on what appeared to work that day, then put the revised procedure into use the next morning without review. Yet we are increasingly designing Agents that can achieve the software equivalent. Learning is change. Change requires evaluation.
Learning is change. Change requires evaluation.
Human behavior changes
This is the least technical dimension and, in many organizations, the least monitored. Humans change how they interact with an Agent over time. Early users are cautious. They provide detailed instructions, check outputs, challenge recommendations, and escalate uncertainty. Then the Agent performs well for a few weeks or months. Trust grows. Inputs get shorter. Review becomes lighter. Exceptions stop being challenged. The human-in-the-loop slowly becomes a human-near-the-loop, and eventually a human-who-assumes-the-loop-is-working. The Agent may not have changed at all. The operating behavior around it did.
Overconfidence can also change the inputs an Agent receives. People begin delegating tasks that are more complex, more ambiguous, or more consequential than the tasks for which the Agent was originally evaluated. A support copilot becomes an automated responder. A recommendation tool becomes a decision maker. A coding assistant gains permission to deploy. A productivity aid quietly becomes part of a critical business control. The validated use case did not fail. The actual use case drifted away from what was validated.
This is why human behavior must be treated as part of the Agent system. If people provide inputs, review outputs, approve actions, override decisions, or choose when to invoke the Agent, changes in their behavior can alter the Agent’s real-world risk even when every technical component remains static.
There is no static error rate
This brings us to the assumption I think technology and risk leaders most urgently need to discard: that an AI Agent has a fixed error rate. It does not.
An evaluation run may tell you that an Agent produced an unacceptable outcome in 2% of test cases at a point in time, under a specific model version, harness configuration, system state, dataset, user population, and set of operating conditions. That number is a measurement. It is not a permanent property of the Agent. Change any of those conditions and the error rate can change too. More importantly, they may change without an explicit deployment you control.
The same is true of efficacy. An Agent that successfully resolves 90% of a class of incidents today may not sustain that performance as incident patterns change, systems evolve, users alter how they describe problems, or the Agent learns from previous resolutions. Production machine-learning guidance has long warned that changes in serving data and feedback loops can create skew and negatively affect performance, which is why explicit monitoring is recommended. Agentic systems expand that problem across far more than the model and its input distribution. So a quarterly eval gives you a quarterly snapshot of a continuously moving target. That is not governance. It is archaeology.
What continuous evaluation looks like
Continuous evaluation does not mean running every possible eval against every transaction. It means designing an evaluation system proportionate to the Agent’s risk, autonomy, and potential impact.
Start with a defined baseline. What does acceptable behavior look like? Which outcomes are unacceptable? Where must the Agent escalate rather than act? What level of precision, recall, safety, fairness, security, and business efficacy does the use case require?
Then instrument the full path. Capture the model and version, harness configuration, prompts, retrieved context, tool calls, system responses, Agent-to-Agent interactions, human inputs, overrides, approvals, outcomes, and feedback. If an eval detects drift but you cannot trace the outcome across these layers, you know that the Agent changed but not why.
Run evals across three operating rhythms:
- On change: Evaluate every explicit modification to the model, prompt, harness, tool, policy, data source, or connected system before production use.
- On a cadence: Re-run representative eval suites even when no known change occurred, because not every dependency change will arrive through your release pipeline.
- On live behavior: Sample and evaluate production traces continuously, with higher coverage for high-risk actions, anomalies, overrides, and previously unseen scenarios.
The cadence should follow risk, not convenience. A low-impact internal summarization Agent may justify daily or weekly sampling. An Agent making customer-impacting, financial, security, healthcare, or infrastructure decisions may require near-real-time evaluation and alerting.
And when the error burn rate crosses its threshold, something needs to happen automatically: reduce autonomy, increase human review, restrict tool access, roll back a model or harness version, route to a safer path, or stop the Agent entirely. A dashboard without a response mechanism is just a better view of the incident arriving.
A dashboard without a response mechanism is just a better view of the incident arriving.
Final thoughts
AI Agents do not sit still after deployment. The model can change. The harness can change. Connected systems can change. Learning Agents can change themselves. Humans can change how they use, trust, and supervise them. And all of this can happen without a single commit to the Agent’s repository.
That is why evaluating Agents only when we intentionally deploy a change is not sufficient. It assumes that the only meaningful changes are the ones we make and the only release cycle that matters is ours. Neither assumption survives contact with a production Agentic system.
The objective is not to prove once that an Agent is safe, effective, or compliant. The objective is to continuously gather evidence that it remains within the boundaries we established for it, and to act before a changing error rate consumes its error budget.
We learned this lesson in SRE: reliability is not a certification you receive before launch. It is an operating discipline. AI risk is no different.

