Before handing a decision to an agent, define how you will assess its quality and consequences. Measurement gives you a basis for deciding whether automation improves the work.
The current wave of AI transformation is mostly about offloading decisions. When you augment a process with an agent, the value comes from handing some of the smaller judgments to a system that can make them without a human in the loop. Verification helps establish whether a system meets its requirements, but it is not sufficient for safe automation. Permissions, failure handling, and human review must also match the consequences of the task. Without a way to assess results, you have little basis for deciding whether to trust or expand the automation.
This is why I find the rush past measurement strange. Data-driven decision-making is not as interesting as agentic AI, and it does not present well at a marketing offsite. But it is the layer that decides whether the agents on top of it mean anything. Without it, an AI transformation is theatre: the demo runs, the dashboard glows, and nobody can tell you whether the new process is better than the one it replaced.
Start with the unglamorous part, and admit it is unglamorous
Before you automate a process, you measure it. You define what success means in terms of numbers, record the manual baseline, and after you deploy the agent, record the same numbers again. None of this is new. The difficulty is making baseline capture part of delivery rather than deferring it until after a demo.
What is worth saying is why this step gets skipped. Baseline capture is boring; it happens before the exciting part, and it requires two groups of people who do not usually sit together. The engineers know what can be measured. The domain experts know what is worth measuring. Resolution time on a support ticket is easy to record and meaningless on its own; it matters only once someone from the business explains which tickets carry cost and why. Getting that right is a cross-team conversation, and those conversations are easy to defer under deadline pressure.
So the foundation is old and unglamorous. The interesting part is what the same foundation lets you do afterward, and that is where the same data can support further decisions.
I find it useful to think of the measurement layer as a ladder. You build the instrumentation once: the traces, the recorded metrics, the translation into the KPIs the business actually cares about. The industry is converging on OpenTelemetry as the standard for this, which means the foundation is increasingly something you adopt rather than invent. I see four useful applications of that investment. Each adds capabilities and responsibilities; teams should choose the level of automation that fits their risks and needs.
Rung One: proving the transition was worth it
The first payoff is the one everybody wants. With a baseline and a live measurement system, you can run the old process and the new one side by side, gate the rollout behind A/B testing or a gradual deployment, and say with confidence whether the agent improved on the manual process and by how much. You also account for costs that did not exist before: tokens spent, inference costs, and the cloud bill for any models you run yourself. Performance minus cost, measured against a real baseline. This is the rung that turns “the agents seem to be working” into a number you can take to a board.
A pilot without a baseline can struggle to justify further investment. A working demo alone cannot establish whether the process improved or whether the improvement justifies its cost.
Rung Two: debugging what you actually shipped
The second payoff arrives the moment something goes wrong, and with software, something always goes wrong. The errors are rarely as easy to find as they are annoying. The same data that let you compare the two processes now gives you the fine-grained signal to locate the failure rather than canceling the whole initiative.
This is the rung the team I am a part of is currently standing on. We built a dashboard on top of traces from our agentic system, joined with other production data, that directly surfaces metrics like resolution time for a given class of tickets. When a process underperforms, the question stops being “the agents feel slow” and becomes “this step, on this category, regressed after this change.” That is the difference between a hunch and a diagnosis. None of it is sophisticated; it is plumbing, but it is the plumbing that converts a vague sense of unease into something you can act on.
These first two uses justify investment and help maintain the running system. The next two extend the foundation into optimization and planning.
Rung Three: closing the loop without a human in it
The third payoff is the one we are building toward, because we see it as a significant step on our journey to embrace an AI-first mindset.
Here is the idea. Once the agent and the measurement run without a person in the middle, you have the makings of a feedback loop that improves the system on its own. The metrics feed back into the agent, and techniques like automatic prompt optimization enable the system to adjust its behavior in response to those metrics over time.
The distinction that matters for a decision-maker is this. In ordinary A/B testing, a human reads the result, forms a hypothesis, and writes the next variant. In a self-improving loop, the system scores its own output, works out what went wrong, and generates the next version itself. The human defines the objective and the evaluation criteria; the optimization runs underneath. This is not a research fantasy. Automatic prompt optimization can improve results on a defined evaluation set. Whether it outperforms a manually designed prompt for your workflow must be tested, including on held-out cases.
Guardrails belong at every stage; a loop that changes its own behavior requires additional controls. A loop that rewrites its own behavior against a metric will optimize for exactly that metric, including the parts you did not mean. The objective, constraints, and limits on what the loop may change must be defined with the business and compliance before the loop runs, not after. For systems subject to regulatory or data protection requirements, these controls also need to support review and accountability. Whether a system is ready to deploy depends on the applicable requirements and its risk assessment. The same trace data that powers the loop is what lets you reconstruct, after the fact, why it did what it did. Measurement is what makes autonomy auditable.
Rung Four: deciding what to automate next
The top rung changes who is steering, and it is worth telling as a progression.
At the bottom, a human is behind the wheel by hand. The responsible stakeholder wants the numbers, so they go to each team, wait on someone to compile a report, and read the answer off a spreadsheet days after it mattered. The information exists, but reaching it costs human effort every single time.
The dashboard removes that cost. Because the metrics and KPIs are computed continuously, the people who need the numbers no longer have to wait on those who hold them. They read the current state in real time and recombine it, slicing the system in ways the engineers who built the metrics never anticipated, surfacing insights nobody designed for. Steering becomes something you do continuously rather than in retrospect.
The top of the ladder is forecasting. Once those numbers exist as a time series, you can project them forward to see where the business is heading, which risks are building, and which capacity is sitting idle. And this closes a different loop than rung three. Rung three uses the data to look backward and keep the running system honest. Rung four uses it to look forward and decide which process to automate next, on evidence rather than on whichever workflow the loudest manager wants modernized. The aim is to combine evidence about potential value with domain judgment, feasibility, and risk when choosing the next project.
Conclusion
That progression, from a hand on the wheel to a live dashboard to a forecast you can plan against, is a steady handover of effort to the system, with the human moving up to the decisions that still require judgment.
None of the four rungs is exotic on its own. The reason to take the foundation seriously is that the same boring foundation carries all four, while the value of each application depends on the workflow. The agents are the visible part of an AI transformation, but the underlying measurement determines whether the visible part means anything. Build measurement into the initiative from the start, alongside the agent and its safeguards.


