Summary

A company can have an AI system that performs exactly as expected and still have a bad implementation.

The model can be accurate. Employees can use it, and the agent can complete its tasks. The pilot can report time savings, and leadership can point to a successful rollout. Then the system reaches ordinary work.

A customer situation falls outside the expected pattern, and nobody is quite sure who is allowed to make the exception. An agent has access to perform an action that no one remembers explicitly authorizing. A manager discovers that the “small percentage” of cases requiring human review has become a steady stream of difficult decisions. An experienced employee leaves and takes much of the practical knowledge for handling unusual cases with her. Six months later, the workflow is still running, but nobody can say with much confidence whether it is producing enough business value to justify everything required to maintain it.

None of that reflects a model-performance problem. Every one of those is an operating problem instead.

That distinction is becoming much more important as AI moves deeper into business workflows. Stanford’s 2026 AI Index reports that organizational AI adoption reached 88 percent in 2025, while agent deployment remained much earlier in most business functions. At the same time, McKinsey’s 2026 research argues that the companies capturing more value from AI are not simply deploying better technology; they are redesigning workflows, decision-making, and operating models around it.

We are getting much better at adding AI to organizations than we are at redesigning the organization around what that AI actually changes.

The False Comfort of a Successful Pilot

Most AI implementations get evaluated through some combination of technical performance, adoption, risk, and ROI, which are all reasonable things to measure. The problem is that they can look good together while important operating questions remain unanswered underneath them.

Who owns the decisions AI influences, and what exactly has been delegated to an agent? Where is human judgment still required, and what happens when the expert, model, vendor, or knowledge source the workflow depends on is no longer available? Once the organization counts the review, correction, leadership, knowledge, and maintenance work surrounding the system, is the business case still as strong as it looked in the pilot?

These questions tend to get handled separately. Governance deals with risk, IT deals with access, business teams deal with workflows, HR deals with roles, and finance tries to measure value. The AI-enabled workflow does not experience any of that separation, though. All of it meets inside the same work.

That is why I think organizations need to start looking at AI Operating Integrity: the degree to which an AI-enabled capability remains bounded, accountable, judgeable, resilient, and economically justified under real operating conditions. It is a different question from whether the AI works. It asks whether the organization surrounding the AI works too.

First, Decide What the AI Is Actually Allowed to Decide

Consider a relatively ordinary use case: an AI system helps a company determine whether a customer should receive a billing adjustment. The organization can spend a great deal of effort evaluating the model without ever making the decision architecture explicit.

Perhaps AI gathers account history, compares the request with policy, identifies similar cases, and recommends an adjustment. A customer-service employee reviews the recommendation and clicks approve. It would be easy to describe that as human decision-making with AI assistance, though a closer look tells a different story. If the AI selected the evidence, interpreted the policy, ranked the comparable cases, and recommended the exact action, the employee may formally own the final click while the system has already exercised enormous influence over the decision itself.

That is why AI participation and AI authority need to be separated. For any consequential decision, leaders should be able to explain what the AI is doing: informing the decision, analyzing evidence, recommending an action, determining an outcome inside predefined limits, or executing a decision that has already been authorized. NIST’s AI Risk Management Framework makes the same underlying issue operational, calling for organizations to clearly define human-AI configurations, differentiate roles and responsibilities, establish oversight procedures, and ensure executive responsibility for AI-related risk decisions.

The practical unit here is not the model. It is the decision. Once the decision is clear, authority, evidence, review, escalation, and accountability can be designed around it. Without that clarity, organizations tend to inherit whatever decision structure emerges from the tool by default, rather than the one they actually intended.

Then Look at the Work the Agent Has Inherited

Agents make the operating problem more visible because they do not simply advise. They can perform sequences of work on their own.

Suppose the same company gives an agent responsibility for monitoring billing requests, retrieving account information, checking eligibility, drafting a response, updating the customer record, and escalating unusual cases. Whether the agent can technically do those things is a straightforward question. What job the company has actually given it is a much harder one, requiring a defined role, scope, systems access, business permissions, handoffs, exception path, supervision, and stop conditions.

There is an important distinction buried in here too: technical permission is not the same thing as business authority. An agent may technically have the ability to edit a customer record because that permission is necessary for part of its role. That does not mean the business intended to authorize every possible change that permission makes technically possible.

McKinsey’s recent work on agent-enabled organizations makes a similar point. As agents become part of real workflows, companies have to reconsider not just tasks but roles, team structures, oversight, and where humans remain responsible for exceptions and outcomes. This is why agent design needs a deliberate Delegation Envelope. Inside it, the agent can work. At the edge of it, the operating model should tell the system what happens next, whether that means asking for more information, escalating, waiting for a person to approve the next step, or stopping entirely. The organization should not discover that boundary for the first time when something goes wrong.

Human Review Only Works if the Human Can Still Judge

Most organizations understand that some AI-enabled work should remain subject to human review, which is good practice on its own, though it falls short of solving the problem by itself. A person can remain in the workflow long after meaningful judgment has quietly disappeared from it.

Imagine an employee who once handled hundreds of customer situations directly. Over time, AI absorbs the routine work, and only unusual cases reach the employee. At first, this looks like pure efficiency: the employee spends less time on simple cases and more time where experience matters. There is a longer-term question hiding underneath that gain, though. Where will the next experienced reviewer come from? Routine cases may have been part of the apprenticeship through which people learned the patterns, policy nuances, and context required to recognize the unusual ones in the first place. If automation removes that experience without replacing the development path, the company can accumulate Judgment Debt, gaining efficiency today while weakening a human capability the operating model still assumes will exist tomorrow.

This is why “human in the loop” is such an incomplete description. What matters is where the Judgment Points sit, what kind of judgment they require, who possesses it, whether that person sees enough evidence to use it well, and whether the organization is still developing that capability going forward. A person is not a control merely because the workflow contains a human being. The control depends on what that person can actually recognize, challenge, and decide.

The Same Efficiency Can Create a Resilience Problem

Now assume the workflow performs well for two years. The organization becomes comfortable with it. Employees stop performing much of the original process, one internal expert understands the unusual exceptions, and the vendor has become deeply integrated with customer systems. The rules behind several decisions are technically documented, although nobody has looked at the documentation recently.

Then the vendor changes the product. Or the model changes, the expert leaves, or a key integration becomes unavailable. Whether the AI performed well last quarter stops being the relevant question at that point; whether the organization can still carry the business capability through the change is.

That is the difference between technical reliability and Organizational Capability Resilience. AI can create tremendous efficiency while concentrating capability in a model, vendor, integration, individual, or knowledge source, and if those dependencies go unexamined, the organization can become more productive and more fragile at the same time. The answer is not preserving every manual process forever. It is deciding what Minimum Viable Independence the organization needs for a critical capability: what knowledge has to remain accessible, what human competence still has to exist, who could take over if the primary expert disappeared, what data and rationale would be needed to migrate, and what the degraded operating mode looks like if the normal AI path cannot be trusted. The right amount of independence will vary with the consequence of losing the capability. What matters is that dependency becomes a design choice rather than an accidental byproduct of successful automation.

And Then There Is the Business Case

By this point, the AI system may be working very well. It may also be more expensive than anyone realizes, and this is where conventional AI ROI can become misleading.

A pilot reports that AI reduced a task from 60 minutes to 20, and the business case credits 40 minutes of savings. Employees may spend part of that time checking the output, though, and difficult cases may now require more senior review. Managers may spend time resolving escalations, while someone maintains the organizational knowledge the system relies on. Model changes require retesting, governance and documentation consume capacity, and the faster workflow may create more output than a downstream team can actually absorb. The AI may still be an excellent investment, and the value calculation still needs to include the whole operating system to know for sure.

This is the purpose of AI Value Assurance. Instead of stopping at time saved or adoption, the value has to be followed all the way through the workflow: AI capability changes the work, the changed work affects human behavior, and that produces an operating effect. The operating effect should change a meaningful business outcome, and the organization then has to determine whether that outcome is large and durable enough to justify the total cost of producing it. McKinsey’s 2026 operating-model research makes this distinction particularly clear, reporting that many organizations remain stuck at task-level acceleration, while companies seeing greater financial impact are more likely to redesign the underlying workflow and operating model. Faster is useful on its own terms. It only becomes valuable once the whole system around it has been accounted for.

These Problems Compound

This is why treating these issues as separate governance exercises creates trouble. An agent receives more autonomy because its performance improved, which changes decision authority. That change reduces the number of cases humans handle directly, which may change judgment development over time. The increased dependency on the agent changes the resilience profile, and the new autonomy reduces some labor while creating different supervision and exception costs, which changes the value equation all over again. One operating change moves through the entire system rather than staying contained to the part it started in.

This is the core of the Tier 1 architecture I have been developing at Sterling Phoenix. Human-AI Decision Architecture governs how AI participates in consequential decisions and where authority remains. Human-Agent Operating Model governs what work agents perform, within what boundaries, and how responsibility moves between people and machines. Organizational Judgment System addresses whether the organization still possesses the human judgment those designs require, and Organizational Capability Resilience addresses whether critical capability survives changes in people, models, vendors, knowledge, and operating conditions. AI Value Assurance System determines whether the resulting capability produces enough credible, durable value to justify continuing it.

Leaders do not need five separate committees for those five questions. They need to understand that each answer changes the others.

A Useful Test Before You Scale

Before expanding an AI-enabled workflow that appears successful, put one real workflow on the table and walk through it from beginning to end.

Start with the decisions. Identify the consequential choices inside the workflow and be specific about what role AI plays in each one; if nobody can distinguish recommendation from authority, fix that before scaling anything further. Then look at delegation, writing down what the agent or AI-enabled system is actually doing, what it can access, what it can change, what it cannot do, and where work returns to a human. Now look at the humans: identify the points where the operating design depends on experience, context, or judgment, and ask whether those people still have the evidence, time, and practice required to provide it.

Then stress the capability by removing the primary model, vendor, expert, or knowledge source in the scenario. Can the business still produce an acceptable outcome? If not, decide whether that dependence is acceptable and what recovery requires. Finally, reconstruct the value case, including the review, correction, exception, leadership, knowledge, governance, maintenance, and recovery work that lives outside the technology budget.

This is where many pilots become much more interesting. The organization may discover that the AI is less valuable than it thought, or considerably more valuable. It may discover that the value is real but cannot be scaled safely until something around the AI changes first. All three answers are useful, which is exactly the point of running the test.

Approval Should Not Last Forever

There is one more operating discipline that ties these systems together. AI changes too quickly for important operating decisions to be permanent.

A different model can change quality and cost. An agent gains a new capability, a vendor changes terms, or an expert leaves. Volume increases, exception patterns change, and new evidence reveals that the business case was too optimistic or too conservative. When something material changes, the organization should ask whether the conditions behind the original authorization are still true.

I call this the Reauthorization Spine. It is not a demand to restart governance from scratch every month; it is a simple operating rule that material change should not silently inherit old authority. That applies to agent roles, decision rights, human review, fallback assumptions, and the business case itself. AI systems evolve, and their operating authorization should be capable of evolving with them rather than staying frozen at launch.

The Real Test Is Operating Integrity

This leads to a more useful definition of a successful AI implementation. It is not merely one where the model performs well. The organization should be able to answer a much harder set of questions: who may decide, what may be delegated, and where human judgment must remain credible. What happens when a critical dependency changes, and what does the system truly cost to operate? What evidence shows that the business outcome is better, and what change would cause the organization to reconsider any of those answers?

NIST’s current risk-management guidance treats AI risk management as continuous across the lifecycle rather than a one-time approval event, and that same lifecycle discipline needs to reach beyond risk and into the operating design itself. Technical performance can remain healthy well after Operating Integrity has begun to deteriorate underneath it: the agent keeps running, the dashboard stays green, and employees keep using it, even as review turns superficial, key knowledge becomes concentrated, authority quietly drifts, recovery grows uncertain, and the original value assumptions stop matching reality. That is the failure leaders need to get better at seeing before it shows up as a crisis.

AI capability is going to keep improving. The harder advantage will be organizational. Can the company design authority as quickly as it deploys capability? Can it change the division of work without losing judgment, and become more efficient without becoming helpless? Can it measure value without ignoring the human system required to produce it?

The organizations that answer those questions well will be doing something more important than adopting AI. They will be building AI systems that hold up in real work.


Practical Tier 1 Review

For one consequential AI-enabled workflow, leadership should be able to answer these ten questions:

  1. What consequential decisions does AI materially influence, and who has authority over each?
  2. What work has been delegated to AI or agents, and where do those boundaries stop?
  3. Can a human intervene, override, or stop the workflow when necessary?
  4. Is the evidence supporting consequential decisions appropriate to their impact?
  5. Does the organization still possess the human judgment the operating design assumes exists?
  6. Can the capability continue in an acceptable degraded mode if a key model, vendor, integration, person, or knowledge source disappears?
  7. Can the organization recover without reconstructing the workflow from memory?
  8. What material changes trigger reauthorization?
  9. Are review, correction, exception, leadership, governance, knowledge, capability, and recovery costs included in the economics?
  10. Is there credible evidence that the system improves a business outcome enough to justify continued operation?

Those ten questions are a much better readiness test than asking whether the pilot worked.

Share The Article, Choose Your Platform!

Get Weekly Fire

One sharp insight. One strategic framework. One idea you can use before your next leadership decision.

The Sparks newsletter delivers clarity, systems thinking, and AI-era leadership insights for ambitious operators.