Summary
Artificial intelligence (AI) assistance helps a person research, analyze, draft, compare, or prepare a decision. Full automation lets a system complete a defined process without case-by-case human direction. That choice determines more than efficiency. It determines who interprets information, handles exceptions, approves actions, and answers for the result when something fails.
Massachusetts Institute of Technology (MIT) researchers reviewed 106 experiments comparing human, artificial intelligence, and combined performance. Human-system combinations did not consistently beat the strongest human-only or system-only approach, and the result depended on how the work was divided. Stanford’s Human Agency Scale reaches a related conclusion from a different angle: tasks inside the same occupation can require very different levels of human involvement.
The unit of analysis is the task, not the job, the department, or the software product. Most valuable workflows contain both models at once, a system handling routine collection while a person retains the consequential decision.
The structural reality: two ends of one operating range
Assistance keeps a person active in the work. The system retrieves information, summarizes records,
You’re staring at a workflow right now, trying to decide whether artificial intelligence (AI) should handle it alone. Maybe someone on your team needs to stay in the loop instead. That decision looks like an efficiency question on the surface. Underneath it, you’re deciding who answers for the outcome when something goes wrong. You’re deciding who catches the exception nobody saw coming, and whose judgment is shaping the result instead of rubber-stamping it.
Researchers at the Massachusetts Institute of Technology (MIT) ran 106 experiments comparing human performance, AI performance, and the two working together. Teaming a person with a system rarely beat the stronger of the two working alone. What decided the outcome, every time, was how the work got split between them. Stanford’s Human Agency Scale found something you’ve probably felt yourself. Two tasks inside the exact same job can call for wildly different levels of human involvement, even though the job title never changes.
Forget the job title. Forget the department. Look at the task in front of you. Most of the workflows worth building run both models at once. A system handles the routine collection, and a person keeps the one decision that carries the real weight.
Two ends of one range, not two separate tools
Think about assistance as keeping you in the driver’s seat. The system pulls the research, summarizes the account history, or drafts the first version, and you decide what to do with it. Picture a sales rep who gets a full account brief handed to them, but still owns the outreach strategy completely. That rep stays central to the work because their judgment is doing something, beyond approving what someone else already decided.
Full automation is different. It takes the wheel entirely, inside a boundary you set. It receives approved inputs, applies the rules you defined, takes the permitted actions, and logs what happened. Your team focuses on design, monitoring, and the exceptions that fall outside the boundary. That boundary has to hold. An invoice-matching workflow can compare records and flag discrepancies for someone to look at. It should never quietly earn the authority to change payment terms on its own.
The real question: does this task earn automation?
Six factors decide that. How often does the task happen? How much judgment does it demand, and how do its exceptions behave? How reversible would a mistake be, how much confidence have you earned in the system, and how serious do the consequences get when it fails? Score each one from one to five, one leaning toward assistance, five leaning toward full automation, then add them up out of 30.

Frequency is a return-on-effort question. If a task happens thousands of times a month, the integration and monitoring work behind automation pays for itself. If it happens twice a year, that investment almost never makes sense, no matter how cleanly you can define the task on paper. Score it a five for steady, frequent demand, and a one for anything rare or unpredictable.
Judgment asks a simple thing: is a person interpreting the situation, or only following a rule? A pricing conversation with your biggest account needs judgment no system has access to, relationships, history, a read on the room. A routine invoice match, once you’ve defined what counts as an acceptable difference, needs almost none of that. Score it a five when your experienced people land on the same answer every time from the same inputs. Score it a one when real judgment is doing most of the work.
Exceptions matter more than raw volume ever will. Every process hits them eventually. What matters is whether they show up often, get caught reliably, and land somewhere safe. It could go the other way too, rare, wildly different from each other, and easy to miss entirely. A low exception rate can still be a real problem if the one case that slips through carries serious weight. Score it a five when exceptions are rare and get caught reliably, and a one when they’re frequent or genuinely hard to spot.
Reversibility asks whether you can undo the whole situation, beyond fixing a database field. A recalled email doesn’t un-read itself in someone’s inbox. Score it a five for actions that are easy to reverse or carry limited fallout, and a one for anything hard to undo completely.
Confidence has to come from evidence you’ve collected: representative testing, stable performance in production, exceptions that route where they should. It should never come from how polished a system sounds, or one generated confidence score dressed up to look certain. Anthropic’s own data found that people approved roughly 93 percent of repeated permission prompts, and their attention dropped steadily the longer that kept happening. Asking someone to confirm something over and over is not the same thing as earning real confidence. Score it a five when tests and production evidence show dependable performance, and a one when performance is still unproven or shaky.
Consequences cover what happens when this goes wrong: revenue, customers, your people, legal exposure, your reputation in the market. A high-value process can still earn automation if its failures stay visible, contained, and recoverable. Score it a five when errors stay contained and recoverable, and a one when they could cause real business or human harm.
Add it up. Six to 14 points means the task needs a person. Build for assistance, and let AI support the research, the drafting, and the analysis, not the final call. Fifteen to 23 points means you’re looking at a hybrid, automating the predictable parts while a person keeps the decisions and exceptions. That middle range wins more often than either extreme, so don’t be surprised when most of your best candidates land there. Twenty-four to 30 points supports full automation, once you’ve built the permissions, monitoring, and exception handling it needs before it ever touches real work.
One thing the total can never do is override a single serious consequence hiding inside it. A workflow can score a 27 overall and still need a person to sign off on one thing buried inside it. An employment decision, a legal exposure, a customer commitment, any one of those still needs a human signature. The total doesn’t get the last word there. You do.
What this looks like on your team
Say your marketing team is turning one approved article into five channel-specific posts. That scores high on frequency and reversibility, but real exceptions show up around claims and executive voice. That combination puts you squarely in hybrid territory: automate the drafting and production work, and keep a real person reviewing anything headed out the door.
Now picture your sales team trying to decide which strategic accounts deserve executive outreach. That scores low on judgment and confidence, because the signals are incomplete and a bad approach can genuinely damage a relationship you’ve spent years building. That calls for assistance: let the system pull the evidence and lay out the options, and let the account owner make the actual call.
Compare that to finance matching standard supplier invoices against purchase orders. High frequency, high reversibility, high confidence, and judgment barely enters into it once you’ve defined what counts as an acceptable difference. That’s a strong candidate for full automation on the matches themselves. Discrepancies and payment approval stay with a person who’s authorized to make that call.
Earn automation in stages. Don’t hand it over on day one.
Nobody should go straight from manual work to full independence in one leap. Build the evidence in stages instead. First, AI prepares the work while a person completes it. Track every correction so your team genuinely learns something. Second, the system runs the whole process under full review. You check honestly whether that review still catches real problems, or has quietly turned into a rubber stamp. Third, the system handles normal cases while people review the exceptions and a sample of the routine work. Fourth, normal cases move without case-level review at all, while your team watches performance, incidents, and permissions through regular audits. The moment you add new users, new data, or new actions, pull the oversight back up until the new evidence earns it again.
OpenAI’s own guidance for building agents follows this exact pattern: start narrow, validate with real users, add controls as real failures show up. Let the evidence lead. Don’t let autonomy become the thing you assume is progress because it feels like moving forward.
Build different controls for each model
Assistance can feel safe because a person is technically involved. It still isn’t automatically safe. It can expose sensitive data, or hand a busy employee a confident-sounding answer they accept without a second look. Name the approved product and account, and the data it’s allowed to touch. Name who’s supposed to use it, what it can’t do yet, and who still holds the final call. Your people need to know exactly what judgment you’re still counting on them to bring.
Full automation removes the person from each individual case, so its controls have to work harder to make up for that distance. Give it a named business owner, its own system identity, and only the permissions it needs. Test it against normal cases, keep detailed logs, build a manual fallback, and confirm the shutdown method works before you need it. The National Institute of Standards and Technology (NIST) recommends naming exactly which capabilities need human oversight, then measuring how much oversight is genuinely happening. A control sitting in a binder somewhere protects nobody.
The same six factors, playing out differently by function
In marketing, the line runs between strategy and production. Choosing the message, setting the direction that shapes your next three campaigns, that needs a person. Turning approved content into different channel formats can run on its own. Claims, sensitive topics, and anything in an executive’s voice still need a qualified reviewer. In sales, the line runs between research and relationship. The same signal can mean completely different things depending on context only your salesperson holds. Research and routine administration can automate once a signal triggers them, but pricing and commitments stay human. In finance, the line runs between processing and authority. Matching records automates cleanly. Variance analysis needs someone who can judge which explanation the numbers support. Payments and exceptions stay with someone who has the authority to approve them, no matter how routine the work looks. In operations, the line runs between routine and disruption. Defined checks can run on their own. A safety issue or a major customer failure needs a real person standing behind it, every single time.
Watch for these signs in both directions
You’ve got too much automation running when your people keep fixing the same errors over and over. The same is true when exceptions you’ve never seen before keep showing up, or when reviewers approve work without genuinely looking at it. You’ve also got too much when the system starts making calls outside what it was approved to do. Fix it with stronger assistance, restored approval, or a tighter boundary. Don’t treat that as a failure. Pulling back autonomy is good management.
You’ve got too little automation when your reviewers rarely change a normal output. The same approved action follows every single review, and your evidence has held steady across time and across different users. Fix that with a limited, tested expansion, and set a clear condition for pulling it back if the evidence ever stops holding up.

Give this 30 days before you commit to anything
Week one. Map the whole process. Follow real cases from the trigger all the way to the finished action. Write down every communication, every record change, and every consequence for a customer or an employee along the way.
Week two. Score all six factors, and put real evidence beside every number. Separate the routine preparation from the decisions that carry real weight inside the same workflow.
Week three. Test both models for real. Run assistance with your actual employees. Run automation against a narrow set of stable cases too, including the weak inputs and known exceptions you’d normally see.
Week four. Approve the model one stage at a time. Set the permissions, the monitoring, the exception routing, and the measurement. Write down exactly what evidence has to show up before this earns more autonomy.
Where this leaves you
Frequency is an economics question. Judgment, exceptions, reversibility, confidence, and consequences are an operations question, and you need both answered before you can defend this decision to anyone. One average score for your whole department will hide the truth. Your best workflows split cleanly along these six lines, one stage and one task at a time.
The teams getting real value from AI in 2026 are scoring the work honestly and letting the evidence set the boundary. They pick the label last, after the evidence is in, not first as something the rest of the process has to justify.

