Summary
Picture the meeting where someone pulls up a slick demo and calls it proof the workflow is ready to launch. Nobody in the room can explain how a request moves from start to finish without that demo running. That is the moment to stop and ask the harder questions instead, before the workflow reaches anyone’s real work.
An artificial intelligence (AI) workflow should enter production only after your team can explain how work moves from request to result. That explanation should hold up out loud, in a room, without anyone reaching for a demo to make the case. Most organizations already use AI, yet nearly two-thirds have not started scaling it across the enterprise. McKinsey found that high-performing organizations were nearly three times more likely to redesign the individual workflow first. MIT research reaches a similar conclusion. AI creates more value when leaders examine task sequences and handoffs across the entire workflow, not only the step getting automated.

These 12 questions are built for you to ask directly, in one working session. Bring the business owner, subject experts, frontline users, and technical and risk partners into the room together. A weak answer becomes a design task with an owner and a due date. That follows the same logic NIST uses to govern, map, measure, and manage AI risk.
Question one: is the workflow stable enough to automate?
Ask three people to explain how the work happens today. Significant differences between their answers usually reveal undocumented decisions or competing definitions of what “done” means. Review several recent cases start to finish, and document every step and correction without cleaning anything up first.
A strong answer identifies the event that starts the workflow and follows a generally consistent sequence. It can also say exactly where variation is acceptable versus where it signals a real problem. A workflow can contain judgment and still be stable, since stability means shared operating logic, not identical treatment in every case. Pause automation when the process changes weekly or completion depends on one person’s undocumented memory. That is a process design problem, not a technology problem.
Question two: are the inputs available, reliable, and usable?
Every AI workflow depends on what feeds it, and teams design around ideal inputs while daily work arrives incomplete, late, or duplicated. Build an input inventory naming each source, its owner, and its known quality problems. Include the information employees currently gather by hand because no formal system holds it.
Test with strong, average, and weak input packages, and record what happens when a required field is missing. A strong answer names which inputs are required versus optional, and which source holds authority when two disagree. It also names what the workflow does when required information is not there. AI can flag a missing or conflicting field. A person or a system still has to close the gap, since AI alone cannot repair an information system nobody owns.
Question three: can your team explain the decisions inside the workflow?
A workflow holds more than tasks. It holds decisions about priority, quality, and the next appropriate action, and most teams only discover those decisions after automation produces a result nobody expected. Build a decision table for every point where the workflow can change direction, naming the condition, the required evidence, and who holds override authority.
“Send qualified accounts to sales” leaves too much room for interpretation. A usable rule names what qualifies, what evidence supports it, and who can override the call. Some decisions should stay judgment-based permanently, and that is fine. The workflow needs to name who makes them and how the decision returns to the process afterward.
Question four: what exceptions can occur, and where will they go?
Normal cases rarely determine whether a workflow succeeds. Exceptions do. MIT research found that one genuinely difficult task can weaken an entire automated sequence. The workflow needs explicit boundaries around the steps AI cannot handle reliably.
Build an exception register from recent work, support tickets, and rejected outputs, grouped by frequency and business impact. Give every exception a real destination: a person, a review queue, or a request for more information. A different customer format is manageable. Conflicting contractual terms usually are not, and your team should know the difference before launch, not discover it live.
Question five: who owns the workflow after launch?
Every production workflow needs one accountable business owner with real authority to approve changes, assign review resources, and stop the workflow outright. Shared interest across several departments is not accountability. Ownership needs to cover five names: the business result, source data and access, review points and escalations, technical maintenance, and risk and compliance.
Smaller teams can assign several roles to one person, but every responsibility still needs a name and available time attached to it. A workflow without an active owner gets harder to change every month. Source systems, models, and business rules shift underneath it while nobody is watching.
Question six: what requires human review?
Review should scale with consequence, uncertainty, and reversibility, not apply uniformly everywhere. McKinsey found that defining when outputs need human validation was among the strongest practices tied to real AI value. Choose a real review model for each output. Full review before release, conditional review triggered by risk, sampled review, exception-only review, or automatic release inside tightly approved limits are all options.
Reviewing everything can create a new bottleneck that quietly erodes the capacity gain the workflow was supposed to create. Automatic release everywhere can create exposure nobody agreed to accept. The design needs real acceptance criteria, a named reviewer, and a response time, not only the phrase “human in the loop” written into a slide.
Question seven: where will the output go, and what happens next?
An AI output creates value only once it reaches the next person or system in usable form. A polished draft sitting in a disconnected app is still unfinished work. Map the actual destination for every output, whether that is a customer system, a service queue, or an employee inbox. Define the format and permissions it needs to arrive in.
Decide whether the output updates an existing record or creates a new one, since duplicate records quietly damage reporting and trust over time. A workflow is incomplete if employees still have to copy, rename, or route every result by hand. That work belongs in the design itself, or it belongs in the value calculation, not left invisible.
Question eight: can the integrations handle normal operations and failure?
Production workflows usually cross several systems, and every connection adds a permission, a limit, and a possible failure point. Document each integration’s authentication method, data direction, and responsible owner, including vendor dependencies sitting behind the visible application.
Test normal volume, peak volume, expired credentials, and duplicate events directly. Confirm the workflow knows whether to retry, pause, or alert someone when something breaks. A model or application update can quietly change output quality without changing anything visible on the surface. Your team also needs a manual fallback for work that genuinely cannot wait for a fix.
Question nine: are security, privacy, and data rules clear?
AI workflows move information across tools faster than most employees realize. Classify the data before connecting any tool, since public content, confidential business information, and regulated records need genuinely different controls. NIST names security, privacy, accountability, and reliability as core characteristics of trustworthy AI for exactly this reason.
Review vendor terms for storage, model training, retention, and subprocessors, and confirm employee permissions match the minimum access the workflow needs. The workflow should prevent sensitive information from entering prompts or logs without approval. Convenience should never quietly become the default security decision. Legal, privacy, and security partners should review the use case in proportion to its actual exposure.
Question 10: will your team see failures before customers do?
A workflow needs visible operating signals after launch, since silent failures produce incomplete records and false confidence at the same time. NIST recommends continuous monitoring across the AI lifecycle, including incident response, override, and recovery.
Define which events trigger an alert: missing inputs, low-confidence outputs, rising correction rates, or unusual volume. Give every alert a named owner, a severity level, and a response expectation. Build a simple incident process before launch, recording what happened and what changes will prevent it recurring. Make sure the workflow has one clear shutdown method nobody has to hunt for during an actual emergency.
Question 11: how will your team measure business value and operating health?
Measurement has to start before the pilot, since your team needs a real baseline before you can honestly claim improvement. NIST recommends measures tied to purpose, known risk, and real deployment conditions, comparing predeployment and postdeployment performance directly rather than reporting activity in isolation.
Track business outcomes, cycle time, conversion, revenue contribution, alongside operating health: completion rate, correction rate, and cost per completed case. Choose a small set of measures leadership can interpret, since a long dashboard hides weak results behind activity counts. A workflow that saves drafting time while doubling review effort has shifted labor somewhere else, not created real capacity.
Question 12: who will maintain the workflow as work changes?
Every AI workflow changes after launch: source data shifts, policies update, vendors revise their products. Deloitte found that only 34 percent of surveyed organizations were reimagining the business around AI. Companies consistently felt less prepared on infrastructure and data than on strategy alone.
Name who reviews performance, approves changes, and retrains users in your maintenance plan, plus a recurring review schedule that happens. Document the workflow in language both business and technical teams can use: purpose, boundaries, decisions, review points, and change history. Set specific triggers for an immediate review: a policy change, a model update, a material incident. Let maintenance respond to real events instead of waiting for the next scheduled check-in.
Score the readiness review
Score each question from zero to two. A zero means the answer is unknown or unmanaged. A one means a partial answer exists but real design work remains. A two means the answer is defined, documented, tested, and owned. Twelve questions produce a possible score of 24.
- A score from 20 through 24 supports a controlled production launch, assuming no high-risk gap remains.
- A score from 14 through 19 supports a limited pilot with tight boundaries and active human review.
- A score below 14 means the workflow needs redesign before automation, not a faster pilot.
The total score can never override one serious security, legal, or accountability gap. Any unmanaged high-impact risk should pause production regardless of what the number says.

Run the session
Start with one real workflow that consumes meaningful time and has an engaged business owner, not the one that made the best demo. Observe the current work before the meeting, following several real cases across the people and systems involved. Bring the business owner, subject experts, frontline users, and technical and risk partners into one working session. Answer each question against real evidence, not memory.
Assign every unresolved item as design work with a named owner and a due date. Test the revised workflow with representative inputs and known exceptions before launch. Run the pilot inside clear boundaries, and expand only once the workflow proves dependable under normal operating conditions.

What you tell them at the end
The strongest AI workflows begin with a clear operating decision, not a tool selection. A model can generate text, classify records, or complete connected tasks well. The workflow surrounding it determines whether that capability becomes dependable business capacity or an expensive demo that never quite ships.
These 12 questions give your team the shared language to make that call together, in one room. That is what happens before automation makes the hidden work harder to see. Production earns its place once the workflow has clear inputs, decisions, owners, controls, and measures behind it, not before.

