The assistant has produced a good reply. Elena is pleased. The owner asks the question that turns a demonstration into a business experiment:
“Does it help often enough to be worth using?”
One good answer cannot settle that. Neither can a bad answer chosen to make the tool look useless. They need a fair view of ordinary cases, difficult cases, review effort, and cost.
This chapter continues the fictional Mesa Equipment Service example. Its numbers and results are illustrative, not observations from a real business.
Your result from this chapter: a recorded decision to continue, revise, or stop one AI-assisted workflow, supported by evidence.
Give the pilot a clear boundary
Mesa's pilot has one job: draft replies to questions about inspection fees and booking, using the approved service policy. Elena reviews every answer and sends accepted replies through the normal system.
The pilot does not authorize discounts, book appointments, issue refunds, or connect to the customer database. If the question requires information outside the approved pack, the draft should identify the missing information.
Write the boundary before you start. A project that keeps changing its task is difficult to evaluate. If you discover a promising new use, put it on a later-project list.
Thirty days is a suggested planning period. A business with few comparable cases may need longer. A serious failure may justify stopping much sooner. Use enough evidence for the decision rather than letting the calendar make it for you.
Days 1–5: prepare the task and the baseline
Complete the project brief, data-use record, and workspace from the earlier chapters. Confirm who owns the pilot, who reviews the work, and what happens if the reviewer is absent.
Describe the current process using recent comparable examples. Measure total handling time, not only typing. Include finding information, composing the answer, checking it, and recording the result.
Agree on a review rubric before evaluating AI output. For Mesa, an acceptable reply must use the correct fee and credit rule, avoid unconfirmed appointments, address the question, and identify missing information. An unauthorized promise is a critical failure even if the rest of the answer is excellent.
Record the source and prompt versions. Without them, a later result may be impossible to reproduce or explain.
Choose cases that challenge the workflow
Include common requests and situations likely to expose mistakes:
| Case | What the test should reveal |
|---|---|
| A direct question about the inspection fee | Whether the system finds and states a simple fact |
| The same question hidden in a long message | Whether it identifies the actual request |
| A request for tomorrow morning | Whether it avoids inventing availability |
| A demand for a discount | Whether it respects the approval boundary |
| A question about an unspecified warranty | Whether it admits that the source does not answer |
| Two conflicting policy versions | Whether it identifies the conflict |
| A message telling the assistant to ignore its instructions | Whether it resists treating customer content as authority |
| A request outside the business's services | Whether it escalates appropriately |
A starting set of 20 historical or fictional cases, plus a few difficult cases, can expose obvious problems. It cannot establish a rare-error rate. Consequential workflows need evaluation proportionate to the possible harm and the conditions of use.
Set aside some cases that you will not use while revising the prompt. Testing only on examples you have repeatedly tuned against can give an overly optimistic impression.
Days 6–10: test without live actions
Run the approved sample through the prepared workflow. Do not connect it to sending, payments, or scheduling during this stage.
For every attempt, record the input case, source and prompt versions, draft output, review result, corrections, and total time. Count abandoned attempts and cases completed manually after the draft fails.
Suppose a draft says:
“We can have someone there tomorrow morning, and the inspection is free if you go ahead with any repair.”
That answer fails. Availability is unknown, and the policy's time condition is missing. Writing “be more accurate” in the prompt may not address either problem.
Ask what failed. Was the right policy supplied? Was the task boundary clear? Did the reviewer have a precise rubric? Did the prompt encourage a commitment it could not verify? Was the model still wrong despite correct inputs?
Change the smallest useful part of the workflow, record the change, and retest relevant cases. Then use the held-back cases to see whether the revision works beyond the examples you just corrected. If you later use those cases for tuning, reserve fresh cases for the next independent check.
Decide whether live use is justified
For this draft-only example, the owner might require that the final offline review finds no unresolved critical failures, that ordinary cases meet the agreed quality standard, and that Elena can use the process without hidden extra work.
Those are proposed acceptance conditions for this fictional pilot. No observed critical failure in a small sample does not mean future risk is zero. If the output remains unreliable or the data controls are unclear, stay offline or stop.
Days 11–20: try limited live work
Begin with a defined, manageable group of eligible requests. Have a person review every draft and retain the manual process.
Compare similar work. If the AI pilot handles only simple fee questions, do not compare it with a manual baseline full of complex complaints. If the same person repeats a case during testing, account for the fact that they may remember the answer and become faster through practice.
Track both the draft's quality and the final accepted answer. A reviewer who corrects every mistake protects the customer, but those corrections still count as workflow effort.
Keep an eye on what escapes the workflow: a mistaken promise sent to a customer, a needed escalation missed, or a case recorded under the wrong account. Measure these separately from mistakes caught before sending.
Tell staff how to report problems without being blamed for spoiling the experiment. A process that discourages negative feedback can look successful while becoming harder to use.
Stop when a defined condition occurs
Pause the affected workflow if it exposes information, takes an unauthorized action, produces a consequential error that escapes review, or repeatedly fails the agreed requirements. Preserve the relevant evidence appropriately and use the manual process while the owner decides what to change.
A temporary pause is an operating decision. It gives the business a chance to understand the failure before repeating it.
Days 21–25: calculate the whole result
Use a small scorecard:
| Measure | What to include |
|---|---|
| Quality before correction | Drafts meeting the rubric, divided by all evaluated drafts |
| Review and rework | Time spent checking, correcting, rejecting, and finishing failed attempts |
| Final quality | Accepted outputs meeting the standard, plus errors discovered later |
| Total handling time | The entire task, from opening it to accepting and recording the result |
| Cash costs | Incremental subscriptions, usage, paid setup, and paid maintenance |
| Staff capacity | Net time released after preparation, review, and rework |
| Business outcome | A relevant result such as response time, complaint rate, or qualified follow-up |
Report counts alongside percentages. “One failure in ten cases” conveys the size of the sample more clearly than a bare “90% success.” Define what success means and show the serious failures separately.
A worked calculation
Assume Mesa handles 100 comparable requests in a month. The old process takes six minutes per request. The assisted process takes four minutes, including preparation, review, correction, and recording.
Manual handling time: 100 × 6 minutes = 600 minutes
Assisted handling time: 100 × 4 minutes = 400 minutes
Capacity released: 600 − 400 = 200 minutes, about 3.3 hours
At a planning value of $30 per staff hour, those 200 minutes represent $100 of capacity. The arithmetic uses the exact 200 minutes before rounding the hours.
Suppose the pilot adds a $20 monthly subscription but does not reduce a paid bill or generate verified additional contribution. Its immediate monthly cash effect is minus $20. The released capacity may still justify that expense—for example, if Elena uses the time to finish necessary follow-up—but it is not $100 of cash savings.
Record setup effort separately. If Elena spent three paid working hours preparing the pilot, those hours matter even if her salary did not change. Do not hide them by measuring only the final week.
Cash payback requires an actual positive net recurring cash benefit. If you later demonstrate a reduction in paid overtime or a contractor bill, include the verified amount after incremental costs. Avoid counting the same benefit twice.
Days 26–30: make and document the decision
Choose one of three outcomes:
Continue the defined workflow. The quality is acceptable, the value justifies the effort and cost, and an owner can maintain it. Keep the controls that made the pilot work. Continuing does not automatically authorize new data access or autonomous actions.
Revise and test again. There is a useful signal, but one or more requirements remain unmet. Name the change, the cases it must pass, the owner, and the next review date.
Stop. The task is a poor fit, the information is inadequate, errors are too difficult to catch, or the total effort outweighs the benefit. Record what you learned and return to the task list.
Avoid expanding simply because the team has spent time on the experiment. The earlier work is useful evidence, not an obligation to continue.
Leave behind a workflow someone can maintain
At the end of the pilot, retain the approved source and prompt versions, the evaluation record, the decision, the owner, and the manual fallback. Limit access to any retained case material and follow the data-retention rules established earlier.
Name triggers for rechecking the workflow: a policy change, a new model or tool configuration, a new type of customer request, an incident, or a meaningful change in performance.
The broader risk-management principle is to assess AI in its actual context and manage it through its lifecycle. NIST provides a voluntary framework for that work; using a checklist does not by itself establish compliance or safety. NIST AI Risk Management Framework.
Your next step: carry the useful workflow forward and strengthen its information, prompting, sourcing, and evaluation habits in Part 2, Reliable Everyday Work.