A dashboard can show hundreds of AI interactions while telling you almost nothing about business value. A draft can appear faster while the reviewer works longer. Employees can use a tool every day without improving the customer experience. Start your measurement with the job the business needs finished.
Your deliverable is an impact scorecard that connects accepted work, complete effort, cash effects, and business outcomes. It should also say what you do not know. A decision made with visible uncertainty is more useful than a confident percentage built on incompatible comparisons.
NIST's measurement guidance emphasizes choosing methods and metrics for the intended use and its risks, including documenting risks that are not measured. The scorecard below applies that principle to an ordinary small-business workflow. Its example figures are constructed teaching data, not findings about a real company. NIST AI RMF Playbook: Measure
Define one unit of finished work
Choose a unit that a reader can recognize without seeing the tool: an accepted product description, a correctly classified invoice exception, or a customer request resolved within a stated period. Specify what counts as accepted, who decides, and when the result is assessed.
For a drafting workflow, acceptance might mean that a reviewer finds the answer accurate, complete for its stated scope, properly sourced, and ready for the next authorized step. This does not establish that the answer was sent, that the customer was satisfied, or that the underlying case was resolved.
Keep those later outcomes separate. An inspection-fee draft can be accepted while appointment availability remains unknown. Counting that draft as a resolved booking request would inflate performance by changing the meaning of success.
Write the definition before looking at the results. If you revise it later, keep the old definition, explain the reason, and avoid presenting the two periods as directly comparable without adjustment.
Keep the whole denominator in view
Use a flow of work that makes exclusions and unfinished cases visible. The Part 8 constructed adoption log contains twenty-five requests. Twenty are eligible for the proposed task: twelve take the assisted path and eight the manual path. Five are excluded. These are different groups from the earlier customer-service and training examples.
Of the twelve assisted tasks, ten pass first review and two need correction. All twelve eventually have final acceptance in the constructed log. No live customer sends are recorded, and the manual-path effort and acceptance results are not supplied.
| Measure | Calculation | What it tells you |
|---|---|---|
| Eligibility | 20 / 25 = 80% | Share of incoming work within scope |
| Assisted uptake | 12 / 20 = 60% | Share of eligible work taking that path |
| First-review acceptance | 10 / 12 = 83.33% | Share accepted without correction |
| Rework incidence | 2 / 12 = 16.67% | Share requiring correction |
| Final acceptance | 12 / 12 = 100% | Final status within this constructed assisted cohort |
The last percentage is not proof of universal quality. The sample is small and synthetic, its acceptance judgments are supplied as fixture data, and customer outcomes are unknown. Treat it as an example of correct accounting rather than a performance claim.
In a real pilot, add unfinished work and missing observations to the scorecard. Do not quietly remove them because they lack an attractive result. State the cutoff date and whether late completions will update the original cohort or appear in a separate follow-up.
Count every part of the work
The same adoption fixture records forty-eight minutes of assisted drafting, twenty-four minutes of review, twelve minutes of correction, and sixteen minutes of maintenance. That is one hundred minutes of total assisted effort for twelve finally accepted outputs.
Maintenance is counted once for the cohort. It should not disappear because no individual request owns it, and it should not be charged again in every row. For a live workflow, include other material work such as input preparation, escalation, data cleanup, administration, and exception handling.
The supplied comparison assumption is nine manual minutes for each of the same twelve tasks: 108 minutes. Subtracting the one hundred assisted minutes gives eight minutes of modeled capacity. It does not give measured time savings because the manual comparison is assumed, not observed.
The eight requests on the manual path do not supply that comparison. Their durations are blank. Treating them as a measured control group would invent evidence absent from the log.
A useful scorecard has a measurement-status column. “Recorded in a constructed fixture,” “assumed,” “observed during a pilot,” and “inferred” describe different strengths of evidence. Keep them visible beside the number.
Translate resource value and cash separately
For this chapter only, add two explicit teaching assumptions: internal time is valued at USD 30 per hour, and USD 6 of software cost is allocated to this same exercise period. These are not a vendor quote, a subscription recommendation, or an extension of the earlier finance forecast.
| Calculation | Modeled result |
|---|---|
| Assisted labor value: 100 / 60 × $30 | $50.00 |
| Allocated software cost | $6.00 |
| Specified assisted recurring resource cost | $56.00 |
| Cost per finally accepted assisted output: $56 / 12 | $4.67 |
| Assumed manual labor value: 108 / 60 × $30 | $54.00 |
| Provisional recurring resource benefit: $54 − $56 | −$2.00 |
The modeled eight minutes are worth four dollars at the assumed hourly value. Six dollars of allocated software cost more than consumes that value. A workflow can therefore create potential capacity and still have a negative modeled resource benefit.
This calculation excludes setup, initial training, and unspecified costs. Adding omitted costs could make the result worse. Do not call this a complete return-on-investment analysis.
For the separate cash view, assume no incremental receipts, no avoided cash expense, and no payroll reduction. With the six-dollar software allocation, the modeled incremental cash effect is negative six dollars. Internal labor value is not automatically a cash expense that disappears when a task gets shorter.
You also cannot compare cost per accepted output across both paths here: accepted quality for the assumed manual baseline is unknown. The fifty-four dollars is a task-effort comparator under an assumption, not a verified cost for twelve equivalently accepted manual outputs.
Connect the scorecard to the business outcome
Choose a small number of outcome measures appropriate to the workflow. Customer-service drafting might connect to correct resolution and avoidable recontacts. Commerce content might connect to profitable orders and returns attributable to misleading descriptions. Internal reporting might connect to decisions made on time with fewer corrections.
Explain the connection as a hypothesis until evidence supports it. Faster replies might help customers, but an incomplete fast reply can create another contact. More product descriptions might expand coverage, but traffic and demand may determine whether the additional pages produce revenue.
Record the observation window. An order can be placed today and returned later. A request can look resolved before the customer recontacts the business. A short pilot can establish that a workflow is usable while remaining unable to establish a durable financial effect.
Avoid forcing every task into revenue attribution. Some useful outcomes concern reliability, accessibility, or reduced operational burden. Define their value honestly and state what level of cost the business is willing to accept for that result.
Improve the comparison before improving the headline
For a stronger live evaluation, collect a comparable manual baseline using the same definition of finished work. Record task complexity, staff experience, workload, and the time period. Count review and corrections on both paths.
Where practical and appropriate, assign comparable eligible tasks between paths using a predefined method that reduces selection bias. Random assignment can help, but it does not repair inconsistent scoring, missing observations, or tasks moved between groups without explanation. For smaller pilots, a carefully documented comparison may be more feasible than a formal experiment.
Watch for changes unrelated to AI: seasonal demand, a new employee, a revised policy, or a different mix of customers. An improvement after rollout does not by itself establish that the tool caused it.
Keep individual staff evaluation separate from exploratory process measurement. Discuss what is recorded, who sees it, and how the results will be used. Otherwise people may avoid difficult cases or hide correction work, leaving a cleaner dashboard and a less reliable business.
Set decision rules before the review meeting
Use essential gates alongside economic tradeoffs. Access boundaries and action authority should not be averaged away by a favorable speed score. Quality requirements should reflect the consequences of error, with a clear response when a serious issue appears.
For a proposed pilot, record the acceptance definition, review capacity, budget scope, stop conditions, and evidence needed to continue. The companion decision-gate file leaves real owners, adopted thresholds, and observed results unresolved. Its examples do not approve a workflow.
A review can end with four useful decisions: continue at the current scope, improve and retest, expand under stated conditions, or stop. Document the reason and the next observation that could change the decision. “We bought it already” is not evidence that expansion is justified.
For the constructed Mesa numbers, a defensible conclusion is that the example demonstrates measurement mechanics and reveals a negative provisional resource result under its assumptions. It does not establish real employee performance, customer benefit, or an approved business case.
Publish only the claim your evidence supports
If you later describe results publicly, state the task, sample, period, comparison, and important limitations. Do not turn a drafting-time result into a claim about total business productivity. Do not present an internal assumption as a measured customer outcome.
For U.S. advertising, the FTC's substantiation policy applies to objective express and implied claims and requires an appropriate basis before dissemination. A precise-looking percentage does not supply that basis on its own. FTC Policy Statement Regarding Advertising Substantiation
Finish this chapter by completing one scorecard and writing a short decision statement: what is known, what remains uncertain, what the workflow costs within the stated scope, and what evidence is needed next. That statement should be understandable without opening the AI product. The next chapter asks whether the business can keep operating when that product becomes unavailable.