New: Boardroom MCP Engine!

People and leadership · Chapter 31 of 40

Use AI to Make Better Decisions Without Handing Over Judgment

Use AI to compare business options, distinguish facts from assumptions, test trade-offs, and document decisions without inventing evidence.

Updated

Open this chapter’s practice pack

On this page

You ask an assistant whether to automate customer replies. It produces a confident recommendation. You change the question to emphasize risk, and it recommends caution with equal confidence.

The problem is not simply that the answers differ. The decision's purpose, constraints, evidence, and trade-offs were never made explicit.

AI is useful when it helps you inspect a decision: identify missing information, compare alternatives, challenge assumptions, and write down why a choice makes sense. Your work product is a decision record that another person can review and that you can revisit after the outcome is known.

This chapter's Mesa proposal and scores are fictional planning material. They establish no purchase, rollout, staffing change, or approved customer action.

Turn the question into a decision someone owns

Replace “What should we do about AI?” with a bounded question. For example: “Which approach should Mesa investigate for routine service-policy replies over a thirty-day pilot?”

Name the decision owner, the deadline, the affected people, and the action the decision would authorize. Choosing an option for further investigation is different from approving its budget or turning it on.

Record what happens if you wait. Delay can have a cost, such as a continuing backlog. Acting can also have a cost, such as training effort or a new failure mode. Include both without inventing figures where evidence is missing.

Define success in the language of the business. A reply workflow should provide supported answers with manageable effort and a reliable path for exceptions. A higher model benchmark score is useful only if it helps meet those goals in the actual workflow.

Separate facts, assumptions, preferences, and unknowns

These four categories often get mixed together in a persuasive memo.

CategoryExample in the Mesa decision
Supplied factSERVICE-01 gives a $45 inspection fee and requires availability checking
Planning assumptionA proposed thirty-day budget limit is $100 in specified incremental cash costs
PreferenceThe decision owner gives reliability more weight than convenience
UnknownWhether the proposed setup will meet quality and access requirements in practice

A fact needs an appropriate source. An assumption needs a label and a way to check it. A preference belongs to the people responsible for the trade-off. An unknown should remain visible until evidence changes its status.

Keep dates and versions. If a supplier quote expires or a policy changes, a once-reasonable decision may need another review. A fluent summary without its supporting record makes that harder to see.

Treat a model's proposed explanation as a hypothesis until supported. “Staff will resist the tool” is not an employee finding. “The vendor will add export next month” is not a contractual commitment unless the actual evidence establishes it.

Give the current process a fair place in the comparison

The practice pack considers three approaches:

  • A: Improve the manual checklist. Keep manual drafting and make the current policy easier to find.
  • B: Test AI drafting with review. Use the approved sources to prepare drafts, with a person reviewing before customer use.
  • C: Send AI replies automatically. Remove the routine review step and let the system send responses.

These are different levels of authority, not merely three price plans. The current exercise does not approve automatic sending. Option C also lacks the evidence and recovery arrangements needed for that action, so it is held outside the scored shortlist.

A useful fourth possibility in real work may be to gather a specific missing fact before choosing. That is worthwhile when the information could change the decision and can be obtained in time. Research without a decision-changing question can become an expensive way to postpone a choice.

Do not make the manual alternative artificially weak. Better source organization might solve much of the problem with less new work. If AI still adds value after that improvement, the comparison is more informative.

Apply essential conditions before preferences

Some requirements cannot be traded away for a higher convenience score. Define these before ranking options.

For the teaching exercise, an approach must preserve required human review before customer use, remain within the proposed specified cash limit, and have an identified manual fallback. These are proposed pilot conditions, not declarations that a real system has passed security, privacy, accessibility, or operational checks.

Option C does not meet the review condition and has an unresolved fallback. Leave it on hold. A low price or impressive speed estimate cannot compensate for that failure by winning extra points elsewhere.

Keep real readiness separate from this paper comparison. Options A and B are eligible for discussion under the supplied planning assumptions. Neither is thereby approved to operate. Required people, workspace controls, source quality, and actual test results must still be established.

Use a scorecard to expose the trade-off

The UK Government Analysis Function's introduction to multi-criteria decision analysis explains how structured comparisons can support decisions involving competing objectives and how sensitivity analysis can test uncertain inputs. The simplified scorecard below is a discussion aid; it is not a fully elicited decision model or a mathematical proof of the best choice.

In the fictional exercise, reliability has weight 0.5, staff effort 0.3, and recovery 0.2. Each criterion uses a one-to-five preference scale, where higher is better. Staff effort therefore scores higher when the expected burden is lower.

Define the endpoints before assigning scores. Reliability ranges from frequent unsupported output to consistently supported output; effort ranges from substantial added work to little added burden; recovery ranges from difficult restoration to a clear, easy fallback. Intermediate scores express teaching judgments. They are not measured failure rates, minutes, or probabilities.

Eligible optionReliabilityStaff effortRecoveryWeighted score
A: Improved manual checklist4253.6
B: AI drafts with review4444.0

For A, the calculation is 4 × 0.5 + 2 × 0.3 + 5 × 0.2 = 3.6. For B, it is 4.0. The practice script checks the arithmetic; it does not verify that the assigned judgments describe a real tool.

This simple calculation treats the score steps as comparable within each scale and combines them additively. Those are simplifying assumptions. Do not turn a small difference into a claim of precise superiority, and do not count the same benefit under several overlapping criteria.

Keep money and omitted costs visible

The example assigns A a $30 initial cash cost and no additional monthly software cash. B has $60 in setup cash plus $30 in software cash during the thirty-day period, totaling $90. Both fit the proposed $100 cash limit.

These are invented planning amounts, not vendor prices. They exclude internal training and operating time, which must remain visible in a fuller resource assessment. The scorecard's staff-effort preference is not a substitute for Chapter 26's cost model.

C's assumed $60 cash total is irrelevant to its unresolved review and recovery conditions. It remains on hold. Avoid presenting a cheap but ineligible option as the winner and leaving the failed requirements in a footnote.

When adapting this record, specify the period, currency, included costs, and missing amounts. A budget check answers whether the specified cash fits the proposed limit. It does not establish affordability across the business or prove positive ROI.

Change the assumption most likely to reverse the result

Suppose B's review burden turns out to be higher than hoped. Reduce its staff-effort score from four to two while keeping the other teaching inputs fixed.

B's total becomes 4 × 0.5 + 2 × 0.3 + 4 × 0.2 = 3.4. A remains at 3.6, so the order reverses.

That reversal identifies a useful next investigation: measure the complete drafting and review effort in representative cases. It does not establish that B will perform poorly. The scenario shows that the initial preference depends on an uncertain input.

Chapter 30's constructed effort example is useful practice for designing that measurement, but its fictional minutes are not evidence validating this scorecard. Do not quietly substitute one teaching fixture for a real test.

Test other material uncertainties as needed: unsupported answers, source upkeep, an outage, low task volume, or a missing reviewer. Use plausible ranges supported by the available evidence. An arbitrary pessimistic number is a scenario assumption, not a forecast.

Ask for a challenge, not a vote

An assistant can argue for the leading option, then identify the strongest reasons it could fail. Ask it to point to source gaps and explain what evidence would change its recommendation.

Do not treat several answers from the same model as independent experts or convert their agreement into a probability. A generated panel may repeat the same unsupported assumption in different voices. It can help explore arguments, but it has not consulted actual employees, customers, or specialists.

Try a short failure exercise: imagine the pilot disappointed the team a month later. What could explain that? For Mesa, possible causes include outdated source files, review effort that erased the benefit, or questions drifting beyond the approved scope. These are hypotheses to investigate, not events that occurred.

Convert the most important concern into an action. Assign an owner to check policy freshness, include difficult cases in the evaluation, or test the fallback. Stop adding hypothetical risks when another round would not affect the decision or its next check.

Write the decision record before the outcome

Prepare a draft decision record from the supplied question, sources,
options, essential conditions, costs, and comparison scores.

Separate facts, assumptions, preferences, and unknowns. Explain excluded
options before ranking the eligible ones. Show which changed assumptions
reverse the ranking and identify the next useful evidence to collect.

Preserve disagreement and unresolved conditions. State the proposed action,
decision owner, scope, review date, and reasons to revisit it.
Do not invent approvals, employee views, test results, probabilities,
vendor commitments, or executed actions.

In the supplied case, a defensible draft recommends investigating B's effort and quality under a controlled evaluation while retaining A as the alternative. It records that the preference reverses under the higher-effort scenario. The actual decision, approver, execution date, and outcomes remain blank.

For a real decision, record who approved what and when. Keep any dissent with its reason. A reviewer who disagrees about effort assumptions has identified a testable issue; the record should not erase it to make the recommendation look unanimous.

Learn without rewriting the original reasoning

At the review date, compare the result with the original expectations and evidence. Separate what was knowable at the time from what became clear later. A good result can follow weak reasoning, and a careful choice can still encounter an unfavorable outcome.

Update the relevant assumption or process. If review time was underestimated, improve the next measurement. If a source owner never maintained the files, address responsibility. If the task did not need AI after all, retain the simpler process and the lesson.

The decision record earns its value when it makes both the choice and the learning visible. AI helps organize that work; the people with responsibility still determine the goals, inspect the evidence, and authorize the action.

Next: Part 9 turns approved workflow ideas into reliable automations, with permissions, recovery, and clear limits on agent actions.