Ready to put this into action?
Get the complete AI Integration Playbook β Practical AI implementation guide β prompt engineering, workflow automation, and ROI frameworks.
Article 063 Β· Part 7
Find the AI Opportunities Worth Your Time
Choose a small project whose value you can measure after review and correction.
By Randy Salars Β· Published
On this page
- Map the work before scoring the idea
- Compare AI with ordinary improvements
- Calculate total effort
- Test the assumption that could reverse the answer
- Measure your task rather than borrowing a headline
- Write a bounded pilot proposal
- For students: evaluate a club or class workflow
- Practice: choose one pilot from five tasks
Choose a small project whose value you can measure after review and correction.
A director asks every department to find an AI use case. Soon the organization has a chatbot proposal, three subscriptions, and a list of ambitious automation ideas. Nobody has measured how long the original work takes.
Without that baseline, the organization can demonstrate activity but cannot tell whether the effort improved anything. A faster first draft may create a slower review. A polished summary may omit the exception that mattered most.
Start with recurring work and a specific problem. The first useful question is: where does effort go, and what would a better result look like?
Map the work before scoring the idea
Choose a process you can observe from start to finish. Record its trigger, frequency, inputs, outputs, people, active work time, waiting time, and recurring errors. Include the person who receives the output; saving the creator five minutes while adding ten minutes downstream is not an improvement.
For a fictional community organization, five candidate tasks might be preparing program descriptions, extracting room requests, calculating expense totals, drafting routine replies, and selecting grant recipients. These tasks differ in both technical suitability and consequences.
Distinguish assistance from action. Drafting a response leaves a review step. Sending it commits the organization to what it says. Extracting a requested room does not reserve the room. Summarizing grant applications does not establish a fair selection process.
Use this prompt to begin:
Analyze these five recurring workflows. For each, identify the actual problem, possible AI assistance, a simpler alternative, required data, verification effort, and the consequence of an error. Separate drafting and extraction from decisions and external actions. Identify missing baseline information before estimating savings.
Workflow observations: [insert authorized notes].
The assistant can organize observations. It cannot supply measurements you never collected.
Compare AI with ordinary improvements
Some workflows are difficult because their inputs are inconsistent. A better intake form can prevent ambiguity before an AI system has to interpret it. Other workflows are repetitive and exact, making a template, spreadsheet formula, or ordinary rule easier to inspect.
For the fictional organization, a qualitative screening table could look like this:
| Task | Plausible assistance | Simpler alternative | Initial decision |
|---|---|---|---|
| Program descriptions | Draft from approved facts | Reusable description template | Compare both |
| Room requests | Extract fields from messages | Required booking form | Fix intake first |
| Expense totals | Explain anomalies | Spreadsheet formulas | Use formulas for totals |
| Routine service replies | Draft from approved policy | Saved response templates | Candidate pilot |
| Grant selection | Organize authorized materials | Existing reviewed process | Exclude from first pilot |
These are judgments for this invented setting, not universal ratings. A different organization may have different records, risks, or bottlenecks.
If you want a score, define its meaning. βReview burden: 1β might mean a short, objective field comparison; β5β might mean expert judgment across many sources. Do not multiply invented scores into an impressive decimal and mistake it for financial evidence. Use scoring to explain a choice, then test the choice.
Calculate total effort
Suppose the fictional organization handles 120 routine inquiries per month. The baseline is eight minutes per completed reply, including its normal checks: 960 minutes, or 16 hours.
The proposed assisted workflow requires two minutes to prepare an AI draft and three minutes to review every draft. Assume 20% of drafts require six additional minutes of correction. These are planning assumptions, not observed results.
| Monthly component | Calculation | Minutes |
|---|---|---|
| Draft preparation | 120 Γ 2 | 240 |
| Review | 120 Γ 3 | 360 |
| Extra correction | 120 Γ 20% Γ 6 | 144 |
| Total assisted handling | 240 + 360 + 144 | 744 |
| Handling time saved | 960 β 744 | 216 |
The apparent saving is 3.6 hours per month. If maintenance takes another hour monthly, the net capacity gain becomes 2.6 hours. If setup takes six hours, recovering that setup effort requires about 2.3 months at the assumed steady workload, excluding subscription expense.
At a hypothetical labor valuation of $30 per hour, 2.6 hours represents $78 of monthly capacity value. Subtracting a hypothetical $25 monthly tool cost leaves $53. This is a planning valuation, not necessarily money removed from payroll or cash profit. Someone must be able to use the released time productively.
Show which costs are included. Training, security review, integration work, and expert supervision may change the case. Avoid claiming a complete return on investment from a calculation that omits material costs.
Test the assumption that could reverse the answer
In the example, correction frequency matters. If 50% of drafts need the extra six minutes, assisted handling becomes 960 minutesβthe entire original workload. Maintenance then makes the new process slower overall.
Including 60 minutes of maintenance, monthly effort is:
600 + 720r + 60 minutes, where r is the fraction of drafts needing correction.
To remain below the 960-minute baseline, r must be below approximately 41.7%. This threshold does not include initial setup or tool expense. It simply shows how sensitive the time case is to correction burden.
Ask AI to identify such turning points, then verify its arithmetic independently. The strongest proposal is not the one with the largest optimistic number. It is the one that explains which observations could invalidate it.
Measure your task rather than borrowing a headline
Research can help challenge assumptions, but its setting matters. METRβs early-2025 randomized study involved 16 experienced open-source developers and 246 tasks in familiar projects; allowing the tools studied increased completion time by 19%. That finding does not establish the effect on routine customer replies or current tools. METRβs early-2025 study.
METRβs February 2026 update described later data as an unreliable signal of current productivity effects because of selection and measurement problems. The researchers believed speedups had likely increased, while emphasizing weak evidence about their size. METRβs study-design update.
For your pilot, record completed-work time, correction work, quality, and unresolved cases. If people multitask while a tool runs, distinguish elapsed time from active labor. Comparing eight minutes of active manual work with eight minutes of mostly unattended AI processing answers a different question from comparing total turnaround.
Write a bounded pilot proposal
The fictional reply pilot can fit on one page:
- Scope: Routine factual inquiries covered by the approved service guide; exclude exceptions and personal account changes.
- Owner: The service coordinator, with authority to stop the pilot.
- Comparison: Current saved templates versus AI drafts, both reviewed to the same standard.
- Data: Authorized inquiry text and the current guide, with unnecessary personal details removed.
- Duration: Four weeks, with case counts and workload mix reported.
- Measures: Active handling time, corrected factual errors, unauthorized promises, unresolved inquiries, and reviewer burden.
- Decision: Continue only if quality meets the agreed standard and net effort falls enough to justify maintenance and cost.
Set thresholds before seeing results. For this example, an unauthorized policy promise is a reason to pause and diagnose the workflow. A proposed continuation threshold might require at least 20% lower active handling time with no deterioration in reviewed correctness. Those are local choices, not industry standards.
Use comparable cases and avoid assigning all easy inquiries to AI. Keep difficult cases visible. If feasible, allocate eligible cases between approaches in advance and record deviations. A small operational pilot informs a local decision; it rarely proves a general effect across an entire organization.
For students: evaluate a club or class workflow
A student organization might test drafting event descriptions from a fixed fact sheet. Measure how long it takes to produce a correct description manually, with a template, and with AI. Include checking and repairs in each time.
For coursework, use synthetic records if real member data is unnecessary or unauthorized. Follow the courseβs AI rules. Your report should distinguish a forecast from a completed pilot and explain what you would measure next.
A result showing little benefit is still useful. You have learned when a tool is a poor fit and how to justify that conclusion with evidence.
Practice: choose one pilot from five tasks
List five recurring tasks and collect a small baseline for each promising candidate. Compare AI with at least one simpler alternative. Write a one-page proposal that includes the owner, boundaries, measurement approach, and continue/revise/stop conditions.
Calculate a favorable scenario and an unfavorable scenario. Explain which assumption most changes the decision. If you have not run the pilot, label all savings as forecasts.
Completion check: Your chosen task has an observable baseline, an accountable owner, a practical review method, and a decision rule. The business case includes correction and maintenance effort, and a reader can reproduce the calculations.
For a stretch exercise, include downstream work and compare the value of freed capacity with actual cash savings. They may support different decisions.
Get the AI Dispatch
Weekly insights on ai & technology β delivered to your inbox. No spam, unsubscribe any time.
Want to choose specific topics? Customize your interests