A new advertisement gets more clicks. A landing page produces more inquiries. The dashboard turns green. Has the business improved?
That depends on what happened next, what the campaign cost, and whether the comparison was fair. AI makes it easy to produce variants and persuasive explanations of results. A useful experiment needs a written decision rule before those results arrive.
Your work product is a campaign experiment brief and a decision memo. The numbers in this chapter are constructed teaching data. No campaign ran, no visitors were observed, and no conversion improvement was measured for Mesa or Salars.net.
Define the decision and the offer
Choose one question that affects a real decision. For fictional Mesa, the question is whether a clearer explanation of the inspection fee should replace the current wording beside the request form.
The service itself remains the same: a $45 inspection fee, credited toward a repair approved within 30 days of inspection, with booking confirmation after an availability check. Both variants must preserve those facts. The experiment is about explaining the existing offer, not quietly adding a discount or guarantee.
A proposed hypothesis is:
Explaining the fee credit more clearly may increase the proportion of eligible visitors who submit a qualified inspection request, without increasing fee misunderstandings or creating more low-quality follow-up work.
Write what would change your decision. If the new version produces more requests but many readers incorrectly believe repairs are free, it has failed an important part of the task.
Pick the outcome before generating variants
A primary metric is the main measure selected for the decision. A guardrail is an additional condition that prevents a narrow improvement from hiding unacceptable damage. Diagnostic metrics help explain what happened but do not automatically decide the winner.
For this proposed exercise:
| Type | Definition |
|---|---|
| Primary | Distinct eligible visitors with at least one human-qualified request within seven days of first assignment, divided by all eligible assigned visitors |
| Qualification | Genuine inspection request, supported service match, usable permitted contact route, and human review under the same written rule |
| Guardrails | Misleading fee expectations, complaints, contact-scope violations, and unmanageable follow-up demand |
| Downstream | Completed paid jobs within thirty days of first assignment and contribution after stated campaign and follow-up costs |
| Diagnostic | Form starts, raw submissions, and clarification questions |
The seven- and thirty-day windows are fictional planning choices. A business must choose windows that fit its actual buying cycle and record later outcomes separately. Do not compare a fully matured group with another whose customers have not yet had time to act.
Count the denominator consistently. People assigned to a variant who leave without submitting still belong in the primary denominator. Counting only people who finish the form would answer a different question and could conceal abandonment.
Make the comparison fair
An A/B test compares two versions. In a randomized design, experimental units are assigned to conditions randomly; this is a basic design principle described in NIST's guidance on completely randomized designs.
For a website experiment, define the unit you can reliably assign and recognize. It may be an eligible visitor ID; in business sales, an account may be a more appropriate unit when several colleagues influence the same purchase. Document identification limitations, including people returning on different devices.
Run versions concurrently with a stable assignment rule. Avoid showing one person a different offer on each page load. Keep other relevant conditions consistent, and record unavoidable changes such as a service interruption or a concurrent advertisement. Comparing last month's email audience with this month's search audience would mix the copy change with audience and timing differences.
Before launch, define exclusions such as internal test traffic or identified bots, along with how they are detected. Apply the same rules to both groups. Do not remove inconvenient non-buyers after looking at the results.
Specify sample size and duration using the baseline rate, the smallest improvement worth detecting, the available traffic, and the analysis method. There is no universal rule that every small business test should run for seven days or collect 100 responses. When traffic cannot support a credible quantitative comparison, use comprehension research or usability observation for a narrower decision and state the limit.
Use AI to prepare an experiment you can review
Ask for a limited number of variants with a clear rationale:
Create two candidate explanations of the approved inspection-fee policy.
Keep the price, credit condition, booking process, page layout,
and next action unchanged. Vary only how the fee explanation is worded.
For each candidate:
- show the wording;
- identify the customer question it addresses;
- explain the hypothesis, clearly labeled as untested;
- identify possible misunderstandings.
Do not introduce urgency, discounts, reviews, guarantees, new services,
or performance claims. Do not predict a conversion lift.
Use the claim ledger from Chapter 15 to review both candidates. Make sure the form and confirmation step still work. Check tracking with synthetic events before real traffic enters, and confirm that the assignment and outcome definitions are implemented as written.
Name the person authorized to approve publication, spending, and any offer changes. A generated plan does not authorize a budget. The practice packet contains an unapproved experiment plan so it cannot be mistaken for a launched campaign.
Read the example without overstating it
Suppose the following summary is supplied for arithmetic practice. The counts are fictional, and the underlying assignment and event logs are not present.
| Measure | A | B |
|---|---|---|
| Eligible assigned visitors | 200 | 200 |
| Visitors with a qualified request | 10 | 14 |
| Qualified-request rate | 5% | 7% |
| Completed paid jobs in the stated window | 8 | 7 |
| Assumed contribution per completed job before campaign/follow-up costs | $90 | $90 |
| Campaign cost allocated to the group | $200 | $200 |
| Follow-up labor cost allocated to the group | $60 | $90 |
The difference in qualified-request rates is two percentage points: 7% minus 5%. Relative to A's 5%, that is a 40% increase in this constructed table. “Up 40%” without the counts and base rate would make the result harder to judge.
Those arithmetic facts do not prove a reliable causal lift. We have no verified randomization, event-level quality checks, approved sample plan, or inferential result in this exercise. Small changes in these counts could alter the story. Do not ask the model to declare statistical significance from a persuasive narrative.
The downstream calculation is also revealing. Define the $90 assumption as revenue less job-level variable costs, before the separately listed campaign and follow-up costs. Do not deduct those same costs twice.
- A: 8 × $90 − $200 − $60 = $460.
- B: 7 × $90 − $200 − $90 = $340.
B has more qualified requests in the table but $120 less contribution after these specified costs. The example does not establish that A will remain better, either. Later jobs, refunds, and other missing costs could change the comparison.
These amounts exclude fixed overhead, initial content production, and any other unlisted costs. They are not net profit, cash flow, or a full campaign ROI calculation. Labor cost allocation may represent staff capacity rather than an actual reduction in payroll. Use explicit definitions so the decision maker knows what has been counted.
Use uncertainty to guide the next action
For a real randomized test, analyze the prespecified metric with a method appropriate to the assignment unit and design. Report an uncertainty interval where appropriate, alongside absolute counts and the smallest effect that matters economically. A statistical result and a worthwhile business result answer different questions.
If the team lacks the needed analytical experience, get help choosing the method before launch. AI can explain a calculation or prepare code, but the reviewer must verify assumptions, denominators, and whether the data supports the method. Never accept a p-value merely because it has several decimal places.
Repeatedly checking results and stopping the first time one version looks better can undermine a conventional fixed-sample analysis. Choose an appropriate stopping rule in advance. Stop promptly for an operational problem or misleading offer, but record that interruption and its consequences for interpretation.
Avoid launching ten variants and reporting only the most flattering pair. Record every planned comparison, failed launch, material change, and excluded period. That history helps explain why the final evidence is or is not persuasive.
Write the decision memo in plain language
A defensible memo for the constructed table would say:
B has 14 qualified requests from 200 assigned visitors, compared with 10 from 200 for A. The descriptive difference is two percentage points. In the supplied downstream window, A has eight completed paid jobs and B has seven. Under the stated cost assumptions, contribution after campaign and follow-up costs is $460 for A and $340 for B. These are fictional practice calculations. They do not justify selecting a real campaign winner or claiming a measured lift.
For a real test, add tracking checks, uncertainty, guardrail outcomes, missing data, and the actual decision owner. If the evidence is inconclusive, say what decision can still be made. You might fix wording that causes a demonstrated misunderstanding while withholding any claim about conversion improvement.
Keep a learning log with the original hypothesis, approved variants, dates, definitions, costs, results, and decision. Include experiments that did not improve the outcome. Over time, that record helps the business avoid recycling unsuccessful ideas under new AI-generated wording.
Next: Part 5 moves into Customer Service, beginning with how to triage questions and draft answers from approved sources.