Short answer
Do not scale an AI pilot because the demo works or staff like it. Scale only when five tests pass: the business problem is named, inputs are repeatable, answer quality is acceptable for the use, a human owner can handle exceptions, and the full cost is visible at the expected workload. Otherwise stop or redesign it.
Your AI demo works. Staff like it. The supplier is asking for a larger budget. That is not yet a case for scale.
For a Kampala operations manager, the right question is more demanding: can this pilot solve a named business problem, with repeatable inputs, acceptable answers, accountable human handling and a cost we can see at the expected workload? If the answer is no, stop or redesign the pilot. Do not scale it because the interface is impressive.
A pilot can fail usefully. It may show that the source data is inconsistent, the approval path is unclear, or the measure chosen at the start is not affected by the tool. That is not wasted work if the organisation records what it learned and changes the next decision. The waste begins when a weak pilot becomes a permanent expense because nobody has agreed how to stop it.
Test 1: Is the business problem named?
“We want to use AI” is not a business problem. A usable pilot starts with a sentence that an operations lead can test:
We are testing whether [tool or workflow] can reduce [specific delay, error, rework or unanswered demand] for [defined process] without increasing [named risk or workload].
Add a baseline and a decision owner. For example, the team might be investigating whether an assistant can reduce the time spent classifying inbound service requests, while keeping escalation errors within an agreed limit. The baseline could be the current handling time and the number of cases requiring correction. The owner is the person who can approve a process change, not only the person who arranged the demo.
If the team cannot name the process, measure and owner, the correct result is stop. A broad promise such as “make the office more productive” cannot tell you whether the pilot deserves another shilling.
The Enterprise AI Strategy section of AI Business Magazine - August 2026 makes a useful distinction: individual employees may complete tasks faster while the organisation sees no corresponding result. A pilot must connect personal convenience to a business measure. If the target measure does not move, staff enthusiasm is evidence of usability, not proof of value.
Test 2: Are the inputs repeatable?
An AI output is only as dependable as the material and instructions supplied to it. Check whether the pilot receives the same kind of input each time, with a known owner and a clear update point.
- What files, records or messages enter the workflow?
- Who creates and approves those inputs?
- What minimum fields or information are required?
- How are missing, duplicated or outdated information handled?
- Where did the source come from, and when was it last checked?
- What happens when the expected input is unavailable?
A demo using a carefully selected folder may look excellent while the live process receives incomplete spreadsheets, scanned documents and inconsistent names. That is a data and process finding, not necessarily a model finding. Redesign the intake, definitions or source record before asking the tool to do more.
The data-foundation literature treats use-case requirements, governance, metadata and operations as part of the system being evaluated. In a small organisation, that does not require a new data platform. It does require somebody to own the input, a repeatable preparation step and a visible rule for exceptions.
Test 3: Is the answer quality acceptable for this use?
“It sounds right” is not a quality test. Build a small review set from real, redacted cases. Include ordinary requests, ambiguous wording, missing information, outdated material and questions that should be refused or escalated.
For each case, record the expected answer or acceptable range, the source the answer should rely on, the error that would matter most, whether a citation is required, the correct human handoff and the time a reviewer needs to check the result.
Quality is use-specific. A draft-summary pilot may tolerate wording changes but not an omitted action. A stock-reorder suggestion may tolerate a conservative recommendation but not an unexplained unit error. A customer-facing answer may need a clear source and an immediate handoff when the question is outside scope.

Test the tool again after changing the prompt, source file, model or workflow. The aim is not to produce a flattering accuracy percentage. It is to know which errors the organisation can accept, which it cannot, and how the system behaves when it lacks the information needed to answer.
Test 4: Is a human owner ready for exceptions?
Every pilot has cases it cannot resolve. A safe pilot makes the boundary visible before the first live user encounters it.
Name the person or role that will receive an escalated case, decide whether the answer can be corrected or must be rejected, communicate the outcome, record a recurring failure and approve a change to the instructions, source material or workflow.
Do not confuse “a human is somewhere in the loop” with actual accountability. The reviewer needs time, access to the source, authority to correct the output and a route for urgent issues. If the pilot increases review work beyond what the team can absorb, record that as an operating cost and redesign the process.
Chapter 8 of Designing the AI-Driven Data Foundations emphasises explicit decision rights, data ownership, escalation paths and controls that match the use case’s impact. That principle is useful at pilot scale: name the owner before asking the tool to act.
Test 5: Is the full cost visible at the expected workload?
Do not compare only the subscription or API price with the old process. Cost includes the work required to make the pilot usable and safe:
- licence, API or hosting charges;
- data preparation and cleanup;
- integration, configuration and monitoring;
- human review and exception handling;
- training and change-management time;
- support, security and access administration;
- rework caused by poor answers; and
- the cost of keeping the old process available as a fallback.
Then test the expected workload, not only a quiet demonstration. A workflow that is affordable for 20 cases may need a different control or budget for 2,000. Set a usage limit and an alert before the pilot starts. Chapter 10’s observability and operations emphasis is practical here: monitor what the system is doing, what it costs and when its inputs or outputs change.
When staff love the tool but the measure does not move
This is the uncomfortable result many teams avoid. Staff may genuinely enjoy a tool because it drafts faster, reduces tedious searching or makes a difficult task feel lighter. Those are valid user signals. They do not prove that the business outcome has improved.
Ask whether the target measure had enough time and volume to move, whether the pilot changed the whole process or only one task inside it, and whether the measure was the right one for the organisation’s actual constraint.
If the measure is wrong, redesign the test and document the change. If the measure is right and does not move after a fair test, stop the pilot or keep it only as a small discretionary tool with an honest budget. Do not rename convenience as return on investment.
The one-page decision table
Use this table at the end of the pilot. A mixed result is normal; the outcome should follow the weakest control that matters to the risk of the use case.
| Outcome | Use it when | Required action before the next decision |
|---|---|---|
| Stop | No named problem or owner; inputs are not repeatable; critical errors have no safe handoff; cost is unknown or unacceptable; or the target measure does not move after a fair test. | Switch off the pilot, preserve the findings, return data or access as agreed, and record what would need to change before a future test. |
| Redesign | The problem is real, but data, workflow, answer-quality thresholds, human capacity or measurement is not ready. | State the repair hypothesis, assign an owner, set a new test set and deadline, and approve only the limited work needed to retest. |
| Scale | The problem and measure are clear; inputs are repeatable; quality and escalation meet the agreed threshold; a human owner is resourced; and full cost is visible at expected workload. | Scale in stages, keep monitoring and a rollback path, and review the decision after the first agreed operating period. |
Do not treat “redesign” as a softer word for “scale”. A redesign should have a bounded repair, a new test and a new decision date. If those are missing, the organisation is only extending the pilot.
End with a short pilot-review meeting
Keep the meeting to 45 minutes and invite the process owner, a working user, the data or system owner, the finance or budget owner, and the person accountable for the outcome.

- 5 minutes — decision and scope: restate the business problem, target measure, pilot period and expected workload.
- 10 minutes — evidence: review the baseline, cases tested, answer quality, escalations, corrections and user experience.
- 10 minutes — operating reality: review input repeatability, human capacity, security or access issues, and support effort.
- 10 minutes — cost: compare actual and expected full cost, including review time and fallback work.
- 5 minutes — decision table: choose stop, redesign or scale; name the owner and deadline.
- 5 minutes — record: capture the evidence, unresolved questions, approved next action and rollback or shutdown steps.
The best outcome of a pilot is not always a larger deployment. Sometimes it is a clear stop that protects the budget. Sometimes it is a smaller, better-designed test. Scale only when the evidence earns it.
Frequently asked questions
When should an organisation stop an AI pilot?
Stop when the business problem or owner is unclear, inputs are not repeatable, critical errors have no safe handoff, the full cost is unacceptable or unknown, or the target measure does not move after a fair test. Preserve the findings so the organisation does not repeat the same experiment without learning from it.
Does staff enthusiasm prove that an AI pilot is valuable?
No. Enthusiasm is useful evidence about usability and perceived convenience. It is not proof that the business outcome improved. Compare the pilot with the agreed measure, workload, review effort and cost.
What is the difference between redesigning and scaling a pilot?
Redesigning means repairing a bounded weakness—such as inconsistent inputs, an unclear handoff or a poor measure—then running a new test by a named date. Scaling expands the operating use only after the problem, inputs, quality, accountability and cost have passed the agreed gate.
How much testing is enough before a pilot decision?
There is no universal number. Use enough real, redacted cases to cover ordinary requests, ambiguity, missing information, outdated material and cases that should be refused or escalated. The test set must match the risk and volume of the intended use.
Sources & the researchers worth crediting
Primary working references: AI Business Magazine - August 2026, Enterprise AI Strategy and Leadership sections; and Designing the AI-Driven Data Foundations: Architecture, Principles, and Practice (Sanjeev Mohan, 2026), especially chapters 4, 8 and 10. This article distils their decision and operating principles without reproducing their text. Check current tool terms, costs, security controls and workload requirements before approval.
Read next
Measuring AI ROI Honestly
Separate AI cost from value when deciding what to keep, fix or switch off.
Before You Buy AI, Clean Your Business Data
Improve the records and definitions underneath an AI workflow.
Responsible AI Without Digital Dependence
Set practical boundaries around data, contracts, capability and accountability.
About the author
Peter Bamuhigire
Software architect and ICT consultant — business management systems across Africa
Peter Bamuhigire helps owners and operations leads turn AI experiments into controlled business decisions. His approach connects the tool to the process, the evidence, the people who must review it and the cost of keeping it running.

