Short answer
Measure an AI feature like any other investment: define the outcome, set the pre-AI baseline, count every recurring and usage-based cost, and compare the result with a credible counterfactual. Keep realised cash savings separate from capacity freed, quality improvement, and future option value. Keep, fix, or cut feature by feature on that evidence.
The CFO’s question is not whether the company is using AI. It is whether each use case earns the right to keep consuming money, attention, data, and risk capacity.
That discipline matters because adoption and return are not the same thing. McKinsey’s State of AI 2025 reported that 88% of respondents regularly used AI in at least one business function, while only 39% reported an enterprise-level EBIT impact. IBM’s CEO study found that only 25% of AI initiatives had delivered the expected return and only 16% had scaled across the enterprise. Those are different surveys with different samples, but they point to the same finance problem: activity is easier to count than value.
A 2026 WRITER survey also reported that 75% of executives admitted their AI strategy was “more for show” than an actual guide to action. Treat that figure as a survey finding, not a universal measurement. The practical lesson is stronger than the headline: a budget line labelled AI is not a business case.
Start with the decision, not the model
Before asking for a return percentage, write one sentence that names what the use case is meant to change. “Use AI to improve productivity” is too broad. “Reduce the average time to classify a supplier invoice from eight minutes to three without increasing exception errors” can be measured.
Give every AI use case four owners:
- Business owner: accountable for the outcome and adoption.
- Finance owner: responsible for the baseline, cost model, and treatment of benefits.
- Process owner: responsible for the actual workflow and human review.
- Risk owner: responsible for privacy, security, quality, regulatory, and reputational guardrails.
If one person carries all four roles, write that down too. A small team can be perfectly sensible; an unnamed owner is not.

Build the full AI cost stack
Most weak AI business cases use a licence price as the cost. The finance view needs the cost of producing the result repeatedly and safely.
Full AI cost = fixed platform cost + usage cost + people cost + integration cost + control cost + failure cost
Use the same period for all six components. A monthly feature review should use monthly costs and a monthly volume denominator.
1. Fixed platform cost
Include subscriptions, minimum commitments, model access, vector or database services, observability tools, and reserved infrastructure. A free trial is not a steady-state cost assumption.
2. Usage cost
Count the events that trigger spend: model calls, tokens, documents, pages, images, audio minutes, retrieval queries, workflow runs, and retries. Frequency matters. An automation that costs a few cents per call can become a material monthly line when it runs on every customer message, invoice, stock movement, or employee request.
3. People cost
Include implementation, prompt and workflow design, data preparation, review time, exception handling, training, support, finance analysis, and change management. The human review queue is part of the operating model, not proof that the AI is free.
4. Integration and data cost
Allow for APIs, connectors, data cleaning, retrieval indexing, exports, backups, and the maintenance of permissions. If the feature needs a new master-data process, cost that work instead of treating it as an invisible prerequisite.
5. Control cost
Privacy reviews, security testing, audit logs, model evaluation, red-team testing, contractual review, retention rules, and incident response all protect value. A use case that cannot pass the organisation’s control threshold is not a bargain.
6. Failure cost
Estimate rework, incorrect outputs, escalation, customer remediation, downtime, and opportunities lost when people stop trusting the feature. You will not know this number perfectly at launch. Track it explicitly and improve the estimate with incidents.

Separate four kinds of value
“Value” is not one number. Put the benefits into separate columns so a capacity gain does not quietly become a cash saving and then a growth claim.
- Realised cash: a supplier cost removed, overtime avoided, contractor hours not renewed, or revenue collected sooner. Show the mechanism and the period.
- Capacity released: hours returned to a team. Show what work absorbed the time: more cases handled, faster close, better follow-up, or simply lower pressure. Do not call it a payroll saving unless payroll actually changes.
- Quality and risk: fewer errors, faster detection, fewer compliance exceptions, safer decisions, or a more complete audit trail. Convert this to money only when the conversion is defensible.
- Strategic option value: data, learning, or a reusable capability that enables a later decision. This can matter, but it is not a reason to hide a negative current return.
IBM’s 2025 CEO research found that 52% of respondents saw AI value beyond cost reduction, including revenue growth, customer experience, innovation, and sustainability. That is a useful reminder for CFOs: cost takeout is not the whole case. It is also a reason to name the kind of value being claimed before a team counts it.
Use a baseline and a counterfactual
A before-and-after comparison is a start, not proof. Volume, staffing, seasonality, pricing, customer mix, and process changes may have moved at the same time as the AI launch.
For each use case, record:
- the baseline period and the exact metric definition;
- normal volume, mix, and service-level conditions;
- the comparison group, old process, or sampled control period;
- what changed besides AI;
- the quality and safety thresholds that cannot be traded for speed;
- the person who signs off the result.
Use the strongest practical comparison available. A randomised test may be suitable for a low-risk message or workflow change. A matched branch, team, queue, or time period may be more realistic. For high-risk finance, health, credit, employment, or customer decisions, keep human approval and report the control performance alongside the productivity result.

Attribute the outcome honestly
Give every benefit an attribution confidence level:
- High: the AI change is isolated, the metric has a stable baseline, and the cash effect is visible in the accounts or an agreed operational record.
- Medium: the direction is clear, but several process or volume changes happened together. Report a range and name the assumptions.
- Low: the claim is plausible but there is no reliable baseline, comparison, or adoption evidence. Treat it as a learning result, not an ROI result.
This protects the CFO from two opposite errors. The first is approving weak tools because every benefit is claimed at 100% attribution. The second is cutting useful capability because finance refuses to recognise anything that cannot be reduced to a single invoice line. Good measurement states what is known, what is estimated, and what is still a hypothesis.
Decide what to keep, fix, or cut
Review each feature at a fixed cadence with the same scorecard. The decision is not “Does the demo look impressive?” It is “Does this version of the workflow earn another period of spend?”
| Decision | Evidence | Next action |
|---|---|---|
| Keep | Positive realised value or a well-supported strategic benefit; guardrails pass; users adopt it. | Set the next target and monitor cost per useful outcome. |
| Fix | The problem matters, but adoption, data quality, workflow design, cost, or accuracy is holding the feature back. | Name one corrective experiment, owner, budget, and review date. |
| Cut | Full cost exceeds credible value, guardrails fail, or the use case has no owner or measurable outcome. | Switch it off safely, document why, and retain reusable learning. |
“Fix” should not become a permanent waiting room. Put a time limit on the next experiment. If the use case still cannot show a credible path to value after the data, workflow, and adoption problems are addressed, make the cut decision.
A 30-day finance review for an AI feature
- Days 1–5: name the outcome, owners, baseline, volume, guardrails, and decision date.
- Days 6–10: list fixed, variable, people, integration, control, and failure costs. Confirm how usage frequency will be counted.
- Days 11–20: measure live work against the agreed comparator. Record exceptions and what staff actually do with time released.
- Days 21–25: finance reviews attribution, cash treatment, capacity treatment, and the quality or risk result.
- Days 26–30: decide keep, fix, or cut. Write the decision in the investment register and set the next review date.
The aim is not to turn every experiment into a quarterly reporting burden. The aim is to make the burden proportionate to the risk and spend. A small internal assistant needs a lighter review than an AI system that approves credit, changes prices, or handles personal data. The discipline is the same: named outcome, full cost, credible comparison, accountable decision.
AI is not exempt from capital discipline
AI can create real value. It can also create a recurring bill for calls, reviews, integrations, and controls before anyone has shown what changed. IBM’s research found that many CEOs were investing ahead of clear value, while most said clearer success metrics would help. That is a finance invitation, not a reason to pause every experiment.
Keep the features that improve an important outcome at an acceptable full cost. Fix the features with a named problem and a time-boxed test. Cut the features that cannot clear the value, safety, adoption, or measurement threshold. A CFO’s contribution to AI is not scepticism for its own sake. It is the evidence that lets the organisation do more of the right work and stop paying for the rest.
For a practical starting point, pair this framework with responsible AI governance and the guide to agentic AI for business owners. If you need to turn one use case into a measured business case, contact Peter.
Frequently asked questions
What is the best way to measure AI ROI?
Start with a named business outcome, a pre-AI baseline, the full cost of running the use case, and a credible comparison period. Report realised cash impact separately from capacity freed, quality or risk improvement, and future option value. If the baseline or ownership is missing, call the result directional rather than proven.
Which AI costs do CFOs most often miss?
The commonly missed items are usage frequency, model and API calls, retrieval or OCR charges, storage, human review, integration maintenance, data cleaning, security controls, training, support, and the cost of failures or rework. A feature that is cheap per call can be expensive when it runs thousands of times a month.
Should time saved count as AI ROI?
Yes, but label it as capacity released unless it becomes a cash saving or supports measurable additional output. Ten hours saved is not automatically ten hours of payroll reduction. Record what the team did with the time and report the cash, throughput, quality, or service effect separately.
How long should an AI ROI pilot run?
Run it long enough to cover normal volume and at least one meaningful operating cycle. A finance workflow may need a month-end close; a customer-service workflow may need several weeks across busy and quiet periods. Agree the measurement window before launch and do not change the baseline after seeing the result.
When should a company switch off an AI feature?
Switch it off when the evidence shows that full cost exceeds realised value, the use case misses its safety or quality guardrails, the workflow is not adopted, or the feature cannot be measured well enough to justify continued spend. Preserve the decision record and keep the underlying data and process available if the feature is redesigned.
Sources & the researchers worth crediting
External figures and recommendations are credited here so you can check the reasoning. The practical frameworks are Peter Bamuhigire’s analysis, not statistics presented as facts.
- McKinsey, The State of AI 2025: adoption and reported enterprise-level EBIT impact
- IBM Institute for Business Value, CEO study: expected AI ROI, scaling, and value measurement
- IBM Institute for Business Value, Benchmarking the AI advantage in finance
- IBM, From hype to high-impact: how business leaders can realise ROI with AI agents
- WRITER, 2026 AI adoption survey: strategy, spend, and reported returns
Read next
Responsible AI Without Digital Dependence
Put data, contracts, capability, and accountability around AI before a promising feature becomes a dependency.
Agentic AI, Explained for the Business Owner
Test autonomy claims with a narrow workflow, clear approvals, and evidence of value before granting an agent more access.
The Real Five-Year Cost of a Business System
A wider total-cost view for software decisions, including support, change, data, and exit costs.
About the author
Peter Bamuhigire
Technology & Business Consultant
Peter Bamuhigire helps African organisations turn technology spending into controlled operating capability. His approach starts with the business owner, the baseline, the workflow, and the evidence that would justify keeping an investment when the initial excitement has passed.

