Measuring AI ROI without fooling yourself or your board

Most AI ROI claims fail scrutiny for the same reason: nobody measured the before. Without a baseline captured in advance, every number afterwards is a negotiation.

By Quality AboveAll · · 8 min read

Analyst reviewing performance figures on a screen
Key takeaways
  • Measure the baseline before you build. It is the cheapest step and the one that makes every later claim credible.
  • Count the whole cost: inference, engineering, review labour, monitoring and the maintenance that never ends.
  • Beware metrics that improve while the outcome does not, such as containment rate rising as satisfaction falls.

Baseline first, always

Before the feature exists, measure how long the task takes today, how often it is done, its current error rate, and what it costs. A week of measurement now prevents an unwinnable argument later about whether anything improved.

This is also where unrealistic expectations get corrected early. Teams frequently discover the task they were about to automate happens less often than assumed, or already takes less time than the story suggested, which is a cheap discovery at this stage.

Direct benefits are the easy half

Time saved per task multiplied by frequency and loaded cost is the standard calculation, and it holds up when the baseline is real. Deflected support contacts, reduced manual review, and faster processing all fit here.

Be conservative about what "saved" means. Time freed only becomes value if it is redeployed to something useful, and five minutes saved across a fragmented day is often not recoverable. Boards discount inflated productivity claims heavily, and rightly.

Indirect benefits, counted honestly

Faster response times, better consistency, improved employee experience and capacity to handle volume without hiring are real, and harder to quantify. State them separately rather than converting them to money with assumptions nobody believes.

Revenue-side effects like better conversion from improved search or recommendations can be measured properly through controlled tests, which makes them the strongest indirect claims available. Do that where you can rather than asserting it.

A business case with one honest number and three clearly labelled qualitative benefits survives scrutiny. One with four converted-to-currency estimates does not.

The full cost picture

Inference cost is the visible part and often not the largest. Add engineering to build, ongoing maintenance, evaluation upkeep, monitoring, the human review the workflow now requires, and the cost of periodically re-validating quality when models change underneath you.

Model deprecation is the cost teams forget entirely. Providers retire models, and re-validating a feature against a successor is real work on someone's roadmap. Assume it happens roughly annually and it will not surprise you.

Metrics that mislead

Usage is not value. A feature people click and then redo manually is negative ROI with excellent engagement numbers. Containment without satisfaction is the support-bot version of the same trap.

Acceptance rate needs care too: a suggestion accepted and then heavily edited is not the same as one used as-is. Where you can, measure the outcome rather than the interaction, and sample real cases rather than trusting aggregates. The evaluation habits in LLM evaluation apply here as much as to quality.

Frequently asked questions

How soon should we expect measurable ROI?

For internal efficiency features, within a quarter of a real rollout. Revenue-side effects take longer and need controlled tests to attribute credibly.

What if we did not measure a baseline?

Reconstruct what you can from historical data and be explicit about the uncertainty. Then measure properly from now, so the next claim is stronger than this one.

Should AI cost sit in engineering or the business unit?

Wherever the benefit is claimed. Cost and benefit landing in different budgets is how features get defunded despite being valuable.

Need to justify AI spend to a sceptical board? A free 30-minute consultation will help you build a case that holds up to questions.

Value you canactually evidence.

Baselines measured before build, full-cost accounting, and outcome metrics rather than engagement vanity numbers.