Every enterprise AI review eventually reaches a slide with a number that goes up every quarter (messages sent, users onboarded, queries answered per week) and a room that cannot say what the number actually bought the company.
That is not a reporting problem. It is a measurement design problem, decided before the pilot ever launched: someone chose to instrument what was easy to count instead of what was hard to isolate. Activity inside a tool and value the tool created look identical right up until someone asks what the activity was worth.
The honest fix is not a better dashboard. It is choosing a different unit before the use case ships, capturing what things looked like beforehand, and being willing to say "we don't actually know" when the two do not line up.
Vanity metrics vs. value metrics
Most AI ROI decks measure activity because activity is free to log and always trends up. None of the numbers below are dishonest on their own. Stacked together, they get dressed up to resemble a return.
| Common metric | What it actually tells you |
|---|---|
| Messages sent to the AI | People opened the tool. Says nothing about whether the answer changed anything. |
| Users onboarded | Accounts exist. Says nothing about who is still using it in week four. |
| Queries answered | The system responded. Says nothing about whether the response was acted on. |
| Response time / uptime | The system worked. Says nothing about whether anyone needed what it produced. |
None of these are worth abandoning as monitoring. A tool nobody opens has failed regardless of anything else. But they measure the tool, not the business. Three movements measure the business.
The three movements worth measuring
Hours returned. The clearest, most defensible win in enterprise AI is recap and compile work that used to eat a person's morning and now does not. A weekly regional sales recap, a monthly reconciliation across three spreadsheets, a status report assembled from five systems before a Monday meeting — these are the tasks worth timing, because they recur on a fixed schedule and someone already owns them. Instrument this by measuring how long the specific artifact took to produce, start to finish, before the AI touched it, and again after.
Decision latency. Question-to-answer time compresses when someone can ask a department-scoped agent instead of calling three people and waiting for a reply. This is not the same as hours returned (nobody was doing "waiting" as their job), but a decision that used to sit for a day waiting on a stock number now clears in minutes, and that has value even when no headcount moves. Instrument it by timestamping when the question was asked and when the action was taken once the answer arrived, not just when the answer appeared.
Exceptions caught earlier. An overdue receivable, a stock anomaly, a margin slip — every one of these has a lag between when it happened and when a human noticed. Shortening that lag is real value even when the underlying problem still has to be fixed by a person. Instrument it by comparing when the event occurred against when someone acted on it, before and after the alert existed.
Baseline first, or the number means nothing
Every one of the three movements above is a delta, and a delta needs two points. The step teams skip, almost always for the same reason (it slows the pilot down and produces no visible progress), is capturing the "before" number with the same rigor as the "after."
Practically: pick the exact recurring artifact, the exact decision type, or the exact exception category the use case targets, and measure it for two to four weeks with the same people doing the same task, before AI touches it. This is unglamorous, and it is also the only way "we cut this from six hours to forty minutes" means something rather than sounding like a guess dressed as a fact. It overlaps with the same discipline an honest data readiness check already asks for: know where a number lives and who can vouch for it before using it to prove anything.
How much of that was actually the AI?
The uncomfortable second half of baseline-first is attribution. AI rarely deserves sole credit for a movement, because it rarely ships alone — a new export button appeared in the ERP the same quarter, a process got simplified, someone on the team got faster for reasons that have nothing to do with software. An intelligence layer that connects five systems is doing real work, but so is whoever finally documented what the fields in those systems mean.
The discipline that keeps this honest is a lightweight change log kept alongside the measurement: what else changed in the same window, who else touched the process, what else shipped. When multiple things moved at once, report a range rather than a point estimate, and say plainly which other factors were in play. A number nobody can dispute is usually a number nobody checked.
What did it cost, and what is the net?
Every movement above is a benefit, and ROI is a ratio. The denominator has been missing so far: what the tool cost to run. A number that measures value returned but never subtracts the spend behind it is half a calculation, and it is the half a CFO will notice is missing.
The cost side has three parts, and only the first is obvious. Licensing or subscription is the visible line. Implementation is the larger and quieter one: the systems integration, the data-cleanup work, the time internal people spent connecting sources and agreeing definitions, all of it real cost even when no invoice names it. Ongoing ownership is the third: whoever maintains the connections, reviews the answers, and runs the enablement as the tool spreads. Summed across a year, this is the total cost of ownership, and it is what the returned value has to clear before the initiative is actually in the black.
There is also a method question hiding inside "hours returned": how do you turn saved hours into a currency figure a CFO will accept? The defensible conversion is the loaded cost of the specific person whose hours were freed, meaning their fully burdened cost per hour, which finance can supply, rather than a headline salary or an invented benchmark rate. Multiply the hours the baseline measured against the after by that loaded rate, and be explicit about one thing: freed hours only convert to money if they were redeployed to something that mattered or if headcount actually changed. An hour saved that turned into an hour of slack is a real quality-of-life gain and a currency figure of zero, and saying so out loud is what keeps the whole calculation credible.
When do you kill a use case?
A use case with a real baseline and a full measurement cycle behind it that still has not moved any of the three numbers does not automatically earn "give it more time." It earns a hard look, using the same discipline that decided which use case to try first when building the roadmap.
Two patterns look similar and are not. Adoption climbing while none of the three movements shift usually means a workflow problem (the tool is being used, but not for anything that changes an hour, a decision, or an exception), and the fix is redesigning where it sits in the workflow, not giving it another quarter. Usage flat from the start usually means the entry point was wrong: the question the use case answers was not one anyone actually had. Killing a use case in that state is not a failure of the AI; it is the measurement doing its job.
Where Nalar fits
Nalar is built to make these three movements measurable rather than assumed: department-scoped agents that carry out the compile and recap work directly, live queries that shorten decision latency because the answer is drawn from current systems rather than yesterday's export, and automations that catch exceptions on a schedule instead of at the next quarterly review.
If you want to see how the numbers behind a use case actually move before committing to one, the interactive demo runs on a realistic mock enterprise across six industries. If you are not yet sure your data can support a credible baseline in the first place, BARI will tell you honestly where that stands — including when the answer is "not yet."
Frequently asked questions
- What is the difference between a vanity metric and a value metric in AI ROI?
- A vanity metric counts activity inside the tool (messages sent, users onboarded, queries answered) and always trends upward regardless of impact. A value metric measures a change in the business itself: hours returned, faster decisions, or exceptions caught earlier.
- How do you measure hours returned from an AI use case?
- Pick a specific recurring artifact the use case targets (a weekly recap, a monthly reconciliation) and time how long it took to produce, start to finish, before AI touched it. Measure the same artifact the same way afterward and compare.
- Why does AI ROI measurement need a baseline?
- Every real ROI claim is a delta between a before and an after. Without capturing the before-state with the same rigor as the after, there is nothing to compare the improvement to, and any percentage claimed is effectively a guess.
- When should a company kill an AI use case instead of extending it?
- When a use case has a genuine baseline and a full measurement cycle behind it and still has not moved hours returned, decision latency, or exceptions caught, it needs a redesign or an end — not an automatic extension on the assumption it will improve with time.