MIT's research, still the reference point a year on, found that 95 percent of generative AI pilots produce no measurable P&L impact, and that the gap is not about model quality. It is about approach: nobody agreed what the return was, in numbers, before the money moved. Arivue builds that case before the pilot starts, priced the way a CFO already prices everything else.
The arithmetic is the easy part. It is the three things underneath it that tend to fall apart in the room.
Twenty minutes back for four hundred people sounds enormous, and by 2026 it no longer clears the bar on its own: finance wants the P&L line, not the productivity estimate. That means saying what happens to the time: fewer contractors, more output, or honestly, nothing yet.
Pilots are cheap because they are small. Usage-based pricing, per-seat licenses, and the people who keep it running all grow with rollout. A single round number for cost hides the one that decides the answer.
Three initiatives counting the same drop in handling time is how a portfolio promises more than the business ever banks. If a case does not say how much of the benefit it claims, the number is a hope.
NPV and IRR are mechanical once the inputs are honest. The real discipline is earlier: knowing what kind of value you are claiming, and what it takes to defend that specific kind. This is the “approach” MIT's research points to: most cases fail here, not at the model.
Cost avoidance, revenue growth, productivity capacity, risk reduction, and experience or retention benefits are not the same kind of number, and pricing them with one blanket multiplier is where most AI cases quietly stop being true. A headcount-hours-saved calculation resolves through a completely different mechanism than a churn-reduction calculation: one only becomes a real cost line if the freed capacity is redeployed or headcount actually drops, the other only becomes a real revenue line if the base rate and the causal link both hold up.
Teams reliably find the obvious benefit and stop there. The second-order value, what the freed capacity, faster cycle time, or better decision actually enables downstream, is usually larger and almost always uncounted, because nobody went looking for it. A rigorous case surfaces this on its own instead of waiting for a reviewer to ask "is that all?"
Adoption and attribution are not one haircut applied at the end. Adoption moves over time, a ramp, not a step function, and attribution has to be argued benefit by benefit: a productivity gain the case owns outright looks nothing like a revenue lift shared with three other initiatives touching the same customer.
A cost that grows with usage, priced like a one-time cost, is how a pilot's economics quietly invert at scale: the number that clears the hurdle rate at 50 users can fail it at 5,000. Build-versus-buy is really a cost-structure decision wearing a technology decision's clothes.
A hard requirement, a latency floor, a data-residency rule, a security posture, found during architecture review instead of during scoping is not a technical delay. It is a valuation error: the case was priced against an approach that was never viable.
The approved case, not the pitch, becomes the reference point. Without it, "did this work" has no answer, and the next AI request restarts the credibility argument from zero instead of building on a track record.
What it costs once everyone is on it, and what it is worth once adoption and attribution are applied, not the number on the slide before either is.
Illustrative structure.
Arivue is not an AI ROI calculator. It is a capital decision platform, and AI initiatives are one of the things it makes defensible. Same financial engine, same checks, whether the ask is a model, a migration, or a warehouse. AI budgets are now facing the same board-level scrutiny every other line item gets, and an AI initiative has to clear the same bar as everything else asking for the money, not a lighter one because it is new.
The same way as any investment, with two extra steps. Work out the benefit for each value driver using the method that fits it, then reduce it for how many people will really adopt the thing and how much of the result this initiative can honestly claim. Set that against the full cost, including what usage will cost at full rollout rather than at pilot size. Net present value, internal rate of return, and payback follow from there.
The problem and what it costs you today. Ranked options with a clear build, buy, or hybrid position. A solution outline. Cost built from its parts rather than one round number. Value drivers with adoption and attribution stated. Hard requirements like response time and data residency checked against the approach. The assumptions and risks the numbers rest on. And a baseline to measure against after approval.
Usually because nobody agreed what the return was before the pilot started. MIT's research on enterprise AI found this is the actual dividing line between the pilots that prove a return and the 95 percent that don't: not model quality, but whether the return was defined and priced before the money moved. Without a baseline for what the work cost beforehand and a measure everyone accepts, there is nothing to compare the result against. You end up with a demo that impressed people and no evidence anyone can fund.
As separate lines with their own rate, volume, and period, so the running cost grows with rollout instead of staying stuck at pilot size. One-off build cost is kept apart from ongoing run cost, and a stated total that does not match its own arithmetic gets corrected and flagged rather than quietly accepted.
They should be judged on the same standard, which is rather the point. AI cases tend to carry softer benefits, faster-moving costs, and more contested credit, so they need more rigor, not a scorecard of their own. Holding them to the same structure is what lets a committee compare an AI request with everything else asking for the same money.
Yes, and it is a common place to start. Bring in what you already have. The platform structures it, shows you what is missing, and produces the case for the scale-up decision, using the pilot's real results as evidence rather than as a projection.
Pick the one you cannot get funded, or the one you cannot justify continuing. Thirty minutes, your real numbers.