Back to Blog Listing

You cannot prove AI value without a baseline

You cannot prove AI value without a baseline
Karol Sobieraj Sep 8, 2026 5 min read

Written by: Karol Sobieraj, Founder & CEO, Digital Colliers

The MIT number everyone is quoting this quarter says 95% of enterprise GenAI pilots deliver no measurable P&L impact. Most of the pushback I read online argues with the methodology. That is the wrong fight. The interesting question is whether your own pilot could survive the same test, and the honest answer for most teams is no, because nobody wrote down what the world looked like before the model showed up.

The 95% number is really a measurement problem

Zero measurable impact is not the same as zero impact. It means the org cannot tell the difference between the pre-AI state and the post-AI state with any confidence. Once you accept that framing, the surrounding data starts to make sense. S&P Global found 42% of companies abandoned most of their AI initiatives in 2025, up from 17% the year before, and the average organisation scrapped 46% of proofs-of-concept before production. IDC's number is similar: 88% of AI POCs never reach widescale deployment. RAND puts the failure rate above 80%, roughly twice that of conventional IT.

Those numbers describe the same underlying pathology. Teams build something, cannot prove it moved a metric, and quietly shelve it. The pilots that survive are almost always the ones where somebody captured a baseline first.

What a baseline actually is

A baseline is not a slide. It is a small set of numbers, captured before the model touches the workflow, that you would be willing to defend in front of a CFO. Three properties matter:

  • It is measured, not estimated. If you cannot pull it from a system of record, it does not count.
  • It is scoped to the exact workflow the AI will touch. Not the department. The workflow.
  • It is captured over a window long enough to show variance. One week is noise. Four to eight weeks is usually the minimum.

If your "before" number is a guess from a team lead, your "after" number is also a guess, and the pilot is already in the 95%.

Five numbers to capture before you ship

The workflows people actually pilot GenAI on tend to share the same skeleton. Whether it is a support triage assistant, a coding copilot, or a contract review tool, these five numbers cover most of it:

  1. Cycle time per unit of work. How long does one ticket, one PR, one contract take today, end to end, including waits.
  2. Cost per unit. Fully loaded, including the people who review and rework it.
  3. Quality or rework rate. What percentage comes back, gets rejected, or fails QA.
  4. Volume and variance. How many units per week, and how spiky is the distribution.
  5. The human tax. Overtime hours, escalations, weekend work. This is the number that usually justifies the project internally, and the one nobody writes down.

For coding assistants in particular, be careful about the quality axis. Recent arXiv work on AI coding assistants found experienced developers reviewing 6.5% more code but showing a 19% drop in their own original code output after adoption. If your baseline is only lines shipped, you will call that a win. If it includes rework and review load, you will see the real shape.

The cost of skipping this step

The cost of not having a baseline is not that the pilot fails. It is that you cannot tell whether it failed. You spend six months, the vendor shows you a nice dashboard, adoption is fine, and you still cannot answer whether the P&L moved. So the project gets renewed on vibes, or killed on vibes, and either way the next pilot starts from the same standing.

The operators I see making it past the pilot phase in 2026 all do the same unglamorous thing. They spend the first two weeks not building anything. They instrument the current process, agree with finance on what the target metric is, and get sign-off on the measurement method before a single prompt is written.

What to do this quarter

If you have a pilot in flight and no baseline, stop the model work for two weeks and go capture one retrospectively from your logs. If you have a pilot on the roadmap, do not let it start until the five numbers above exist in a shared doc with the CFO's finance partner named on it. And if you are about to renew a vendor contract on the strength of a dashboard nobody can tie to cash, that renewal is the moment to ask the awkward question. The 95% is not a technology problem. It is a bookkeeping one.

Related Posts