Skip to content

Why Most Organizations Cannot Prove AI Is Working

McKinsey’s 2025 research found that 88 percent of organizations use AI in at least one function. Yet only a small fraction reported a meaningful share of profit attributable to it.

That gap, between deploying AI and proving it works, is the Impact dimension of the 6 I’s framework. And it is the most neglected dimension across organizations.

Impact is the lowest-scoring dimension more often than any other in readiness assessments. Not because organizations do not care about measurement. Because measurement is the last thing they get around to; and by then, the baseline data they needed was never collected.

The measurement gap

Think about what happens in a typical AI deployment. A team identifies a use case. They build or configure the AI system. They pilot it with a subset of users. The pilot works. They scale it. Six months later, someone asks; “Is this actually saving us money? Time? Errors?”

And nobody can answer with evidence.

The problem is not that the AI is not working. It might be working beautifully. The problem is that nobody measured the process before AI touched it, so there is no baseline to compare against. Nobody defined what success looks like before deployment, so there is no criteria to evaluate against. And nobody built a feedback loop, so the data that could answer the question is not being collected.

This is the Impact gap. And it is expensive in a way that does not show up on any dashboard; because the dashboard does not exist.

What Impact actually measures

In the 6 I’s framework, Impact covers five areas.

Success definition. Has the organization defined what success looks like before deploying AI? Not “improve efficiency”; specific, measurable outcomes tied to business objectives.

Baseline measurement. Do baselines exist for the processes AI is augmenting? Can the organization quantify performance before AI so it can measure the delta after?

Operational KPIs. Are KPIs connected to business outcomes (cycle time, error rate, customer satisfaction, revenue impact) rather than just technical performance (model accuracy, uptime)?

Feedback mechanisms. Does AI performance data flow back to teams that can act on it? Is there a defined cadence for reviewing AI system performance?

Evidence-based decisions. Can the organization answer “is this AI initiative worth continuing?” with evidence? Are there clear criteria for scaling, iterating, or sunsetting?

Organizations that score a 1 on Impact deploy AI without defining success and have no baselines. Organizations that score a 5 have a comprehensive measurement system with evidence-based portfolio management.

Why Impact is the dimension that closes the loop

Impact is positioned last in the 6 I’s sequence for a reason. It is the dimension that validates everything else.

Intent defines why. Intelligence identifies where. Infrastructure provides the platform. Integrity provides the guardrails. Ingenuity prepares the people. Impact answers the question; was all of that worth it?

Without Impact, organizations operate on faith. They believe AI is working because the system is running, because the vendor said it would, because the pilot looked promising. But belief is not evidence. And when the next budget cycle arrives; or the next leadership change, or the next competing priority; belief is not enough to sustain investment.

The organizations that sustain AI investment over multiple years are the ones that can walk into a budget meeting and say; “This AI initiative reduced contract review time by 62 percent, saving 14,000 hours in Q2, with a measured ROI of 340 percent. Here is the dashboard.” That sentence requires baselines, KPIs, and feedback loops; the components of Impact.

The cost of skipping measurement

When AI initiatives lack measurement, three things happen.

First, good initiatives get killed alongside bad ones. In a budget squeeze, unproven initiatives are the first to go. An AI system that is actually delivering value but cannot prove it gets cut at the same rate as one that is not delivering value. Measurement is how good initiatives survive organizational scrutiny.

Second, bad initiatives persist. Without evidence, there is no mechanism for identifying AI deployments that are not working. They continue consuming resources because nobody has the data to justify sunsetting them. The cost is not just the resources they consume; it is the opportunity cost of resources not allocated to initiatives that would have worked.

Third, organizational learning stalls. Without feedback loops, the organization does not learn from its AI deployments. Each new initiative starts from scratch instead of building on the evidence of what worked and what did not. The learning curve that should flatten over time stays steep.

Where to start

If Impact is your critical gap, start with the initiative that matters most.

Define success before you deploy. For the highest-priority AI initiative, write down what success looks like in specific, measurable terms. Not “improve efficiency.” Something like “reduce invoice processing time from 14 minutes to 4 minutes with less than 2 percent error rate.” If the initiative is already running, define success retroactively; it is better than not defining it at all.

Measure the baseline. Before AI touches a process (or right now, if it already has), measure the current state. How long does the process take? What is the error rate? What does it cost? These numbers are the denominator against which AI’s contribution will be measured.

Build one feedback loop. Pick the single most important metric and set up a way to track it over time. A monthly review. A dashboard. A spreadsheet someone updates every Friday. The mechanism matters less than the cadence; AI performance data that flows to a team that can act on it is infinitely more valuable than data sitting in a log file.

Measurement does not have to be sophisticated to be useful. A simple before-and-after comparison is better than no comparison. A basic monthly review is better than no review. Start small. Get the loop running. Refine later.

The full Assessment toolkit uses 30 scored questions with detailed rubrics, stakeholder interviews, document review, gap profiling, and a 90-day roadmap. It is designed to work whether you are assessing your own organization or helping someone else assess theirs. Get it in link below

The AI Readiness Diagnostic book covers the complete assessment methodology. The AI Readiness Assessment Toolkit includes the book plus 8 professional files; the Scoring Workbook, Interview Guide, Document Review Checklist, Findings Report Template, Executive Summary, 90-Day Roadmap Template, Findings Presentation, and Re-Assessment Tracker. Explore both at builttooperate.com.