A practical framework for enterprise leaders
Where your AI investment is actually going - and how to prove what you're getting back.
Every department just bought its own AI tools. Almost none can show what they returned.
That sentence would have sounded aggressive eighteen months ago. Today it's just an accurate description of the average enterprise: Salesforce Einstein in Sales, Copilot seats across Engineering, Gong and Clari layered onto the revenue org, a scattering of point solutions in Legal, Finance, and Ops - each purchased on its own business case, each measured (if measured at all) by its own vendor's dashboard.
Boards have caught up to the gap. "What's our AI ROI?" has moved from a curiosity to a standing agenda item. The problem is that most organizations have no reliable way to answer it - not because the AI isn't working, but because nobody built the measurement layer that would tell them either way.
This guide is about that measurement layer: how to find where AI and software value is actually leaking out of your operations, how to put a real dollar figure on it, and how to close it - without a six-month instrumentation project standing between you and the first answer.
Don't start from a use-case matrix or a feasibility score. Generic prioritization frameworks - value-vs-feasibility matrices, use-case scoring models, process-mining tools - rank hypothetical initiatives. They tell you where to look. They don't tell you which dollar figure is real today.
Start instead by reading what your own systems already show:
Part 3 below walks through exactly this, on a real fintech's Salesforce + Gong + Clari stack - a $77,235/yr leak found across four findings, none of which would have shown up on a feasibility matrix.
Spending on AI is easy. Turning it into measurable value is hard. In practice, the gap between the two almost always comes down to one of four barriers.
Manual, high-volume work remains embedded in operations despite being squarely suited to automation or AI augmentation. This is the barrier everyone assumes they've already solved and almost nobody has. It survives because it's invisible in the tools that would normally surface it - a rep who bypasses an AI coaching summary to write notes from memory doesn't show up as a support ticket or an error log. It shows up nowhere, which is exactly the problem.
If you pulled the raw usage logs for your three most expensive AI tools right now, would you know what percentage of the workflow they're actually touching versus how much still happens by hand around them?
Enterprise systems accumulate unused AI capabilities, under-adopted features, and misconfigured tools at a rate that outpaces anyone's ability to track by hand. A Copilot license sitting at 40% activation isn't a rounding error - on a 500-seat deployment at $30/seat/month, that's over $100K a year in fully paid, unused capacity, and it's rarely the only tool in that state.
Can you name, right now, the activation rate of every AI-tagged line item on your software budget? If the honest answer is "no," this barrier is active in your organization today.
Isolated AI initiatives launched without a shared value roadmap, business case, or operating model. This is the barrier that makes the first two worse over time: when Sales, CS, and Finance each run their own AI pilots with no shared measurement standard, the organization ends up with five different definitions of "working" and no way to compare them, prioritize between them, or learn from one to accelerate the next.
If two different departments each claimed their AI initiative delivered value last quarter, could you compare those two claims on the same basis?
No reliable way to connect AI and operational initiatives to measurable improvements in revenue, productivity, or performance. This is the barrier that shows up in the boardroom, but it's a symptom of the first three, not a separate root cause - you cannot prove ROI on value you can't see, on waste you can't quantify, across initiatives you can't compare.
The next time someone asks what your AI investment returned last year, will the answer be a number tied to a source system, or a sentiment ("teams seem to like it")?
Wherever your organization sits today, value is leaking through the gaps between these stages - and most organizations significantly overestimate which stage they're actually in.
| Stage | What It Looks Like | What's Missing |
|---|---|---|
| 1. Ad Hoc | Isolated experiments. Individual employees use AI on their own initiative. | No shared strategy, no governance, no policy or budget attached to any of it. |
| 2. Emerging | Pockets of adoption. A few teams run trial projects, usually driven by one or two internal champions. | Still uncoordinated - what works in one team doesn't transfer to the next. |
| 3. Scaling | Coordinated rollout. Shared standards, training, and policy start appearing across departments. | ROI is just starting to be measured, usually inconsistently. |
| 4. Embedded | AI is core to how the organization runs. It's a default step in workflows, not an add-on. | Continuous optimization - spend tied directly and immediately to outcomes. |
Two things are consistently true across every organization we've measured against this model:
First, maturity is not a single number. An enterprise doesn't have one AI maturity stage - Engineering might be operating at Stage 3 while Finance is still at Stage 1. Any maturity assessment that produces one company-wide score is averaging away the information you actually need, which is where to focus next.
Second, the stage-to-stage transitions are where the money is. The leap from Emerging to Scaling is where fragmented pilots either get a shared measurement standard or calcify into permanent silos. The leap from Scaling to Embedded is where "ROI is being measured" turns into "spend is tied directly to outcomes" - and that's usually a bigger jump than it sounds, because it requires the underlying data to actually support it.
Frameworks are only useful if they produce a number. Here's what one real scan found - a pre-IPO fintech, $1.5B/yr in originations, running Salesforce, Gong, and Clari.
The approach was straightforward: pull 14 business days of historical logs across the three systems, model the behavioral patterns in those logs, and quantify exactly where capital was leaking. Pro-forma annual margin leak uncovered: $77,235/yr, across four findings:
None of these four findings would show up in a standard usage dashboard - a dashboard would have shown Gong as "adopted" (it was installed and generating briefs) and Clari as "active" (forecasts were being generated on schedule). The leak was only visible one layer down, in the gap between what the tool produced and what the humans around it actually did with it.
One of these four findings, investigated one level deeper, turned into something larger. The 4.2-day Sales-to-CS gap was worth asking a single, specific question about: what actually happens between a signed contract and Customer Success? The answer - a person manually emailing a broker by hand - led to reconstructing 14 related workflows in 10 days, from just 2 stakeholder conversations, surfacing three additional AI opportunities worth $420K in combined annual value (the largest single opportunity worth $180K on its own, with an estimated 4-month payback). The lesson generalizes: the size of a leak you can see from the outside is rarely the size of the opportunity underneath it.
Score each item honestly, 0-2 (0 = not at all, 1 = partially, 2 = fully in place):
Scoring:
The instinct, once a leak is suspected, is to launch a full instrumentation project. That's almost always the wrong first move - it's slow, it requires broad system access before anyone has proven the hypothesis is worth the access, and it puts the security review in the critical path before there's anything concrete to review.
A tighter sequence, staged so that each step only asks for as much access as the previous step's finding justifies:
Test the framework above against your own numbers, using synthetic or estimated data. Zero system access required. The goal isn't a real answer yet - it's identifying which department and which metric is worth testing against real numbers.
Export an aggregated dataset yourself - ticket volumes, call brief open rates, override counts, whatever's relevant to the hypothesis from Sandbox - and run it through the same framework. No live connection, no IT ticket. This step alone usually produces a real, defensible dollar figure.
If the manual sample justifies it, set up a narrow, read-only, time-boxed connector to one system in one department. Never broader than the hypothesis requires.
With a validated, dollar-quantified finding and a named budget owner, decide whether to widen scope. If yes, the same connector pattern extends to more departments - a scope change on an already-reviewed grant, not a new security review from scratch.
The organizations that get through this sequence fastest aren't the ones with the most sophisticated data infrastructure. They're the ones that resist the urge to ask for full access on day one.
A usage dashboard tells you a tool was opened. It doesn't tell you that a rep opened it, decided it wasn't worth trusting, and quietly did the work by hand anyway - which is the actual leak. Usage and value are correlated, but the gap between them is exactly where the money is.
A consulting engagement is accurate on the day it's delivered and stale the day after - the findings don't update as the organization changes, and drift creeps back in silently. A measurement layer that stays live catches drift as it re-forms, not just the snapshot from six months ago.
No. The 90-day path above is specifically sequenced so that the first real, dollar-quantified finding comes from a manual data export - no live connection, no IT ticket. Deeper access is only requested once a finding justifies it.
Most existing maturity assessments (including well-known analyst models) score inputs - governance, strategy documents, training completion - rather than outcomes. This framework is deliberately outcome-first: the question isn't "do we have an AI governance policy," it's "can we show, in dollars, that AI changed how a specific workflow runs." The two are complementary, but they answer different questions.
Using the self-assessment above to identify a target, followed by a manual data sample (Part 5, Days 7-21), most organizations get a real, dollar-quantified finding inside three weeks - without any live system access.
Both. The same measurement layer that surfaces a $77K/yr leak is what makes the case for the next AI investment defensible - because it's the only way to show, after the fact, whether that investment actually worked.
Everything above is a framework anyone can apply by hand, with a spreadsheet and a few hours. The constraint most organizations hit isn't the framework - it's the labor of pulling logs from five different systems, reconciling them, and re-running the analysis every time something changes.
That's the specific gap Melt is built to close: reading the real system logs (not self-reported usage), running the Four Barriers and Four Stages framework above continuously rather than as a one-time exercise, and surfacing findings - like the $77,235/yr example above - tied to a real log, a real dollar figure, and a named owner. Every organization we've measured against this framework had a leak nobody had found yet. The self-assessment above is the fastest way to find out where yours is.
Start with the self-assessment above. Once you identify your constraint, the 90-day path is designed so you'll have a real number - not a sentiment - within three weeks.
Request a Melt Score