60% of companies that invested in AI report no material value from that investment. In the same body of research, only 5% have taken AI all the way to value creation at scale. This piece works through what those numbers are actually measuring, and — separately — what LLL sees as the structural difference between landing in the 60% and landing in the 5%.
It’s written for a decision-maker who has already run a PoC, confirmed that something works, and is now stuck on whether — or how — to take it to production in a way that shows up in the numbers.
”60%,” “88%,” “7%” — Different Questions, Not the Same Study
Two large 2025 surveys produced numbers close enough in timing that they get cited interchangeably. They aren’t measuring the same thing.
| Source | Sample | Numbers | What it measures |
|---|---|---|---|
| BCG, “Are You Generating Value from AI? The Widening Gap” (2025) | 1,250 companies | 60% report zero material value / 35% scaling partially / 5% creating value at scale | Value creation from AI investment |
| McKinsey, “The State of AI” (2025) | ~2,000 organizations | 88% use AI in at least one function / 32% in a scaling stage / 7% fully scaled / 39% report EBIT impact (most under 5%) | Breadth of adoption and economic impact |
McKinsey’s 88% answers “are you using AI somewhere.” Within the same survey, “fully scaled” drops to 7%, and only 39% of organizations can point to a measurable EBIT impact at all — most of that group under 5%. There’s a real step between “using it” and “it’s showing up in the P&L.”
BCG’s question is more direct: it asks about value creation from the investment itself. 60% report no material value, 35% are scaling partially, and 5% are creating value at scale. Even among companies that have adopted AI, converting that into money is the minority case.
Put the two studies side by side and the same shape appears: a large majority using AI, a much smaller group getting a priced result, and that group narrowing further the higher the bar gets set.
The “6%” Figure Circulating Online Is Answering a Different Question
A statistic that shows up often — “only 6% of companies are getting value from AI” — does not hold up against its primary source.
The 6% comes from HBR Analytic Services, sponsored by Workato and AWS, surveying 603 business and technology leaders in July 2025. What it measures is the share of companies that fully trust AI agents to autonomously run core business processes — not whether AI investment produced value. 43% trust agents only with limited, routine tasks; 39% restrict them to supervised or non-core use. That’s a statement about trust in autonomous execution, not a statement about ROI.
Two adjacent survey results get mixed together often enough that the citation is wrong even when the underlying number is real. Checking that before using a figure is part of the point of a piece like this one — which is also why the rest of this article sticks to the BCG and McKinsey numbers above for anything about value creation.
What Actually Separates the 5% From the 60%
Neither BCG nor McKinsey drills into the mechanics of what separates the two groups in much operational detail (BCG points to broader patterns among “AI Future-Built” companies — workflow redesign, heavier investment in upskilling). What follows is not from either study. It’s LLL’s own structural read, built from operating private LLM infrastructure and running PoCs — an operating hypothesis, not a finding from the cited research.
Split 1: Is where the data goes decided before the build starts?
“Using AI” isn’t one technical choice — it’s a decision between three configurations with genuinely different data-flow characteristics: a private LLM (inference completes inside infrastructure you control), a frontier API (prompts go to a provider’s servers — OpenAI, Anthropic, Google), or a hybrid setup (sensitive material routed to the private path, everything else to the frontier model).
Leave that decision vague through the PoC stage, and it tends to surface right before production — a legal or security question about whether this specific data is actually allowed to leave the building through a third-party API. That’s where a number of projects stall. Our read is that the 60% group skews toward building something that works technically first, and treating data governance as a later step rather than a starting condition.
Split 2: Is the investment made after proving it holds on your own data — not after a demo?
Generative AI demos tend to perform well on generic benchmarks or vendor-supplied sample data. That doesn’t guarantee the same accuracy on your actual documents and your actual question patterns. Skip that check and move straight to hardware or a long-term contract, and the failure mode shows up later — accuracy falls short, or the workload turns out to need a bigger model than planned — after the money is already committed.
The order matters: validating against your own workload before the large financial commitment, not after it.
Split 3: Is model size chosen against the workload, or by default?
“Bigger model, better outcome” and “start small to stay safe” are both defaults, not answers. A 32B-class dense model is a practical floor for business use; a 70B-class dense model is a reasonable default sized to approximate a company-wide ChatGPT-equivalent; 122B-A10B fits heavier concurrent usage; and a 35B-A3B tier is sized for a triage and guardrail role specifically — not as an answering model. Quantization follows the same logic: going below FP8 (to 4-bit, for instance) produces a measurable accuracy drop on Japanese-language business tasks in testing, which is why it isn’t shipped as a default recommendation regardless of the hardware savings it offers on paper.
Pick a size without a workload-based reason behind it, and the result tends to land on one of two failure modes: paying for capacity that isn’t needed, or falling short on accuracy that was needed.
What the Numbers Are Actually Saying
The gap between 60% and 5% is probably not a gap in what AI itself can do. It looks more like a gap in whether the decisions that need to happen before the investment actually happened before the investment. Both studies agree on the same underlying shape: there’s an unfilled step between “we started using this” and “it’s showing up as money.”
If a PoC already worked and the project has stalled since, the next move is probably not “try a bigger model.” It’s more likely to be settling the three questions above — data flow, validation order, model sizing — against the actual workload, specifically.
Taking an AI Investment to the Next Step
Where this shows up in what LLL builds: private LLM infrastructure, AI development work more broadly, and a 30-day PoC built specifically to test whether an investment holds up on your own data before it gets bigger.
Private LLM Infrastructure AI Development Services 30-Day Private LLM PoC
Sources: