A featured contribution from Leadership Perspectives: a curated forum reserved for leaders nominated by our subscribers and vetted by the CIOReview Advisory Board.

Merit Data Tech
The Intelligence Gap: Why More Data Is Not the Same as Better Intelligence


Daniel Dennis
There is an assumption sitting underneath most enterprise data strategies that rarely gets challenged directly. The assumption is that more data produces better AI outcomes. More sources, more volume, more coverage. If the model is not performing as expected, the answer must be to feed it more.
In my experience working with organisations across automotive, construction, energy, healthcare, maritime, and retail, this belief causes more commercial damage than almost any other in enterprise technology. More data does not produce better intelligence. Better data does. And the two are very different things.
The research reflects what I see on the ground. Informatica's 2025 CDO Insights survey found that 43% of organisations cite data quality and readiness as their number one obstacle to AI success. Not model capability. Not infrastructure. Not talent. Data quality. The same survey found that only 12% of organisations report data of sufficient quality and accessibility to actually run AI applications effectively. That means the overwhelming majority of enterprise AI programmes are being built on a foundation that their own data leaders know is inadequate.
The IBM Institute for Business Value puts a financial figure on the same problem. Their 2025 research found that more than a quarter of organisations lose over $5 million annually due to poor data quality alone, before any AI investment is factored in. When you layer AI spend on top of a poor-quality data foundation, those losses compound rather than diminish.
What I see consistently when tracing the root cause is that the data feeding these AI programmes was collected for a different purpose. Automotive pricing records structured for dealer management workflows. Construction planning documents assembled for project tracking. Healthcare formulary data gathered for clinical reference. Each was fit for its original use. None was designed with the accuracy, freshness, and domain specificity that a production AI system demands. The model works. The outputs quietly mislead. Nobody immediately knows why.
What makes this particularly difficult to diagnose is that volume disguises the problem. An organisation with ten million records feels like it has a data advantage. Often it does not. What it has is ten million records of varying accuracy, varying freshness, and varying relevance, processed by a model that has no mechanism to distinguish the reliable records from the ones skewing its outputs.
McKinsey's 2025 AI research makes the strategic implication explicit. Organisations reporting significant financial returns from AI are twice as likely to have redesigned their end-to-end data workflows before selecting modelling techniques. The winning pattern is not better models on top of existing data. It is better data, approached as a first-principles decision before the modelling conversation begins.
The organisations I work with that have genuinely closed this gap share a common characteristic. They stopped measuring their data strategy by volume and started measuring it by quality. Specifically, they asked a harder question: is the data we hold accurate enough, current enough, and specific enough to our sector that we would trust a commercial decision made on the back of it?
That question has a concrete answer. AI-ready data has verifiable properties that most generic datasets do not meet. It is validated at point of collection rather than corrected downstream, which according to Gartner costs organisations an average of $12.9 million annually when that correction happens after the fact. It is enriched by people who understand the sector it describes, not just the technology that harvested it. It reflects the actual terminology, classification conventions, and data structures that matter in a specific industry. And it is maintained on refresh cycles that match how quickly the underlying market moves.
The commercial consequence of getting this right is not incremental. Organisations operating on high-quality, domain-specific data make faster decisions because they trust the intelligence without needing a manual validation layer between the data and the action. Their AI systems perform against their original design intent because the foundation was right from the start. And the data asset compounds in value over time in a way that raw volume never does.
I have explored this dynamic at length in a whitepaper on AI-ready data harvesting. But the observation worth making plainly here is this: the intelligence gap in most enterprises is not between organisations with more data and those with less. It is between organisations that built their data to be useful and those that accumulated data and hoped it would be.
For any CIO or CDO reviewing their data strategy right now, the most useful question is not whether you have enough data. It is whether the data you have was built to the standard your AI actually needs.