Our last issue examined cybersecurity, where the proliferation of AI agents is pushing enterprises to redesign security models suited to a former era. This edition closes out our three-part look at the changing cloud software sector with data infrastructure, the foundational layer supporting everything else.
While cybersecurity defends against an adversary, data infrastructure contends with a structural problem: data is accumulating faster and across more systems and formats than many organizations can organize, govern, or trust. AI is raising the stakes – models and agents are only as good as the data that feeds them. IDC projects the global datasphere will surpass 700 zettabytes by 2030, an increasing share of it created not by people, but by machines and AI systems generating data as a byproduct of their own operation.
The industry is converging on the open lakehouse: a shared, open data foundation that combines the management and performance of a warehouse with the flexibility of a data lake. Popularized by Databricks, its core principles have now been adopted across the major cloud and data platforms. Feeding, moving, and governing that data at the speed its newest consumers – machines – now require is the current and more challenging task. Three forces are reshaping the category:
- The explosion of ‘unstructured’ data: Gartner estimates 70-90% of enterprise data is now unstructured (documents, images, audio, and code); these are formats that legacy infrastructure was never built to store or search.
- The shift to streaming architectures: Enterprises are abandoning scheduled batch systems for pipelines that ingest and process data continuously, as it is created.
- The complexity of hybrid and multi-cloud governance: Per Flexera’s 2026 State of the Cloud Report, 73% of organizations now run hybrid cloud environments, with a further 14% operating multi-cloud without a private cloud component – obscuring where data lives, whether it can be trusted, and who is permitted to access it.
A related shift is underway alongside these operational pressures: the business model underneath the category is being rewritten. Software has historically monetized the seat; an agent doing the work occupies none, and that is the risk markets priced into SaaS multiples this year. However, agents are heavy, unpredictable consumers of data – every action pulls records, checks permissions, and writes results back. Each of those actions is billable under a usage-based model, so the same dynamic undermining per-seat software works to the advantage of consumption-based pricing instead.
The data infrastructure businesses we find most compelling sit at that intersection. They charge for usage rather than seats and turn data governance from an operational burden into a product capability, making data as legible to machines as it has long been to people. Across frontier AI, cybersecurity, and data infrastructure, we’ve seen the same pattern hold: the companies that make enterprise systems more secure, data more trustworthy, and outcomes more predictable are the ones that become harder to displace.