The article is clearly relevant to enterprise AI – it’s about managing AI costs at enterprise scale, tokenomics, FinOps for AI, etc.
Let me score newsworthiness:
– Timeliness: 3 (fresh McKinsey survey from May 2026, highly current)
– Significance: 3 (93% of organizations exceeding AI budgets – massive finding)
– Novelty: 2 (FinOps for AI is somewhat known but the specific data is new)
– Actionability: 3 (very actionable with specific recommendations)
– Audience fit: 3 (directly relevant to CIOs, CTOs, CFOs, etc.)
Total: 14 – well above threshold.
Now let me write the article following the persona instructions:
– HTML output mandatory
– Senior analyst at Enterprise AI publication
– Voice of Ben Thompson at Stratechery: opinionated, analytically sharp, takes positions
– Write for CIOs, CTOs, CISOs, CMOs, CROs, CFOs, CHROs, CDOs, COOs, CEOs
– Every paragraph in
tags
– Every heading in
The AI Budget Crisis Is Already Here — And Most CIOs Are Flying Blind
McKinsey’s analysts note that CIOs instinctively reach for the cloud analogy when trying to contextualize the AI cost problem. That instinct is understandable and partially useful: governance gaps, shadow spending, fragmented procurement, and the eventual emergence of FinOps disciplines all rhyme with what happened when AWS and Azure started eating enterprise budgets in the 2010s. But the analogy breaks down in ways that matter enormously.
Cloud costs, at their core, scaled with infrastructure utilization — servers, storage, data transfer. Predictable units with predictable price-performance curves. AI costs scale with behavior. The same task, executed by the same agent, against the same model, can generate token volumes that vary by a factor of thirty depending on how a prompt is written, how a context window is populated, or whether an agent loop hits an unexpected retry condition. No cloud CFO ever had to explain to a board that a single autonomous workflow consumed the equivalent of an annual software license in a single afternoon because a developer forgot to cap output length.
This non-determinism is the core of the problem — and it is why the traditional FinOps playbook, while necessary, is not sufficient. The McKinsey report introduces the term “enterprise AI tokenomics” to describe the emerging discipline of actively and continuously predicting and managing AI model usage to maximize ROI. The framing is right. What enterprises need is not just cost tracking — it is a demand-shaping capability that operates at the architectural level.
Shadow AI Is the Iceberg Below the Waterline
The budget overruns visible to the CFO are only part of the story. McKinsey estimates that 20 to 30 percent of AI spend in most enterprises is entirely unaccounted for — fragmented across cloud providers, foundation model vendors, software platforms with embedded AI features, business-unit-sponsored experiments, and what the report aptly calls “citizen developers” and “vibe coders.” That last category deserves attention from every CISO and COO in the room.
The same democratization of development that has made AI exciting — the ability for a non-engineer to build a functional AI-powered workflow in an afternoon — has also created a new class of unmanaged infrastructure risk. An employee building an autonomous agent with a poorly structured loop does not intend to consume millions of tokens per day. They are almost certainly unaware that they have. The organization will find out when the invoice arrives. By then, the agent has been running for three weeks, the employee has moved on to the next project, and no one in IT has visibility into what the workflow is actually doing.
This is not a hypothetical. It is the operational reality McKinsey’s survey describes when it notes that AI spending remains fragmented across “enterprise copilots, foundation-model contracts, AI-enabled software features, API-based services, experimentation environments, and business-unit purchases.” The CIO who believes they have a handle on AI spend because they control the central cloud contracts is likely missing a significant fraction of actual consumption.
The Four Levers — And the One That Most Organizations Skip
McKinsey’s framework for managing AI consumption organizes around four actions: gaining visibility, optimizing spend, modernizing sourcing, and embedding governance into architecture. Each deserves scrutiny, because the sequencing and emphasis matter as much as the actions themselves.
Visibility is table stakes and the correct starting point. You cannot optimize what you cannot see, and an organization that does not have a consolidated view of AI spend — a single control plane spanning cloud providers, model vendors, embedded software features, and business-unit purchases — is not managing AI costs. It is reacting to them. McKinsey’s finding that only 20 to 25 percent of companies have mature AI FinOps practices is striking. It means roughly three-quarters of enterprises are in the reactive camp right now, even as their spend is about to accelerate.
Spend optimization is where the tactical opportunity is most immediate. McKinsey identifies approximately 40 levers, but the highest-impact ones are more accessible than many CIOs realize. Prompt caching alone — reusing static context rather than re-sending it with every inference call — can reduce input-token costs by up to 90 percent for retrieval-augmented generation workloads and agentic applications with large, stable system prompts. Model routing — directing tasks to the lowest-cost model capable of delivering acceptable quality rather than defaulting to frontier models — represents another significant lever that most organizations are not systematically pulling. The survey finding that roughly a third of organizations have already achieved 20 to 30 percent savings through active optimization should be provocative for the two-thirds that have not started.
Sourcing modernization is the strategic dimension that will separate AI-mature organizations from the rest over the next 18 months. McKinsey frames this as the death of the buy-versus-build question. The correct framing is now buy, build, host, route, and switch — a continuous portfolio management discipline that evaluates model performance, cost, risk, and business value against a rapidly evolving landscape. The organizations that lock themselves into rigid vendor relationships today, attracted by the apparent safety of committed spend discounts, may find themselves unable to shift workloads to better-performing or lower-cost models that will emerge in the next technology cycle. In AI infrastructure, optionality has compounding value.
The fourth lever — embedding governance directly into architecture — is the one most organizations will be tempted to defer, and the one that McKinsey’s analysis suggests is actually the most durable solution. AI gateways, policy engines, and automated guardrails that enforce budget thresholds, route requests to appropriate models, and monitor agent behavior at runtime are not nice-to-have additions to an AI platform. As adoption scales, they become the only practical mechanism for governing consumption. Manual oversight does not scale. Architecture does.
What This Means for Every C-Suite Stakeholder — Not Just the CIO
It is telling that McKinsey addresses this report primarily to CIOs, because the budget and governance ownership naturally lands there. But the strategic implications extend well beyond IT leadership, and the organizations that treat AI cost management as a CIO problem will almost certainly manage it less effectively than those that distribute accountability appropriately.
For CFOs, the immediate priority is establishing a cost-attribution model that connects AI consumption to business outcomes rather than treating it as an undifferentiated technology line item. McKinsey’s recommended evolution — from tracking token consumption to measuring cost per claim processed or revenue generated per AI-enabled workflow — is exactly the right framing. It also happens to be the framing that enables rational capital allocation decisions about which AI investments to expand and which to rationalize.
For CMOs and CROs with significant AI-powered customer-facing workflows, the agentic complexity described in this report has direct implications for how you model the unit economics of AI-assisted selling, service, and personalization. A customer service agent that retries failed tool calls or expands context windows unpredictably is not just a cost problem — it is a latency and reliability problem that affects customer experience. Understanding the token economics of your AI-powered workflows is increasingly inseparable from understanding their operational performance.
For CHROs and COOs, the “citizen developer” dimension of this problem deserves a specific governance response. The organizations that will manage AI costs most effectively are not those that restrict access — that path leads to shadow AI and the very fragmentation that makes costs invisible. The organizations that win are those that channel the energy of distributed AI development through platforms and architectures that make responsible consumption the path of least resistance. That is a change management challenge as much as a technical one.
The Deeper Strategic Issue: AI Is Becoming an Operating Expense, Not a Capital Investment
The most important long-term implication of this report is one McKinsey does not state explicitly but that the data makes unavoidable. Enterprise AI is completing a transition from a project-based capital investment model to a continuous operating expenditure model — one where costs are variable, usage-driven, and deeply sensitive to design and behavior choices made by thousands of individual contributors across the organization.
This transition has profound implications for how AI is planned, governed, and measured. The project-based model — approve a budget, build a capability, measure ROI against the original business case — is simply incompatible with the reality that the same capability can cost ten times more to operate depending on how users interact with it and how agents are designed. The organizations that will navigate this transition successfully are those that build the forecasting, attribution, and optimization muscles now, before spend reaches a scale that makes the problem politically as well as technically difficult to manage.
McKinsey’s final framing — that managing AI economics will be “the defining challenge for CIOs in the age of AI” — is accurate, but perhaps undersells the stakes. The organizations that develop genuine competency in AI tokenomics will not merely save money. They will create the financial headroom to reinvest in the highest-value AI applications while their less disciplined competitors are managing budget emergencies. In a technology landscape where competitive differentiation increasingly depends on AI capability, the ability to sustain and scale AI investment without financial crisis is a strategic advantage, not just an operational virtue.
The room that went quiet when someone asked about cost has a choice. It can stay quiet and hope the next budget cycle is more forgiving. Or it can build the systems, disciplines, and cultural expectations that make the answer to that question something every executive in the room can give confidently. The survey data suggests most organizations are still in the first camp. The window to move to the second is open — but it will not stay open indefinitely.
Based on reporting from Enterprise AI spend and cost management, originally published 2026-07-19 20:00:00.

