What I love about AI transformation right now is that everyone is sharing lessons learned - not just the wins, and lately it’s all about the eye popping bills!
As every company across nearly every industry seeks to empower their teams, guardrails are hard to set on tokens as every user aims to push the edges of what’s possible, often with the latest, most expensive models.
In parallel, the sovereignty conversation running through enterprise tech right now is bringing up many important AI architecture stack decisions, but isout of reach for most of the operators I talk to. They will never own a GPU cluster. Control has to come from somewhere else.
In this post, I’ll cover:
Why the recent AI failures are failures of economics, not capability
Why the standard answer, own your own infrastructure, doesn’t work for most organizations
Where control actually sits and the ownership question that separates real control from a change of landlord
I want to start with the failures, because they are the least ideological evidence available.
In August, Canva told investors it expected roughly 20% revenue growth for 2026 rather than the 30% it had planned on, according to Fortune, after its AI features turned out to cost far more to run than the company had modeled. Users loved the features. That was the problem.
Uber ran through its entire 2026 AI coding budget in about four months after rolling Claude Code out to thousands of engineers, and CTO Praveen Neppalli Naga has since said the company is "coming to the end of the so-called tokenmaxxing era" (Fortune, August 2026). Adoption was the success metric. Adoption broke the budget.
Neither of those is a failed AI project. The technology worked as designed in both cases. What nobody owned was the economics underneath it, and that turns out to be a separate thing you have to own on purpose. The FinOps Foundation's State of FinOps 2026 puts scale on how new this is: 98% of practitioners now manage AI spend, up from 31% two years ago. A discipline built to govern reserved cloud instances is now forecasting token consumption.
Software companies hit this first because they shipped AI to millions of users first. Physical operations in industries like energy, construction, and infrastructure are not exempt, they are just behind but not by far
David Vellante and Amit Eyal Govrin laid out the sharpest version of the argument I have read in SiliconANGLE's Breaking Analysis this month. Vendors want progress measured in tokens and model calls, which are their revenue metrics rather than anyone's business outcome. The alternative they describe is control: over routing, over policy, over the ability to change providers without rebuilding the application. Their line for it is "Cheap is a price. Sovereign is a position." I think that is exactly right.
Now look at who the examples are: Uber, Canva, Microsoft, and the Swiss federal research system, which trained its own open model on a national supercomputer. Every one of them could write a check for GPUs and build out their own infrastructure stack.
The reality for most operators is that they cannot afford to do that, and they shouldn’t try. For them, the practical version of the SiliconANGLE argument is not "own your infrastructure." It is a harder question: If the floor of the stack is going to be rented no matter what, which layer is worth owning outright?
Vellante and Govrin get close to the answer and then move past it. In one paragraph on context management, they sketch an agent that queries a knowledge graph for just-in-time state instead of re-deriving the situation from scratch, and note that it would produce more deterministic results on a fraction of the tokens. Then the piece moves on to gateways and routing policy.
That paragraph is the whole thing, for an operator.
Everyone is converging on routing as the control point, and Stripe's $7.5 billion for OpenRouter this month says where the industry thinks the leverage is. But a router only decides which model gets the work. It has no opinion about what the work is. Ask it whether a question about a delayed switchgear delivery is routine or load-bearing and it cannot know, so the safe default is to follow set algorithms.
That is not a pricing failure. It is an ignorance tax, paid per request, forever.
The reason I’m not neutral here is because I have built one of these before. My platform teams at Atlassian began our work on Teamwork Graph, the context layer that powers Atlassian's AI more than 5 years before GenAI became hot. Teamwork Graph was meant to build a complete picture of the SDLC for an organization from across many tools and systems of record so analytics on where the bottlenecks were could be clearer to diagnose and address. When GenAI came along, the Teamwork Graph not only made every query response get better, but with fewer tokens, because an agent that knows how a piece of work connects to everything around it does not have to re-derive the company on every prompt. If you’re curious, I wrote about what that taught me when we started Ontollo in my last blog.
So here is the claim I would defend. An ontology, a plain definition of what the things in a business are and how they relate, is usually described as setting the ceiling on what AI can eventually do for a company. But it also sets the floor on what that AI costs to run.
Nobody has pushed the own-your-alpha argument harder than Alex Karp, CEO of Palantir. As Vellante and Govrin summarize his position, a company should control its compute, its models and the value created from its own proprietary knowledge, and consumption pricing on frontier models works as a tax on the enterprise. Palantir has built a lot of what that requires: orchestration, governance, deployment, and an ontology. That last item is not a simple technology feature - it requires understanding of the business, workflows and some amount of FDE integration and data hygiene that cannot be avoided. Investment in this, though, creates value for clients exactly where the leverage sits.
Vellante and Govrin's objection is not to the advice. Rather, it’s to the route. Karp's solution runs open-weight models on GPUs in a data center the customer controls, underneath an operating layer that is Palantir's and is not open. Their read is that this swaps one dependency for a harder one: the meter goes away, and what replaces it is a layer that is considerably more expensive to leave than a contract is. They argue the sovereign version of the same advice would run on something inspectable and forkable, so that the crown jewels are actually held rather than deposited.
That objection is worth taking seriously by anyone selling into this, and I include Ontollo in that. A model of how a business's work connects is only an asset if the business can take it somewhere else. In my last post, I included a question that every organization should ask their provider – and this article really stresses the importance of it: If you decide to change platforms in three years, does the model of how your business runs and own work connects come with it?
Underneath our customers' operations, Ontollo maintains a deterministic layer that tracks what is true in the business rather than a model guessing at it. AI reasons on top of it, inside boundaries the customer sets, and every output carries a trail showing how it got there. I used to call the trail a compliance feature, but it really isn’t. It is what lets a company argue about its own AI spend from a position of fact.
None of this argues against renting frontier capability. For most operators, that is the only reasonable option.
Instead, I’d argue that there are three things that should not be rented alongside it.
The firms that come out of the next two years ahead will not be the ones who negotiated the best rate per million tokens. They will be the ones who could tell, in advance, which questions were worth the expensive answer.
Noah and I are building that layer for operators who are never going to own a data center, and we think it belongs to the operator. If you are in the middle of one of these vendor conversations right now and interested to see how Ontollo could make an impact, we’d be happy to give you a completely customized demo tailored to your exact business case.