Is your architecture preventing you from calculating AI value? – CIO

I argued in my last column, The case and model for real-time AI cost visibility at the infrastructure layer, that the AI measurement problem looks to be finally solved as a technical matter, and that companies can now manage AI as a strategic investment rather than a pay-and-pray experiment. That’s the defining shift for the next era of AI, but with one clarification: it’s only true for organizations whose architecture permits it.
My work puts me inside enterprise engineering teams, helping practitioners measure and optimize their cloud and AI costs. The biggest roadblock is rarely budget or expertise. It’s that AI spend, unlike cloud spend, doesn’t attach to anything you can tag. A token call has no owner, one API key can carry a dozen workflows across several teams, and the thing spending the money is usually an agent rather than a person.
The answers these teams need are only available by looking underneath the application, where you can watch what’s running. Getting there isn’t a complex engineering project, but it does require access to the machine the workload runs on, and that access is structurally unavailable through managed services.
So, the first thing I ask now is: “What does your inference actually run on?” The response determines the accuracy, completeness and usability of the AI cost insights we can get.
The implication of the platforming decision is easy to state and expensive to reverse. You run AI either on services your cloud provider operates for you, or on compute you operate yourself, meaning VMs or container nodes you control. Almost every team picked one many months ago on platform engineering merits, for reasons that had nothing to do with AI measurement. It just wasn’t obvious then that the choice would also determine how much they could understand about their own business later.
But here we are. Gartner expects the average Fortune 500 enterprise to run more than 150,000 agents by 2028, up from fewer than 15 in 2025. A blind spot you could live with across a handful of workloads becomes a serious headache across six figures of them.
To get ahead of claims of bias, I recommend managed services regularly. They get you to a working agent much faster, and they take scaling, patching, availability and capacity off your plate, none of which differentiates anybody. If you come to me with three engineers and an AI roadmap, my recommendation is to use serverless and managed services as much as you can. Running your own infrastructure requires a platform team and the maturity to go with it, which makes this a question of size and stage.
Those advantages are settled science. The other trade-offs materialize later, and while CIOs have weighed most of them before, they’re worth a brief mention.
Model selection is where this shows up in dollars. Open-weight models often run on infrastructure you control, while managed catalogs can lag availability or limit model choice. That matters when a newer model materially changes inference economics or performance. If you’re using a managed service that doesn’t offer the model you want, you can’t route to it through that service.
Stripe agreed in August to buy OpenRouter for more than $7 billion, a strong indicator of what the market thinks routing is worth. Of course, you can still route between models on a managed service, but you have to build the router into your own application code and run it yourself, outside the service you bought so you wouldn’t have to do that stuff.
The provider’s schedule also determines when you get access to new capabilities. For example, plenty of teams on their own infrastructure started building against MCP within days of its release, while stateful MCP server support didn’t reach AWS’s agent runtime until this past March, more than a year later.
It’s also worth noting that agents handing off to each other can introduce another execution or initialization cost, whereas on your own nodes you can keep resources warm. Enterprise authentication is not always one of the methods on offer, so you may have to build that path yourself anyway. And in regulated industries, proving where data can be a dealbreaker.
None of this is news to anyone who has run a platform, and none of it is disqualifying. You can put a number on each and decide it’s worth paying. The next one doesn’t work that way.
Let’s start with exploring what exactly your managed service provider tells you about your AI spend. The bill arrives on the provider’s schedule and can give you a detailed view of what you spent on a service, often by account or API key. That tells you where spend went up or down. It doesn’t necessarily tell you which workflow did it, which customer or which employee triggered it (not just created it). It also doesn’t enable you to map that spend to outcomes, productivity or revenue. A bill is not a measurement of value.
Going further, an AI agent isn’t a single service either. It’s the model call plus MCP servers, vector databases, APIs, prompt management and whatever else the workflow leans on, and that supporting cast can represent a substantial share of the cost. Without the ability to review those costs in isolation, you’re evaluating on estimates. You can’t say what a feature costs to serve, which accounts are profitable at the terms you signed, or whether the expensive part is the model or everything around it. So, you ballpark, and everything downstream inherits the inevitable errors.
The standard workaround has been to instrument the application, tracking each model call and carrying cost data through the workflow. I’ve built that many times and, on infrastructure you control, it works. But inside a managed runtime, you’re instrumenting someone else’s execution model. You can track the calls you make, but you may not be able to see the initialization you’re paying for, the orchestration between steps or the retries the platform runs on your behalf. You get detailed numbers for your own code, but less visibility into everything happening around it, and that’s often where the surprises are.
The expensive AI failures are also episodic rather than steady or predictable. An agent might loop on a retrieval it can’t satisfy, a workflow could revert to an expensive frontier model or a prompt change adds context and cost to every downstream call. If you’re reviewing the bill on a monthly cadence, you see that the AI number on the bill is bigger than last month, but don’t know why.
With real-time, per-request granularity you can instantly identify the workflow that’s misbehaving and it’s usually a quick fix. It also sets how soon you can intervene.
The way to get that granular, real-time view is to watch the kernel, using the same eBPF technology that security and observability tools use to see system calls without touching the applications above them. In an implementation like this, it can map compute and network activity back to the relevant process and request, then join outbound model calls to provider cost data. The advantage is that the view doesn’t depend on developers remembering to instrument every call.
The catch is that it only works where you control the host machine, a narrower set of places than most people assume. On AWS it means EC2, and container workloads on EC2-backed ECS or EKS nodes. It doesn’t mean Fargate, where you don’t control the host kernel. The same principle holds across clouds: customer-controlled VMs and Kubernetes nodes can give you access to the kernel, while serverless and fully managed runtimes generally don’t.
It’s a firm line in the sand. You either have the option, or you don’t. There are no workarounds here. That alone doesn’t make managed services a mistake, and provider telemetry and application instrumentation still matter, they just can’t give you the same infrastructure-level view. If your inference runs where you don’t control the machine, you get as much infrastructure-level visibility as your provider chooses to share, and that ceiling remains capped even as the number of agents and the value of visibility grows.
Today, price per token is the standard unit for AI spend, but it’s the wrong one for agents. Agents burn far more tokens than chat; input drives most of the cost, and token use on the same task swings widely between runs, so cheaper tokens can still produce more expensive tasks. Cost per completed task is what tracks value.
But pricing a task means matching spend to its specific workflow and a measurable outcome. On infrastructure you control, you can get that visibility. On a managed runtime, you’re mostly inferring from an aggregated bill.
The last two years rewarded shipping. The next phase rewards knowing what each task costs and what it’s worth, and companies that can’t see the work will be competing against ones that can.
Eduardo Mota is principal cloud architect and AI/ML specialist at DoiT, where he works embedded with enterprise engineering teams on cloud and AI architecture, cost and optimization, with particular depth in Google Cloud and large-scale AI systems. His experience spans DoiT and AWS, focused on architecture and cost optimization for large organizations. He writes and speaks about the practical side of running AI in production: How teams attribute and control spend, where FinOps practices built for the cloud era break down with token-based workloads, and what it takes to turn AI from an unpredictable line item into a measurable, strategic investment.
Sponsored Links

source
This is a newsfeed from leading technology publications. No additional editorial review has been performed before posting.

Leave a Reply