Agentic AI
Token Costs
Global Business Services

Token costs are the new budget risk in agentic AI. Here's how to plan for them

Sally Fletcher
July 21, 2026
4
min. read

Token pricing can vary 20x between models. Hypatos CEO Uli Erxleben explains why outcome-based pricing and not token tracking is the safer bet for GBS leaders.

Shift your operation teams to high-value tasks
By enabling Autonomous Finance
Free test demo

Token costs are the new budget risk in agentic AI. Here's how to plan for them.

Every CFO conversation about agentic AI eventually comes down to the same question: What is this actually going to cost us? Not the license fee. Not the implementation project. The ongoing, invisible line item running underneath every agent's every action: tokens.

We put the question to Hypatos CEO and Co-Founder Uli Erxleben, because it's the one coming up in every webinar, every business meeting, every conference hallway conversation right now.

What a token actually is

Large language models don't read words the way we do. They convert language into numerical representations (i.e., tokens), process those tokens, and hand back an answer in tokens. Every token, in and out, has a price attached. That price is the currency of working with an LLM.

For casual use, for example chatting with an assistant, asking a quick question, the cost is negligible. Nobody needs to think twice about it.

Agentic workflows are a different story. As Erxleben explains, once you're running multiple agents across a multi-step, end-to-end process, which could include parsing long contracts, pulling from master data, chart of accounts, and procedure manuals, the volume of words being processed, and therefore tokens being consumed, climbs fast.

Why the token bill gets away from people

Token pricing isn't uniform. Different models carry different costs, and the gap isn't small — Erxleben puts it at up to a 20x difference between models for the same underlying task. The highest-performing, most capable reasoning models from the frontier labs also tend to be the most expensive.

That creates a moving target. New model versions ship constantly. Each one needs to be evaluated for performance and re-evaluated for token cost. Retiring an old model isn't optional either; providers eventually deprecate it, forcing a migration regardless of whether you were ready.

And it's not just about cost. A new model version can interpret the same prompts and the same underlying data differently, which means every migration also demands regression testing across your workflows.

Now scale that across a real deployment — Erxleben's example is a workflow running 50 different agents, each potentially on a different model. Tracking performance, cost, and behavior across that surface area isn't a side task. It's a standing engineering function. Companies that have tried to run this in-house are the source of what Erxleben calls the "horror stories" — organizations that went live on agentic workflows and woke up to token bills in the millions, with no real visibility into how they got there.

The case for paying by outcome, not by token

This is where the build-versus-buy conversation gets concrete. Hypatos prices by outcome rather than by token consumption. The logic: different tasks carry different complexity, and the price should reflect that, not the underlying token mechanics.

A low-complexity task, such as extracting data from an invoice, requires little reasoning and few tokens, so it costs very little. A high-complexity task like tax coding or tax compliance demands more capability and more tokens, so it's priced higher. Either way, the price per outcome holds for the life of the contract, whether that's three years or five, with agreed limits on any cost movement at renewal.

The upside isn't just predictability. It's that the burden of tracking model versions, optimizing token usage, and re-testing after upgrades shifts to the vendor. As Erxleben puts it, that work needs to be handled by a specialist, not by an internal team already stretched across other priorities.

What happens if token prices actually go up?

It's a fair objection: if the underlying token cost rises, won't that get passed straight through? Erxleben's answer starts with the trend line — token prices have been falling, driven by competition among the frontier labs. But he doesn't dismiss the risk entirely. Data center capacity and energy are finite, and demand for tokens keeps climbing. A supply-side shock is plausible.

There's a second risk that gets less airtime: access. Erxleben points to a real recent case — a government restricting access to certain large language models outside its own borders. That's not a pricing problem. That's a system going dark overnight.

Hypatos's answer to both scenarios is as follows. The default is running efficient, competitively priced commercial models from the frontier labs. But the company also maintains self-hosted, open-source alternatives — distilled from those same frontier models, running a few months behind on raw capability but more than adequate for high-volume, document-based transactional work. If commercial access is disrupted or pricing spikes, there's a fallback.

Most companies can't replicate this on their own. Running open-source models at scale means owning GPU infrastructure that most organizations have no interest in building. But for the transactional processes that can't afford downtime, vendor payments, for instance, that contingency isn't optional. If the system is down, the business is down.

The takeaway

Token costs aren't a footnote in the agentic AI business case. They're one of the biggest variables in it and one of the least visible until it's too late. The organizations getting this right aren't the ones trying to master token optimization internally. They're the ones asking vendors to own that risk, price by outcome, and build in contingency before a shock forces the issue.

Unleash the potential of your people and business

Dial up results for any team with agentic transaction processing

Further stories from our blog