Sat. Aug 1st, 2026

FOREMY INSIDER BRIEFING · ECONOMICS OF AI

The Invoice That Stopped Making Sense

For two years, buying AI meant buying tokens. A finance team would open a vendor bill, see a line for “input tokens” and “output tokens,” multiply by a rate card, and call it a budget. It was tidy, it was auditable, and it was almost completely disconnected from the question every executive actually cared about: did the thing we bought make the business better?

That disconnect has now become impossible to ignore. As frontier labs have pushed inference costs down at a pace that would have sounded like marketing fiction three years ago, the per-token invoice has started to look like the wrong unit entirely — a bit like billing a law firm by the word instead of by the outcome of the case. Inside boardrooms and finance functions this summer, a different phrase has started to circulate: useful intelligence per dollar.

Why the Old Metric Broke

Per-token pricing worked reasonably well when models were expensive, scarce, and roughly interchangeable in capability. Under those conditions, the token was a decent proxy for cost, because cost was the main variable that moved. That world no longer exists. Inference has gotten dramatically cheaper across nearly every serious provider this year, driven by a combination of more efficient model architectures, smarter routing between large and small models, and a wave of new accelerator hardware coming online. When the underlying cost of a token falls by an order of magnitude but the value delivered by a well-orchestrated agent rises by more than that, billing by the token stops telling anyone anything useful.

Worse, per-token pricing actively rewards the wrong behavior. A verbose, meandering model that takes twenty thousand tokens to solve a problem a sharper system solves in two thousand looks, on a token invoice, like it did ten times more “work.” In practice it did less. Finance teams who have spent a year reconciling AI spend against actual productivity gains have started to notice that their token counts and their business outcomes tell almost unrelated stories.

What “Useful Intelligence Per Dollar” Actually Tries to Measure

The emerging framework asks a simpler and harder question: for every dollar spent, how much verified, usable output did the organization get back? That reframing sounds obvious once stated, but operationalizing it is where the real work begins. A serious scorecard along these lines typically has to account for several layers at once:

  • Task completion, not token count. Did the agent finish the ticket, close the loop on the customer request, or produce a document that a human accepted without heavy rework?
  • Error and rework cost. A cheap answer that a human has to fix afterward is not actually cheap; the correction time has to be priced back into the total.
  • Latency and throughput. Intelligence that arrives too slowly to be used inside a workflow has a lower effective value than the same intelligence delivered in real time.
  • Model routing efficiency. Increasingly, production systems don’t call one model for everything; they route simple queries to small, cheap models and reserve frontier-class reasoning for the subset of tasks that actually need it. A good cost metric has to capture the blended economics of that routing, not just the sticker price of the flagship model.

None of this is trivial to standardize, and that is precisely the point of contention right now. Every vendor has an incentive to define “useful” in the way that flatters its own product, which means the metric risks becoming another marketing battleground before it becomes an actual industry standard.

The Provider Response: Racing Down the Cost Curve

Whatever the eventual accounting standard looks like, the underlying trend it is trying to measure is unambiguous: the cost of running capable models has been falling fast, and providers are competing openly on that axis for the first time. Where model announcements used to lead with benchmark scores, more recent releases have leaned just as heavily on cost-per-task comparisons, cheaper context windows, and pricing tiers aimed squarely at high-volume enterprise workloads. Open-weight models trained by well-funded challengers have added additional downward pressure, since any enterprise unhappy with a frontier lab’s pricing now has a credible, high-quality alternative to threaten to switch to — and increasingly, to actually switch to.

This has created a genuinely unusual dynamic: capability and price are both moving in the buyer’s favor at the same time, something that rarely happens together in mature technology markets. Normally a market matures by capability plateauing while price competition intensifies. Here, capability is still climbing quickly even as price falls, which is part of why the old billing models are struggling to keep up.

What This Means If You’re the One Signing the Check

For teams evaluating AI spend right now, the practical implication is to stop treating the token rate card as the primary decision input. A model that costs twice as much per token but resolves a task correctly on the first attempt, with less orchestration overhead and less human review, will usually beat a cheaper model that requires three retries and a human editor. The organizations getting real value out of AI budgets this year are the ones that have built their own lightweight version of a useful-intelligence-per-dollar scorecard internally, even an imperfect one, rather than waiting for an industry-wide standard to arrive.

A few practical questions worth asking before the next renewal cycle:

  1. What percentage of AI-generated output currently ships without human correction, and how has that trended over the last two quarters?
  2. Is spend concentrated on a single flagship model, or is there a routing layer sending simple tasks to cheaper models?
  3. How does the fully loaded cost per completed task compare across vendors, not just the headline per-token rate?
  4. Are open-weight alternatives being benchmarked against the incumbent contract on a recurring basis, or only at renewal time?

A Short History of Metrics That Outlived Their Usefulness

This is not the first time an industry has had to abandon a billing convention that made sense at one stage of maturity and stopped making sense at the next. Cloud computing spent years billing by raw server hours before shifting toward consumption-based and then outcome-oriented pricing as workloads became more elastic and more automated. Telecommunications moved from per-minute billing to flat-rate data plans once voice stopped being the primary product. In each case, the shift lagged the underlying technical reality by a few years, because billing systems, procurement habits, and vendor contracts are sticky even after the economics underneath them have moved on. AI’s transition away from per-token billing looks like it is following the same lag, just compressed into a much shorter window because the underlying cost curve is moving so much faster than those earlier technology transitions ever did.

What makes this transition unusually difficult to standardize is that “usefulness” is far more context-dependent than “server hours” or “minutes of talk time” ever were. A customer support agent and a scientific research assistant both consume tokens, but what counts as a successful, valuable outcome for one looks nothing like the other. Any serious industry-wide metric will likely need to be more of a framework that companies adapt to their own workflows than a single universal number that can be quoted on an earnings call, which is part of why so many vendors are currently proposing competing versions rather than converging on one.

There is also a quieter cultural shift happening alongside the technical one. Teams that have started measuring outcomes instead of tokens report that the exercise itself changes how they build, not just how they budget. Once a team has to define what “success” looks like for a given AI-assisted workflow well enough to measure it, that definition tends to sharpen the workflow’s design, surface hidden failure points, and expose tasks that were never well suited to automation in the first place. In that sense, the shift away from per-token billing is doing something more valuable than fixing an invoice; it is quietly forcing a level of rigor about what AI is actually being asked to do that many organizations had skipped during the initial rush to adopt it.

The Foremy Take

Metrics shape behavior, and the industry is mid-transition between two very different ones. Token-based billing optimized for volume; outcome-based measurement optimizes for judgment. The providers who win the next eighteen months will not necessarily be the ones with the lowest sticker price, but the ones whose systems make the fewest mistakes per dollar spent chasing an answer. Enterprises that keep score by tokens will keep negotiating the wrong number.

What to Watch Next

  • Whether a genuinely independent, cross-vendor benchmark for “useful intelligence per dollar” emerges, or whether every lab publishes its own favorable version.
  • Continued downward pressure on flagship model pricing as open-weight competitors mature.
  • Growth of routing infrastructure that blends cheap and expensive models inside a single workflow, which will make simple token accounting even less meaningful over time.
  • Procurement teams beginning to write outcome-based clauses into AI vendor contracts rather than pure consumption-based ones.
  • Whether analyst firms and industry consortiums step in to certify a shared standard before the term gets diluted by competing vendor definitions.

This report is part of Foremy's ongoing AI Insider Report series, tracking the economics, infrastructure, and policy decisions shaping the AI industry. Foremy Team, foremy.com/.

By Foremy

Foremy

Leave a Reply

Your email address will not be published. Required fields are marked *