The short answer. Most AI cost reviews establish whether a customer is profitable today. Gartner's August forecast adds a harder question: whether that contract stays profitable until the company is allowed to change the economics.
A software contract fixes the price for a term. The company's ability to respond is not continuous, and arrives only on renewal dates, uplift clauses, and entitlement changes.
The Repricing Gap measures the interval between today and the next date the company can change what a customer pays, against the cost it will be carrying by then.
Table of Contents
Research Grounding
Gartner published the forecast on August 17. Inference costs per agentic workflow will increase more than fivefold through 2028. Gartner calls the mechanism the Inference Paradox and defines it as better unit economics escalating overall AI costs without a clear path to matching value.
Three trends produce it, in Gartner's account:
Foundation model cost economics are improving rapidly. Price per token continues to fall.
That efficiency makes more capable and more expensive models affordable to deploy.
Sophisticated workflows consume far more tokens than simple chatbot interactions. An agent reasons, replans, and questions itself across many calls where a chatbot answers once.
Gartner states the consequence directly: "Product leaders cannot rely on more efficient token economics to rationalize AI costs. Each successive generation of AI capability will necessitate more, and often more expensive, tokens." Gartner adds that routing a task to an agentic reasoning model increases inference costs at least fivefold, and by more as the task becomes more complex.
A separate Gartner forecast published August 10 shows where the money is going. Worldwide spending on AI-optimized infrastructure as a service reaches $42.3 billion in 2026, growth of 96%. Inside that figure, inference passes training for the first time, at $23.3 billion against $19 billion.
The PE Translation
The July 14 issue covered the measurement problem, and BCG's estimate that gross margins for AI-enabled software are resetting into a 65% to 80% band. On its own, that reads as a single step down to a lower level, and a company can plan against a level. Gartner's forecast says the level keeps moving through 2028, making it a question of timing rather than size.
A software contract fixes what the customer pays for a term. The cost of delivering the AI inside it isn’t fixed, and the company's ability to respond isn’t continuous either. It arrives on particular dates: a renewal, an annual uplift clause, an entitlement change, the introduction of a new tier. Between those dates, the price is settled, and the cost is not.
A company can change the model a workflow routes to within a sprint. The date it can change what the customer pays was set at signature.
For a portfolio company, the exposure is therefore the length of that interval and where the cost curve reaches by the end of it. A three-year enterprise agreement signed in 2026 at a flat fee with unlimited AI usage holds its price across the whole of the period Gartner is forecasting, and in many cases past the exit.
A buyer's diligence looks at both halves. The technical team asks what a completed workflow costs, and the commercial team reads the contracts; together, they produce a schedule of when the seller can act. A seller who has not built that schedule sees it first in someone else's data room.
Where inference is consumed to deliver a customer-facing feature, management should evaluate it as part of the cost of delivering that revenue rather than leaving it in a general AI or R&D line. The accounting treatment depends on the nature of the spending and the company's policies, so the controller should confirm it.
Operator Experience

Three things make the cost curve harder to hold than a plan usually assumes.
First, the company's own product decisions push it up. Many of the quality improvements a team reaches for first also raise the cost of a completed task, because accuracy often comes from retrieving more documents, using longer context, adding a verification pass, or using a more capable model at the step that was getting it wrong. Teams ship those changes in response to accuracy complaints, which are the loudest signal a customer sends, so cost rises fastest on the accounts that were already expensive to serve.
The second is what optimization can and cannot do. Routing simple steps to cheaper models or even SLMs (Small Language Models), caching repeated retrieval, batching, and distilling a smaller model for a narrow task all work, and a team that has not done them has real headroom. A capable CTO will object that the company can optimize its way down, and that is a fair point. Optimization changes where the cost curve starts, not the date the company is allowed to alter the commercial terms. The plan should therefore model the optimized cost against the contract calendar rather than treat optimization as the answer to both.
The third is enforcement. Features sold as unlimited were generally built without a counter, a quota, or a throttle, because nothing in the original design needed one. Adding metering later means changing behavior customers already have and have not been told is finite, so that work belongs on the roadmap well before the renewal it supports.
A usage limit in the contract does nothing until the product can count what a customer uses.
Board Question
For our AI-bearing contracts, what happens to gross margin if cost per completed workflow rises fivefold before we have the right to reprice?
The Repricing Gap

The METER Test in the August 11 issue asked whether a company can change what it charges for at all. This asks a narrower question: when the paper allows it, and whether the cost arrives first. The gap is the interval between today and the next date the company can change what a customer pays. Four entries measure it.
Entry | What it records |
|---|---|
Cost today | Cost per completed workflow for that customer, produced by the One-Bill Test in the July 14 issue |
Sensitivity | Contract gross margin at the optimized inference cost, then at two, three and five times that cost |
Next date | The next date price, entitlement or usage limit can be changed, and by what mechanism |
Floor | The cost per workflow at which the contract falls below the company's gross margin threshold |
A contract belongs on the action list when the floor is reached before the next date arrives. The floor should be the company's board-approved gross margin target. BCG's 65% to 80% range for AI-enabled software is an external benchmark for that conversation and not the target itself.
The exercise produces a list rather than a score. Where the next commercial action date arrives first, the contract can be handled in normal account planning.
A contract with a three-year term, a flat fee, and no usage language is the one to look at first.
Three Decisions
1. Sort the AI-bearing contracts by their next repricing date
This is a single pass through the contract set by commercial and finance together, and it can be finished before the FY27 plan is approved. The output is a schedule of when each contract can be changed and by what mechanism.
2. Build the enforcement ahead of the renewal that needs it
Metering, quotas, and throttles take engineering work to add to a shipped feature, and longer to introduce commercially to customers who have used it without limits. Starting that work when the contract comes up for renewal leaves very little room to change the product, test the controls, and prepare the customer.
3. Change what a new contract assumes about cost
For agreements signed from here, the terms should assume the cost of delivery moves during the term. An annual repricing right, a usage cap with overage, or an AI component on a shorter term than the base contract each achieve it. Choosing none of them is an acceptance of the margin risk and should be recorded as one.
One Number for the Next Operating Review

55%.
Inference is 55% of Gartner's 2026 forecast for AI-optimized infrastructure as a service, $23.3 billion of $42.3 billion, and it passes training spending this year for the first time.
Training is comparatively episodic. Inference is consumed every time the product performs the work, for as long as the customer holds the contract. The crossover marks the point at which AI spending starts behaving like a cost of delivering the service, scaling with customer consumption and running for the length of the contract. That is the form in which it reaches gross margin and the exit.
Board Takeaway
Gartner expects the cost of running an AI workflow to rise more than fivefold through 2028, while the price in a multi-year contract signed in 2026 does not automatically move with it.
Before the FY27 plan is approved, management should be able to produce a schedule of when each AI-bearing contract can next be repriced, and say what gross margin it would be on the contracts where that date arrives after the cost does.
Portco Brief is a weekly briefing for PE operating partners and portfolio company executives focused on technology, AI, and value creation. If this was forwarded to you, subscribe at portcobrief.com.
Sources
Gartner, "Gartner Predicts AI Inference Costs Per Agentic Workflow Will Increase More Than Fivefold Through 2028", August 17, 2026.
Gartner, "Gartner Forecasts Worldwide Artificial Intelligence-Optimized IaaS Spending to Grow 96% in 2026", August 10, 2026.
BCG, "Return on AI: What CEOs Need to Know About the True Cost of Intelligence", July 1, 2026.

