The short answer. Almost any company can tell you what it spends on AI, because the invoices arrive monthly and Finance already tracks them. Far fewer can tell you what came back. The operating problem shows up in the systems: nothing reliably joins a product event to a line in the ledger, and no function owns that join.
It matters at exit. A buyer prices realized earnings from the financials, and will only underwrite the incremental AI value a seller can trace to a baseline.
The Attribution Tree runs five questions over every AI line in the FY27 plan, before the budget is approved rather than after.
Research Grounding
BCG surveyed 152 chief executives at companies with at least $500 million in revenue and published the results on July 22.
Nearly nine in ten said their companies are seeing cost or revenue benefits from AI somewhere in the business.
56% named an unclear link between AI initiatives and specific financial outcomes as what stops those benefits from reaching the P&L. It came first of the seven barriers BCG listed, ahead of workflows and incentives that nobody had redesigned, and ahead of gaps in technology and data.
14% said they have clearly defined the P&L impact for all of their AI initiatives.
So 56% of these CEOs name the measurement problem, and 14% have solved it. That is a 42-point gap.
The pattern isn't confined to software or chief executives. Gartner reported on August 5 that 55% of chief supply chain officers are unclear on the return from their AI investments, even as 67% of supply chain digital spend now goes to AI. A different function, surveyed by a different firm, produced the same answer.
The PE Translation
You could read that 42-point gap as management talking about financial rigor without practicing it; in that case, the fix is to tell them to define the P&L impact.
That doesn't work because of how the systems are built. Defining the P&L impact of an AI feature means connecting a business outcome back to that feature, and that connection exists only if somebody built the product to record it. Where nobody did, no amount of asking produces the number.
BCG's own ranking points the same way. Weak value tracking sat at 34%, below both of the organizational barriers above it. That is what you would expect if the problem starts before the tracking does.
For a sponsor, putting more money into the FY27 AI plan does not solve the measurement problem. A larger check adds initiatives to a portfolio that has no way of retiring the ones producing nothing, so pilots accumulate.
The costs are easy to see, because the invoices already show up in the monthly financials. The benefit side arrives as somebody in the room saying it worked.
During the hold period, some uncertainty is manageable. At exit, it becomes much harder. If AI has already reduced costs and those savings show up in the financials, the buyer does not need an “AI” label. The earnings are already there. The problem starts when management asks the buyer to give credit for value that is not yet visible in the historical numbers — another $2 million of run-rate savings, higher productivity, or revenue growth attributed to an AI feature. Now the buyer wants proof. What was the baseline? What would have happened without the AI? Can the operating result be tied back to the financials? Without that evidence, the $2 million is still a management claim, not value a buyer can confidently underwrite.
The AI benefit in the FY27 plan arrives as a claim, and a buyer underwrites only what the seller can trace.
By the time a sale process opens, the company holds years of documented cost against no comparable record on the other side of the ledger, and has to reconstruct the benefit from data nobody collected at the time.
Operator Experience
Underneath the governance question there is an engineering question, and it is narrow enough for a platform lead and a finance analyst to answer together.
Why the number cannot be produced today
To attribute financial value to an AI initiative, four things have to connect:
The product event. Did the AI feature actually run?
The operating outcome. What work changed because it ran?
The baseline. What would have happened without it?
The financial result. What did that change in revenue, cost, margin, or cash flow?
The product event is the easiest part. Most applications already log whether a feature ran.
The operating outcome is where the chain breaks first. Recording what the feature replaced is a separate design decision, and somebody has to make it while the feature is being built. An AI summarizer logs that it generated a summary. It does not log the eleven minutes an agent did not spend reading the thread, and that second recording is what the P&L claim rests on.
The baseline has to match the granularity the initiative operates on, and it has to exist before the work starts. Take a company that measured average handle time across a whole queue, then shipped an AI assist into one workflow inside it. The queue number moves by less than its own weekly variance, so the measurement exists and still cannot answer the question.
The last problem is connecting the data. Usage lives in the product database, revenue in the billing system, and model cost may only exist on a vendor invoice, with no shared key tying them together per feature until somebody builds one.
Why the pilot number is not the scaled number
A pilot runs on the population most likely to succeed, so it shows the mechanism works without showing the size of the effect on everyone else. Two costs also behave differently once real volume arrives:
Exception handling and manual review climb with volume. The share of cases a model handles cleanly is a rate, so the absolute number of exceptions grows with throughput.
Inference cost per transaction is easy to measure in a pilot. It is also the figure most likely to be renegotiated or re-tiered before the deployment finishes.
What the instrumentation costs
A holdout or a staged rollout is a product change, which means a flag, a stable cohort assignment, and a reporting path that survives the rollout. The join from product event to financial outcome needs a shared key and a warehouse model built for it. The work takes about a quarter, and no customer sees anything at the end of it, which is why it loses roadmap fights unless somebody funds it as its own line with an owner and a date.
Who owns it
Finance owns the P&L. Product owns the feature. The measurement sits between them and belongs to neither.
That leaves two ownership problems, and both sides have a reason to leave the work where it is:
Finance owns the number and cannot specify what should be logged. The telemetry decision belongs to somebody else.
Product can specify it and is not measured on it. It gets measured on features shipped, not on evidence that a shipped feature worked.
So both functions agree the work is necessary and neither has a reason to schedule it. I would treat that as a staffing problem. One analytics engineer, named, for two quarters, owning the baseline and the join and reporting to whoever will be asked for the number. The flag and the cohort assignment stay with the product team shipping the feature, because that is where they have to be built.
Board Question
Which AI initiatives in the FY27 plan have a named business metric, a baseline that predates the work, and a way to separate their effect from everything else changing at the same time?
An initiative that fails all three is not ready for funding. What is ready for funding is the measurement.
The Attribution Tree

If I were reviewing an FY27 plan, I would put every AI line through five questions in order, and each one comes out with one of five verdicts.
1. Is there a named business metric this initiative is supposed to move?
Gross margin, average handle time, win rate, churn, revenue per employee, cost per invoice processed. A capability does not count. "Improve the customer experience" fails here.
If nobody can name one, park it. Holding an idea costs no money this cycle.
2. Is there a baseline for that metric that predates the work?
The baseline has to match the granularity the initiative operates on, and it has to exist before the initiative starts. If the team reconstructs it after seeing the result, the evidence is much weaker.
If there is no baseline, instrument first. The measurement gets funded this cycle and the initiative waits.
3. Can the effect be separated from everything else that changed?
One of these has to exist:
a holdout group
a staged rollout
a matched cohort
a period when nothing else moved
Choosing between them is a product decision that carries an engineering cost.
If there is no way to isolate the effect, instrument first again, though this fix is the narrower one.
4. Has the effect been observed at the intended scale?
A pilot tells you the mechanism works without telling you how large the effect is. The question here is whether the number held once the population stopped being the one chosen for it.
If the effect has only been seen in a pilot, prove it in one place. Fund the smallest deployment that produces a real number, and hold the rest of the request.
If it has been seen at scale, the initiative has a number, and the fifth question decides what that number is worth.
5. Does the verified benefit clear the investment hurdle after fully loaded cost?
Set the measured benefit against everything the initiative consumes at the intended volume, and then against the return required to justify continued capital and management attention. Fully loaded means three things:
model and vendor spend
the engineering and platform time to keep it running
the human review the workflow still needs
Exception handling climbs with volume, so the pilot's cost per transaction is the wrong figure to use here.
If the benefit does not clear the hurdle, stop it. A benefit of $1.05M against $1M of fully loaded cost passes the arithmetic and fails the capital decision, and a sponsor can now retire it with a number attached.
If it clears, scale it. This is the initiative that goes into the value creation plan with a figure attached. It will hold up when a buyer asks where the figure came from.
The Attribution Tree in one screen
1. NAMED METRIC? No → PARK
2. BASELINE THAT PREDATES THE WORK? No → INSTRUMENT
3. EFFECT SEPARABLE FROM EVERYTHING ELSE? No → INSTRUMENT
4. OBSERVED AT SCALE RATHER THAN IN A PILOT? No → PROVE
5. CLEARS THE HURDLE AFTER FULLY LOADED COST? No → STOP · Yes → SCALE
How this differs from benefits tracking. Value tracking measures an initiative after it ships and reports what appears to have happened. The Attribution Tree runs before the money is committed, and its second verdict funds the measurement instead of the initiative.
Three Decisions
1. Run the Attribution Tree over every AI line in the FY27 plan before the budget is approved.
The output is a count of initiatives in each of the five verdicts. A first pass weighted toward park and instrument is not a bad result. It shows which initiatives are not yet ready for more capital.
2. Fund instrumentation as its own line, separate from the initiative it measures.
Give it an owner, a budget, a completion date, and the name of the executive who will be asked for the number later. Anything attached to an initiative gets cut the first time that initiative runs late, which tends to be exactly when the measurement matters.
3. Add one question to technical diligence: show me an AI initiative you stopped, and the number that stopped it.
Any company can produce a slide describing how carefully it governs AI. What is harder is showing an initiative the team stopped because the economics did not work, along with the number they used to make the decision. If the team cannot name one, I would want to understand whether its measurement process is influencing capital allocation at all. The Attribution Tree produces that answer for any plan it is run over, because the fifth question is where a stop verdict and its number come from.
The answer also arrives quickly. A team with the number names the initiative and the date, and a team without it describes a review process.
One Number for the Next Operating Review
Money is the least-cited barrier, so increasing the AI budget is unlikely to solve the underlying problem.
18% is the share of CEOs in BCG's survey who cited insufficient funding as a barrier to turning AI into financial impact, which put it last of the seven barriers listed.
So a sponsor who answers weak AI returns by writing a larger check is spending against the barrier management complains about least.
The figure worth asking for is the portfolio's own. Take the FY27 AI budget and work out what share of it supports initiatives with a named metric, a baseline, and a way to isolate the effect. That share is what AI plan management can evaluate with evidence. For the rest, the company may be spending money without a reliable way to decide whether to stop or scale.

Board Takeaway
Nearly nine in ten companies report AI benefits somewhere and 14% can put a P&L number on all of them. The gap sits between product telemetry and the ledger, where no function owns the join. That is a measurement architecture gap rather than a discipline gap.
Sort the FY27 AI plan into park, instrument, prove, stop, and scale before approving it. Fund the measurement as its own line, and expect a buyer to underwrite only the AI value that traces to a baseline.
Portco Brief is a weekly briefing for PE operating partners and portfolio company executives focused on technology, AI, and value creation. If this was forwarded to you, subscribe at portcobrief.com.
Sources
BCG, "Nearly Nine in Ten CEOs See Some Cost or Revenue Benefits from AI in Targeted Areas, But Most Are Struggling to Scale It", July 22, 2026, and the accompanying report, CEOs Are Starting to See Value from AI. Now Comes Execution., July 2026. Survey of 152 chief executives at companies with revenues of at least $500 million.
Gartner, "Gartner Survey Finds Majority of Chief Supply Chain Officers Unclear on AI Investment Returns", August 5, 2026. Survey of 394 supply chain professionals at organizations with at least $250 million in annual revenue, conducted November 2025 through February 2026.

