Benchmarking across nearly 3,000 coding sessions, report finds ~80% of the AI bill is spent re-sending cached context, not generating answers
NEW YORK, Aug. 6, 2026 /PRNewswire/ -- New research from AI Efficiency OS company PointFive challenges a common assumption about AI cost management: that using fewer tokens automatically means spending less. An empirical study titled Token Reduction Is Not Cost Reduction charts the path toward more strategic AI budget allocation and use at scale. The study reviewed 2,908 paid coding sessions and found that reducing tool-output tokens by 38.4% actually increased billed costs by 6.8%, highlighting the disconnect between token consumption and the true cost of AI workloads.
Managing the cost of AI is the top forward-looking priority in the FinOps Foundation's 2026 State of FinOps survey. According to McKinsey's July 2026 Enterprise AI FinOps Survey, 93% of companies have exceeded their AI budgets while only 26% of companies have real-time visibility into the cost of running AI (KPMG). PointFive's report sheds light on what is happening between AI usage visibility and budget expenditure and why the most logical attempt at reducing AI costs – using fewer tokens – is failing.
"AI efficiency is a brand-new field, and most teams are making cost decisions without evidence," said Alon Arvatz, CEO of PointFive. "We did this research to understand how AI workloads actually behave, whether in the cloud or in the coding agent, and how we can bring those efficiencies to our customers. This is the first of many research reports we will publish. Research-backed evidence is how we build our product, and we look forward to sharing more discoveries."
Proper AI Measurement
PointFive's report spanned across 2,908 paid coding sessions, and incorporated 103 specific software tasks, using seven open-source repositories and tested against three models. The study produced surprising findings:
Why Cutting Prompts Raises – Not Saves – the AI Bill
The study identifies two ways that efforts to reduce token use can backfire:
The Path to Lower AI Costs: Visibility, Not Compression
The study points engineering leaders toward visibility and governance, not prompt compression, as the path to controlling AI costs. McKinsey research found that:
Costs are driven by how sessions unfold, the retries, re-reading and extra loops, and those stay invisible without per-team reporting, budgets and policy.
The benchmark is open source
The full paper, Token Reduction Is Not Cost Reduction, is a free download on arXiv (arxiv.org/abs/2607.12161), and the AI Efficiency Benchmark behind it is open source at github.com/PointFiveLabs/ai-efficiency-benchmark, so any savings claim, including PointFive's own, can be run through it. For more information, read the paper, download the field guide, or read the launch blog at pointfive.co. Further research in the series will follow, published the same way.
Methodology
The study was authored at PointFive and evaluates the company's own experimental build, an unmodified open-source compressor, alongside Headroom, a third-party tool. It is not independent research. Method, per-session data and limitations are published in full. Every cost figure was read from the actual provider bill rather than estimated from a token counter.
About PointFive
PointFive is the AI Efficiency OS, from the cloud to coding agents. It continuously improves efficiency across cloud infrastructure, data platforms, AI workloads, and coding agents, helping teams understand spend in plain language and ship optimizations at scale. Organizations like Fanatics, H&M, Hertz, Nubank, Citizens Bank and other Fortune 500 companies trust PointFive to create greater transparency into AI spend to uncover opportunities to maximize AI ROI. Learn more at https://pointfive.co.
Media contact:
+1 (339) 242-0393
View original content:https://www.prnewswire.com/news-releases/pointfive-research-finds-cutting-ai-tokens-can-actually-increase-costs-302844453.html
SOURCE PointFive