Council Post: The Hidden Cost Of AI Coding: Why Context Matters More Than Tokens
tags:Kumar Chivukula leads Opsera's innovation in Agentic DevOps and AI-SDLC, empowering secure, AI-driven software delivery.

getty
Over the past year, my conversations with technology leaders have shifted noticeably. Back then, we asked: Can large language models write faster and better code, generate documentation and run tests?
Today, it’s more compelling: What is all of this actually costing us?
Early adopters are finding that AI costs can climb much faster than expected. Uber, for example, said it exhausted its 2026 budget for Claude Code by April, while one tech investor reported that his agentic AI workflows were generating roughly $300 per day in API costs.
These stories have triggered an important conversation about AI spending. The issue is not how many tokens organizations consume; it’s whether that consumption translates into meaningful software delivery outcomes.
Many organizations are discovering how AI increases activity more quickly than it increases value. Code generation, analysis, testing and documentation all happen faster. Yet engineering teams still struggle with rework, alignment and delivery bottlenecks. As a result, token consumption is rising while the returns remain difficult to measure.
Tokens As A New Production Resource
Enterprise software organizations already have mature frameworks for managing labor costs, cloud infrastructure and software licensing. Every investment is tied to expected outcomes, and leaders know what they spend and expect to receive in return.
AI poses a different dynamic. Every prompt, model interaction and automated workflow consumes tokens, creating a new operational spend category. As adoption expands, token consumption is becoming as measurable as cloud usage or infrastructure costs.
Building A Value Framework
Reflecting on the cloud era, when adoption accelerated, many organizations prioritized consumption over outcomes. The result was sprawl, rising costs and elusive return on investment. Over time, organizations learned the goal was not simply to reduce spend, but to understand the value each unit of compute produced.
The same shift is happening with AI.
Organizations that succeed with AI will not be those that consume the most tokens, but the ones that understand how token consumption translates into business outcomes.
Right now, most organizations have no answers.
More Output Vs. More Value
Most organizations measure AI progress today without knowing if it’s actually working. They measure the number of users, licenses activated, prompts run and lines of code generated. These are useful signals, but they say nothing about performance.
The output-to-value gap shows up in sprint reviews. Sure, a team can generate thousands of lines of code in a week, but it takes two weeks to review, rewrite and realign it with existing architectures and security requirements. As a result, token consumption increases while delivery velocity remains low.
The problem is that AI doesn't reduce the cost of rework; it accelerates the rate at which rework occurs. Every ambiguous prompt can produce confident-sounding output. But it has to be validated, corrected or discarded downstream. This waste is invisible to usage dashboards that track what went in, but not what came out.
We must start measuring AI yield: the percentage of generated output that reaches production without rework.
The Root Cause: Context Loss
Understanding why AI yield is lower than expected requires looking at something not yet examined: What do AI systems actually know when they are asked to perform work?
Most AI coding tools operate with limited awareness of organizational context. They may have access to a codebase, but not legacy architectural decisions, compliance requirements or the business rules. Most importantly, AI often doesn’t understand the intent of feature requests.
As a result, context has to be recreated throughout the software delivery process. Engineers reexplain requirements. Teams reestablish architectural constraints. Domain knowledge gets reintroduced into prompts and workflows. The result is an organizational context reconstructed repeatedly.
This becomes an economic issue because too much token usage is spent rebuilding context that already exists. This cost doesn’t appear as a separate line item—it accumulates across teams, projects and delivery cycles.
That’s why “AI tokenomics” cannot be solved through budget controls alone.
Why The Real Bottleneck Was Never Typing Speed
The first wave of AI coding tools solved a real problem: developers could produce code faster. But software delivery was never primarily bottlenecked by code generation speed; it’s always been an alignment issue. It’s hard to ensure everything else remains connected throughout the delivery process.
Closing those gaps consumed far more time and effort than writing the code itself.
Software can now be produced faster than ever. However, AI has created a growing disconnect between activity and outcomes. More code is generated, but not necessarily more value.
Organizations that address this early can approach AI differently by focusing less on generation speed and more on preserving intent, reducing rework and ensuring context remains connected to execution.
When Context Is Infrastructure
Measuring token consumption alone is incomplete. Instead, a more useful signal is how AI activity is converting into production-ready outcomes. How much rework is needed, how quickly do features move from intent to deployment and how often must teams recreate context before meaningful work begins?
The organizations seeing the greatest returns from AI are beginning to look beyond usage metrics and focus on outcome metrics.
Recognizing that context can no longer be treated as something that lives in meetings, documents, chat threads or individual employees' heads is now vital.
As AI becomes part of the delivery process, context increasingly functions as infrastructure.
All existing software delivery elements must remain accessible so they don’t have to be repeatedly reconstructed, and so organizations don’t pay for re-contexting tax through rework, misalignment and unnecessary token consumption.
The Next Phase Of AI Adoption
As AI becomes more deeply embedded in software delivery, three capabilities will become increasingly important: persistent intent that endures beyond individual prompts and sessions, shared context that can be reused across teams and workflows and governed execution that links AI-generated outputs to approved objectives and constraints.
Organizations that address these issues can gain real value from every token they consume.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?