In the rapidly evolving landscape of software development, the deployment of autonomous coding agents has moved from experimental novelty to a cornerstone of modern consultancy workflows. At Straight Up AI, the development of an internal "control plane"—a specialized architecture designed to orchestrate these agents—has yielded significant productivity gains. However, a recent seven-week internal audit has revealed a sobering reality: while the agents are undeniably capable, the underlying economic architecture of these systems is currently unsustainable.
As the industry pivots toward agentic workflows, the hidden costs of "context bloat" and inefficient agent orchestration are emerging as the next great hurdle for AI-integrated engineering.
The Economics of Agentic Orchestration
At Straight Up AI, the latitude granted to coding agents is traditionally determined by three primary factors: complexity of the task, developer oversight, and project constraints. Yet, in the pursuit of building a functional system, one critical factor was initially omitted: the cost of construction.

Operating as a lean consultancy, the team utilized a Claude Max subscription (priced at £200 monthly) to prototype their control plane. While this flat-fee model provided a low-risk environment to prove the system’s efficacy, it masked the true cost of operation. Without the discipline of per-token API billing, the team was essentially building in a vacuum. To assess viability, the team conducted a deep-dive analysis of 44 development cycles executed between July 15 and September 4, 2024.
The results were stark. By converting actual usage into equivalent API token costs—accounting for cached reads (0.1x), cached writes (1.25x), and output tokens (5x)—the analysis revealed that at enterprise API rates, the current system would be 22 times more expensive than the flat-rate subscription. For a growing consultancy, this represents a financial burden nearly equivalent to the salary of a mid-level engineer, rendering the current iteration of the system commercially non-viable at scale.
Chronology of an Audit: Uncovering the Efficiency Gap
The audit was triggered by more than just financial concern; it was a matter of latency. Developers noticed that certain segments of the agentic workflow felt sluggish, particularly the "adversarial review" process—a mechanism designed to spin up a secondary coding agent to challenge every commit.

The Adversarial Overhead
The team hypothesized that the delay was caused by a single, monolithic, and computationally expensive review agent. However, the data told a different story.
The overhead was not caused by the intensity of one agent, but by the sheer volume of agents. The data showed a 1.19x ratio of reviewers to implementers. When an adversarial reviewer flagged an issue—which occurred in 26% of cases—a remediation agent was dispatched, triggering yet another adversarial review cycle. This created a recursive loop that often led to "rabbit-holing," where agents focused on superficial code fixes rather than stepping back to evaluate structural requirements.
The Controller Bottleneck
The control plane architecture is centered on a "Controller" agent that remains active throughout the job, acting as the arbiter of truth between the implementer and the requirements. While this architecture successfully prevents "hallucinated progress"—a common issue in standard plan-and-build setups—it became the primary engine of cost.

The analysis found that 75% of runs never triggered the compaction protocols intended to keep context windows lean. Consequently, the Controller’s context grew linearly with the duration of the session. By the end of the audit, the largest context window observed reached nearly one million tokens.
Supporting Data: Where the Money Goes
The audit shattered several internal assumptions regarding what makes an AI system expensive.
| Metric | Adversarial Review | Implementation | Ratio |
|---|---|---|---|
| Agents Dispatched | 242 | 204 | 1.19x |
| Cost Units | 234.8M | 339.6M | 69% |
| Agent-Hours | 28.0 | 39.2 | 71% |
While the team assumed the planning documents (architecture specs, technical briefs) were the primary drivers of context usage, the data proved otherwise. The entire stack of planning documents accounted for a mere 2% of the overhead.

Instead, the culprit was the Controller’s own output. Tool-call arguments and "dispatch briefs"—the instructions sent to agents—accounted for 37.7% of the total accumulated context. Because the system was designed to persist these briefs for human readability, the Controller was effectively re-reading thousands of tokens of completed work on every single increment.
Implications for System Architecture
The findings from Straight Up AI suggest that the path forward requires a fundamental shift in how agentic memory is handled. To move toward a sustainable model, the team identified three strategic pivots:
1. Passing by Reference, Not by Value
Drawing inspiration from low-level programming, the team proposes a "pointer-based" memory architecture. Rather than forcing the Controller to hold the entire history of a run in its context window, it should hold only the most recent state and a reference to the persistent run record. This ensures the agent maintains full situational awareness without the exponential cost of redundant data.

2. Compaction at Boundaries
The audit revealed that compaction was being implemented in the wrong place. While individual workers were compacting, the Controller was carrying the weight of every worker’s history. By implementing "sawtooth" compaction—where the Controller resets its context at the boundary of every terminal increment—the team projects a reduction in per-turn context from 360,000 tokens to approximately 100,000.
3. Scoped Review Funnels
To solve the adversarial overhead, the team intends to shift from an exhaustive review model to one scoped by attempt. By ensuring the second review only examines the diffs generated by the remediation agent, the team expects to significantly lower the agent-to-reviewer ratio, preventing the system from falling into the recursive loops that characterized the audit period.
Conclusion: The Necessity of Honest Analysis
The experience at Straight Up AI serves as a cautionary tale for any organization rushing to integrate LLM-based agents into their production pipelines. The "magic" of AI often obscures the mechanics of the engine, leading to systems that are performant in the short term but financially ruinous in the long term.

The audit was a humbling exercise. The team’s initial hypotheses—that compaction was occurring or that planning documents were the primary bloat—were both proven incorrect. By setting aside these "priors" and conducting a data-driven analysis, they uncovered the true source of inefficiency: the system’s failure to shed completed metadata.
As the consultancy looks to the future, the goal is clear: to maintain the high-quality output of their coding agents while enforcing a discipline of ephemeral memory. For AI-driven development to become a standard tool rather than a luxury, the industry must stop treating compute as an infinite resource and start building systems that respect the economic reality of the token.








