The rapid adoption of AI coding agents—tools like Claude Code, Cursor, and various GitHub Copilot integrations—has transformed the software development lifecycle. Developers are shipping features faster than ever, yet a silent crisis is emerging within engineering departments: the "Token Tax."
For most AI coding agents, the primary activity is not "thinking" or high-level architectural reasoning. It is massive, repetitive Input/Output (I/O) operations. Whether it is scanning five sprawling files to answer a single question about a method, generating boilerplate test files that mirror existing patterns, or updating documentation after a sync, the cost is mounting. Each of these actions consumes thousands of tokens, frequently offloaded to frontier models that are significantly overqualified for the task.
As engineering leaders look at their monthly burn rates, the seat license cost is increasingly becoming a secondary concern. The real financial hemorrhage is the token consumption. By 2028, Gartner projects that AI coding costs will surpass the average developer’s salary, a trend driven by the unrestrained use of high-end LLMs for low-level grunt work. With a quarter of engineering leads already reporting monthly spends of $200–$500 per developer—and some exceeding $2,000—the imperative to optimize these workflows has never been higher.
The Chronology of the "Shunt" Strategy
The realization that developers were burning expensive frontier tokens on mundane tasks led to a surge in internal "routing" experiments. Initially, developers attempted to manage this through documentation—specifically by adding advisory rules to CLAUDE.md files.
These early attempts followed a predictable, yet flawed, trajectory:
- The Advisory Phase: Developers would add instructions to project-level configuration files, asking the AI to "self-route" specific tasks to cheaper models or manual processes.
- The Failure of Soft Enforcement: These rules were inherently advisory. High-performance models like Claude 3.5 Sonnet or Opus, when prompted to solve a complex problem, often found it easier to "ignore" the routing rules in favor of brute-forcing the context window.
- The Infrastructure Burden: The alternative—building custom "platform teams" to manage internal routing servers—proved too costly and slow for the fast-paced nature of modern dev-ops.
The breakthrough came with the emergence of "ephemeral runtimes"—specifically the AiKA Modes feature within Spotify’s Portal. By treating agents as declarative, serverless functions rather than long-running, stateful servers, developers finally had a way to offload tasks to specialized, lower-cost models like Gemini 2.5 Flash without the overhead of managing infrastructure.
Supporting Data: Why Routing Matters
The mathematical case for routing is compelling. When an AI agent reads multiple large files to understand a codebase, it incurs costs proportional to the size of those files. By routing that I/O to a leaner, "worker" model—a model designed for extraction and summarization rather than deep reasoning—the token overhead can be slashed by as much as 90%.
The Efficiency Benchmarks
Testing conducted across a Java monorepo revealed stark differences in efficiency:
- Bulk-Reading: When Claude was instructed to delegate file-scanning to a
bulk-readermode, mean token consumption for code analysis dropped by roughly 90%. - Code Writing: For boilerplate generation, the savings were even more profound. By offloading the generation of test stubs and config files to a secondary model that writes directly to the disk, the primary agent (Claude) never enters the loop for the generation process, effectively removing the output token cost entirely.
However, the data also highlights the "latency tax." Each delegation involves a network round-trip. While the worker model is faster at processing, the total time for a task can increase by 10 to 30 seconds due to the orchestration overhead. For small tasks, this is inefficient; for large-scale analysis, it is a massive net gain.
Official Perspectives and Technical Implementation
The industry is beginning to coalesce around the idea of "Model Routing." At the heart of this movement is the shunt plugin for Claude Code. This plugin introduces a multi-layered approach to delegation that shifts the burden from the human developer to the infrastructure.
Layer 1: The Pre-Tooling Hook
The shunt plugin operates at the hook level. It fires before every Read tool call, analyzing the file size. If a file exceeds a threshold (e.g., 350 lines), the hook blocks the read and redirects the agent to the bulk-reader skill. This ensures that even if the agent "forgets" its instructions, the system enforces the token-saving protocol.
Layer 2: Ephemeral Scripting
The system utilizes bash scripts that wrap Portal CLI calls. These scripts are strictly one-shot: they build the request, invoke the action, unwrap the output, and report usage data. Crucially, the "context" never persists on the server. The primary model (Claude) remains a stateless orchestrator, while the worker model (Gemini 2.5 Flash) does the heavy lifting in an isolated, ephemeral environment.
Layer 3: Skill-Based Guidance
Skills are essentially metadata files that provide the "what and how" for the agent. By embedding these skills into the development environment, developers can ensure that when a block occurs, the agent is immediately given the syntax required to invoke the correct worker mode.
Implications for the Future of AI Engineering
The shift toward model routing has profound implications for how companies will manage their AI budgets in the coming years.
The "Overqualified Model" Problem
The current industry standard—throwing a frontier model at every task—is unsustainable. We are seeing a shift where "Reasoning Models" (the expensive, highly capable LLMs) are increasingly relegated to the role of "Architects," while "Worker Models" (cheaper, faster LLMs) act as the "Laborers." This division of labor mirrors real-world engineering teams, where a senior engineer (the frontier model) dictates the strategy, and junior or specialized tasks are delegated to automated pipelines.
The Death of Infrastructure Bloat
Perhaps the most significant takeaway is that routing no longer requires a specialized platform team. Because tools like AiKA Modes allow for declarative configuration, the "platform" is essentially the code itself. Developers can fork, share, and update their agent modes just as they would any other piece of software. This democratizes the ability to optimize AI costs, moving the power from the finance department to the individual contributor.
Limitations and Critical Boundaries
It is essential to acknowledge what this architecture cannot do:
- Editing remains human-proxied: The worker models used for summary and analysis often lack the precision to suggest specific line-level edits. If the agent needs to refactor code, it must still read the specific section directly.
- Reasoning is off-limits: Architectural decisions, debugging subtle race conditions, and safety-critical logic are areas where routing fails. These require the full "thought" process of a frontier model.
- The Latency Floor: For small tasks, the overhead of the network round-trip makes routing counterproductive. A successful routing strategy requires intelligent thresholds—knowing when to stop being "clever" and simply let the agent do the work.
Conclusion: The Path Forward
The "Token Tax" is a symptom of a maturing industry. As we move past the novelty phase of AI coding agents, we are entering the era of optimization. The tools to build these routers are already available, and the financial incentives are becoming impossible to ignore.
For the modern engineering leader, the strategy is clear: Stop paying for "thinking" when you are really paying for "reading." By implementing tiered routing and embracing ephemeral, model-agnostic runtimes, teams can drastically lower their operational costs while simultaneously improving the performance of their AI workflows. The future of software engineering isn’t just about using AI—it’s about managing the intelligence budget with the same rigor we apply to our cloud infrastructure.








