Context Engineering: Managing Token Budgets and Attention in Long-Running Agents
How to manage AST-level code contexts, sliding execution traces, prompt compression, and attention degradation during complex autonomous refactors.
The "Infinite Context" Myth
With modern LLM context windows expanding to 1 million and even 2 million tokens, many developers assumed context management was obsolete. The prevailing assumption was simple: just dump the entire codebase into the prompt and let the model figure it out.
In practice, this approach fails in three critical ways:
- The "Needle in a Haystack" Attention Tax: While models can locate specific facts in massive contexts, their reasoning degradation across multi-step dependency graphs increases linearly with context length.
- Economic and Latency Penalties: Feeding 200,000 tokens on every step of a 30-turn agent execution consumes hundreds of millions of tokens and introduces intolerable 15-second latency spikes.
- Recency Bias and Hallucination Accumulation: As execution logs, tool returns, and intermediate diffs accumulate, the model's self-attention disproportionately weights stale intermediate failures over the initial design contract.
Context Engineering is the practice of dynamically constructing, pruning, and optimizing the agent's context window on every turn to maximize reasoning density while minimizing token waste.
The Core Principles of Context Engineering
Raw Codebase (50,000 LOC / 400,000 Tokens)
│
AST Symbol Extractor & Dependency Graph
│
▼
Target Context Envelope (~12,000 Tokens)
┌────────────────────────────────────────────────────────────────────────┐
│ [System Rules & Safety Guardrails] ~1,200 tokens │
│ [Contract & API Schemas] ~2,000 tokens │
│ [AST Type Interfaces of Touched Files] ~3,500 tokens │
│ [Surgical Code Window (Active Focus Area)] ~2,500 tokens │
│ [Sliding Rolling Tool History (Last 3 Steps)] ~2,800 tokens │
└────────────────────────────────────────────────────────────────────────┘1. AST Skeleton Extraction vs. Raw File Dumping When an agent needs to consume a file to understand how to invoke an exported class or hook, it almost never needs the internal implementation of private helpers.
Instead of reading the raw 500-line file, our context engine strips function bodies, preserving only: - Exported interfaces and types - Public method signatures with docstrings - Dependency imports
// AST Skeleton (32 tokens) - 77% compression with zero loss in interface utility export function calculateTax(amount: number, region: string): number; ```
By feeding AST skeletons for 90% of repository dependencies and loading raw file bodies only for active edit targets, agent token footprints plummet without degrading semantic comprehension.
2. Sliding Tool Return Buffers
During a software build, tool outputs produce massive amounts of transient information: directory listings, grep outputs, package installation progress bars, and test runner outputs.
If left unmanaged, a single npm install log can consume 8,000 tokens.
The Compression Hierarchy
- Immediate Turn (t=0): Full diagnostic output with highlighted errors.
- Previous Turn (t-1): Summarized result (e.g., ✓ 42 tests passed, 0 failures, 1.2s).
- Historical Turns (t > 2): Compressed event ledger entry (e.g., Step 3: Ran Jest suite -> Passed).
By pruning the output of completed steps and retaining only their high-level semantic status in the history buffer, the model's attention remains laser-focused on the task currently underway.
3. KV-Cache Alignment & Static Prompt Prefixes
In modern inference APIs, Prompt Caching (KV-cache reuse) provides up to 90% discount on cached tokens and dramatically lowers Time To First Token (TTFT).
To exploit KV-cache reuse, context engineering requires strict ordering:
- Frozen Prefix (Immutable): Core system prompt, base tool schemas, environment definitions.
- Semi-Static Tier (Rarely Changed): Project specification, API contract schemas, directory tree.
- Dynamic Tail (Mutates Every Turn): Sliding conversation history, latest tool outputs, active diffs.
[ FROZEN PREFIX ] ─── Cache Hit (100%) -> Instant TTFT, 90% cost savings
[ SEMI-STATIC TIER ] ─── Cache Hit (95%)
[ DYNAMIC TAIL ] ─── Evaluated dynamicallyIf you accidentally prepend dynamic timestamps or changing session IDs at the top of the prompt, you bust the entire KV cache on every turn. Structuring context into stable prefix blocks is essential for real-world production viability.
Lessons from Qbit's Multi-Agent Context Loops
In Qbit's multi-agent architecture, context engineering is what enables our specialized agents to work seamlessly in parallel.
When the Central Hub Agent delegates a task to the Frontend Agent, it doesn't pass the entire project history. It synthesizes a Scoped Context Envelope: - Exactly the relevant UI components - The typed backend schema generated by the Backend Agent - The design system tokens
This guarantees the Frontend Agent doesn't hallucinate non-existent API routes or drift from the design tokens.
Conclusion
Context engineering is software architecture applied to LLM reasoning. As models grow more capable, the systems that win won't be those with the largest brute-force prompts—they will be the systems that orchestrate high-density, low-entropy context with mathematical precision.
Related: [Evaluating Coding Agents Beyond SWE-bench: Why Unit Tests Lie](/blog/engineering/evaluating-autonomous-agents).
Autonomous software engineering in practice.
Every architectural principle described in this dispatch—deterministic microVM sandboxing, contract synthesis, and specialized multi-agent coordination—is active in Qbit. Build production Next.js apps with natural conversation.
Launch Qbit SystemWhy We Built Qbit: Moving Beyond Code-Completion to Autonomous Product Engineering
Why code autocomplete hit a ceiling, and how we engineered an autonomous multi-agent harness to turn natural language conversations into production-grade Next.js applications.
What is an Agent Harness? The Missing Infrastructure in Coding Agents
Why LLMs alone fail at real-world software engineering, and how execution sandboxes, state management, and tool feedback loops form the harness that makes autonomous coding reliable.