← BlogAI Product Development

Context Engineering, Not Prompt Engineering: How Long-Running AI Agents Actually Manage What's in the Window

By Aditya JhaOctober 6, 20269 min read

Context Engineering, Not Prompt Engineering: How Long-Running AI Agents Actually Manage What's in the Window

An agent is given a real multi-step task, audit forty vendor contracts for a specific liability clause, and it runs clean for the first ten, correctly flagging the ones that violate the rule it was given at step two. By step thirty, it's missing contracts it should have flagged, because the exact instruction that mattered got buried under twenty-eight rounds of tool outputs and is now sitting somewhere in the middle of a context window stretched far past where the model attends to it reliably. The model didn't get dumber. Nobody engineered what stayed in its context and what got pushed out, and Anthropic's own engineering team now treats that as a distinct discipline from prompt writing, calling it context engineering.

Why does a long-running agent degrade even though the model itself never changes?

Because of a real architectural constraint, not a bug: transformers create roughly n² pairwise relationships for n tokens in context, so as the context grows, attention stretches thinner across all of it, producing a performance gradient rather than a hard cliff. Anthropic calls this context rot, and the practical effect is exactly what the forty-contract example shows: the model's ability to accurately recall a specific earlier instruction decreases as the surrounding context volume grows, even though nothing about the model's weights changed mid-task.

This is one layer up from the basic lost-in-the-middle problem covered in why your AI agent forgets mid-conversation: that post is about a single long conversation losing track of earlier turns, this is about a multi-step agentic task accumulating tool outputs, reasoning traces and retrieved documents fast enough that the same rot sets in within a handful of steps, not dozens of conversation turns.

What's the actual difference between prompt engineering and context engineering?

Prompt engineering is about how you write and organize the instructions in a single system prompt. Context engineering is the broader discipline of curating and maintaining the optimal set of tokens, system instructions, tool definitions, MCP connections, retrieved external data and message history, at every step of a multi-turn task, not just at the first one.

The distinction matters because a perfectly written system prompt doesn't help once an agent is forty tool calls deep and that prompt is competing for attention against forty rounds of accumulated tool output. Context engineering is the ongoing decision, repeated at every step, of what earns a place in the window and what gets compacted, stored externally, or handed to a sub-agent instead.

How does compaction actually keep a long task from drowning in its own history?

By summarizing and reinitiating instead of letting history accumulate forever. When a conversation or task approaches its context limit, the agent's message history gets passed back to the model to summarize and compress, preserving the details that still matter, architectural decisions, unresolved issues, specific constraints, while discarding redundant tool outputs that already did their job.

The hard part isn't running the summarization step, it's deciding what counts as safe to discard. Anthropic's own coding agent, Claude Code, implements this by explicitly preserving unresolved bugs and architectural decisions through every compaction pass, because losing either of those silently is far more costly than losing a verbose tool output that's already been acted on.

What is structured note-taking, and why can't compaction alone solve this?

When does it make sense to split work across sub-agents instead of one long-running context?

When the task is genuinely parallelizable and the cost of extra tokens is worth the accuracy gain. In an orchestrator-worker pattern, a lead agent decomposes a task into independent subtasks and spawns sub-agents that each get their own clean context window, their own system prompt and their own scoped tool access, then each one reports back only a condensed summary, typically 1,000 to 2,000 tokens, instead of dumping its full working trace into the orchestrator's context.

Anthropic measured this directly: a multi-agent system with Claude Opus 4 as the lead and Claude Sonnet 4 sub-agents outperformed a single Opus 4 agent by 90.2% on internal research evaluations, because the parallel sub-agents could each explore a different angle in their own uncluttered window instead of one agent trying to hold every angle in one increasingly rotted context.

That gain isn't free: the same architecture uses roughly 15 times more tokens than a single chat interaction does, which is exactly the kind of cost that has to be scoped deliberately rather than discovered after the fact, the same scoping discipline covered in why an AI agent's production bill is never the pilot's bill. Sub-agent orchestration is a tool for tasks that genuinely benefit from parallel, independent exploration, not a default architecture for every agent that runs long.

Source: Anthropic Engineering, 2026 — Claude Opus 4 multi-agent system (lead + subagents) vs single-agent Opus 4 on internal research evaluations.
Source: Anthropic Engineering, 2026 — Claude Opus 4 multi-agent system (lead + subagents) vs single-agent Opus 4 on internal research evaluations.

How AIBOOTSTRAPPER helps

There isn't a single published AIBOOTSTRAPPER case study that's specifically a sub-agent research system, so we won't force one. What we will say plainly: every AI Product Development engagement we scope treats context budget as an explicit architecture decision made before the first line of code, the same discipline behind building traceable, evaluable agents covered in how to evaluate an AI agent before it ships, rather than something discovered after an agent starts contradicting itself on step thirty of a real task.

If you're scoping a long-running agent and aren't sure whether compaction, note-taking or sub-agents is the right call for your actual task, book a call, or see our AI product development work for how we approach it.

Want this done for you?

Book a free strategy call and we'll show you how to build and market your business with AI.

FAQ

Questions, answered

Everything you might want to know before we hop on a call.

The ongoing discipline of curating exactly which tokens, system instructions, tool outputs, retrieved data and message history, stay in an AI agent's context window at every step of a task, rather than writing a good prompt once and leaving the rest to accumulate.

Because of context rot: transformers create roughly n² pairwise attention relationships for n tokens, so as context grows, the model's ability to reliably recall any specific earlier detail decreases, a gradual performance gradient rather than a sudden failure.

Compaction summarizes and compresses the existing conversation history in place when it nears the context limit. Structured note-taking writes specific facts to a persistent store entirely outside the context window, so that information survives even a compaction event that would otherwise compress or drop it.

No. Anthropic measured a 90.2% improvement on research tasks from an orchestrator-plus-subagents architecture versus a single agent, but that architecture also uses roughly 15 times more tokens, so it's worth it specifically for tasks that genuinely parallelize, not as a default for every long-running agent.

Keep reading

Let's talk

Ready to build and sell with AI?

Book a free 30 minute strategy call. We'll map the highest ROI AI move for your business, no pitch, just value.