A team scoping a support agent looks at a provider's million-token context window and makes a call that feels like it saves engineering time: skip building a retrieval pipeline, paste the entire product knowledge base into the prompt on every call, let the model sort it out. It works in the demo. But a context window isn't free storage, every token in it is billed, and self-attention means the computation itself grows as the window fills, so the model is now re-reading and re-paying for the same few hundred pages on literally every turn of every conversation. The bill, and in a lot of cases the accuracy, both move in the wrong direction from what the team expected.
Why does a huge context window feel like it should replace RAG?
Because the headline numbers are genuinely large. 1,000 tokens is roughly 750 words or 3 pages of English text; 128,000 tokens is around 384 pages; a million-token window is in the neighborhood of 3,000 pages, which is more text than most companies' entire support documentation. When a window that big is available, stuffing the whole knowledge base in and skipping retrieval engineering looks like the simpler, more future-proof choice.
The part that intuition misses: a context window isn't a filing cabinet you pay for once. In a standard transformer, every token has to relate to every other token in the window through self-attention, so the window costs you twice, more compute to run, and, per several 2026 long-context studies, often worse reasoning as it fills. Size and usefulness are not the same curve.
What does RAG actually do differently, mechanically?
It replaces "give the model everything and let it find the answer" with "find the answer first, then give the model only that." A RAG pipeline splits source documents into chunks, converts each chunk into a vector embedding, a numeric representation of its meaning, and stores those vectors in a vector database. At query time, the user's question gets embedded the same way, and the system runs a cosine similarity search to find the chunks whose vectors sit closest to the question's vector in that embedding space, the mathematical proxy for "most semantically relevant."
Only the top handful of matching chunks, typically 3 to 10, get inserted into the prompt alongside the question, often after a reranking step that re-scores those candidates with a more precise model before the final few are sent. The model never sees the other 99% of the knowledge base on that call. It's the structural difference behind why our own ComplySpark build grounds every generated document in the client's actual policy library instead of a model trying to hold the whole library in working memory at once.
Does a bigger context window at least perform better, or just feel safer?
- **Raw volume hurts accuracy on its own.** One 2025 study found Llama 3's HumanEval coding accuracy dropped by roughly half at 30,000 tokens compared to its short-context baseline, with similar declines on GSM8K math reasoning, and the drop tracked total input volume, not where the relevant content sat in the prompt.
- **Position matters too: the "lost in the middle" problem.** Multiple transformer models studied show a U-shaped attention pattern, attending more reliably to content at the very start and very end of the context while underweighting the middle. If a long-context prompt buries the one relevant paragraph on page 140 of 300, the model is structurally less likely to use it well, regardless of how much room is technically available.
- **Returns diminish long before the advertised limit.** A study across 13 long-context models found most peaked around 20,000 tokens for in-context learning and got no better past that point, meaning a 1-million-token window advertised capacity and its useful capacity are two very different numbers.
What does the cost actually look like, turn by turn?
Because the model is stateless, every turn has to resend the full conversation and context, so total session cost tracks turns multiplied by average context size, not just how much useful work got done. Take an illustrative 300-turn agent session where context grows roughly 1,000 tokens per turn up to 300,000 tokens, an average of 150,000 tokens resent across the session: that's 45 million input tokens total. At Claude Sonnet 5.5's published rate of $2 per million input tokens, with no caching, that single session costs about $90 on input tokens alone.
Prompt caching helps, a lot: Sonnet 5.5's cache-read rate is $0.10 per million tokens, roughly 5% of the base input rate, which brings that same 45-million-token session down to around $4.50. But caching only lowers the per-token rate, it doesn't change the fact that cost still scales with how much history is in the window, every turn. A RAG-based version of the same 300-turn session, retrieving only the top 5 relevant chunks, roughly 2,500 tokens, per turn instead of resending the growing history, runs about 750,000 input tokens total, around $1.50 at the same rate. The full-context path costs 20 to 60 times more than retrieval for the same conversation, and per the accuracy data above, isn't even reliably more accurate.

So does that mean you should never use a long context window?
No, and treating this as a rule rather than a tradeoff is its own mistake. A single, bounded document under roughly 20,000 tokens, a contract, a report, a one-off Q&A session, often doesn't justify the engineering cost of a retrieval pipeline at all; just put it in context and ask. RAG introduces its own real failure modes too, a chunk boundary that splits a critical sentence in two, or a retrieval step that misses the one paragraph a multi-step reasoning task actually needed, the exact brittleness that breaks naive single-pass RAG on genuinely multi-step agent tasks.
The actual decision is about shape, not size: a single bounded document favors long context, a corpus that grows over time, gets updated independently of any one conversation, or is larger than a handful of documents favors retrieval, because that's the point where resending the whole thing on every turn stops being a convenience and starts being the single largest line item in the system.
How AIBOOTSTRAPPER helps
ComplyNexus, the RAG-powered compliance platform we built, is the architecture this post describes in production: a vector database holding the client's full regulatory and policy corpus, with only the handful of genuinely relevant clauses retrieved and surfaced per query, not the whole library resent on every call. That design is a direct contributor to the 92% cut in manual review time and the drop from a 3-week to a 2-hour regulatory change turnaround the client has seen, because the system stays fast and cheap precisely because it never asks the model to re-read everything it already knows.
If you're scoping an AI product and trying to decide whether it needs a real retrieval layer or can get away with a long context window, book a call, or see how we approach it in our AI product development work.
Want this done for you?
Book a free strategy call and we'll show you how to build and market your business with AI.
Sources and further reading
- 1.Redis — Context windows in AI: why every token is a budget decision (2026)
- 2.FlowHunt — What an AI Agent's Context Window Actually Costs You
- 3.ACL Findings 2025 — Long-context accuracy degradation study (Llama 3, HumanEval/GSM8K)
- 4.arXiv — Lost in the Middle: How Language Models Use Long Contexts
- 5.arXiv — Long-context model in-context learning performance study
- 6.Claude — Pricing (Sonnet 5.5 input/output/cache rates)
