← Back to blog
David18 min read

What Happens When Your Project Knowledge Base Hits Context Limits

Stuffing every document into a Claude or ChatGPT project knowledge base eventually hits context limits — answers degrade first, uploads fail later. Symptoms, causes, and moving knowledge out of the context window.

ClaudeKnowledge BaseRAGMCP

A while back I came across a thread on r/ClaudeAI. Someone asked: what happens if I put close to 200k tokens of material into project knowledge?

One reply stuck with me: then there’s no room left for the conversation. Your model is chatting with you while carrying two hundred pounds of documents on its back.

This is such a common scene. I suspect a lot of people are living it right now without realizing: you upload document after document into the project, things work fine for a while, and then the answers quietly start going wrong.

The first symptom isn’t an error, it’s the model getting dumber

The first sign of hitting the wall is usually not a “context limit exceeded” message. It’s the model slowly getting dumber.

You ask about something the knowledge base clearly covers, and the answer comes back vague. The output format you specified weeks ago stops being followed. You ask the same question twice and get two different stories. Most people’s first reaction is “the model got worse,” or they go fiddle with the prompt. Almost nobody suspects the knowledge base.

Then one day you try to upload a new file, it gets rejected, and you finally see the hard line. By then the degradation has been piling up for weeks.

This isn’t mysticism. A 2023 paper, Lost in the Middle, measured exactly this: information at the beginning or end of the context gets recalled fine; information in the middle gets recalled noticeably worse, a U-shaped curve. The more you stuff in, the more “middle” you create, and the more gets missed.

The context window isn’t a disk, it’s RAM

Understanding this only takes one swap of the metaphor.

200K tokens sounds like storage. In practice it behaves like memory: every time the model answers you, it re-reads the entire context from start to finish. Every document you stuff in is baggage it re-carries on every single turn. A 50-page PDF parses into tens of thousands of tokens; a few dozen of those and most of the window is gone, with conversation history, tool output, and system prompts squeezing into what’s left.

The metaphor isn’t mine. In 2023 Karpathy posted the LLM OS framing: an LLM is the kernel process of a new operating system, the context window is RAM, tool calls are peripherals, and agents are the processes running on top. Anyone who has managed RAM knows the rule: keep resident data minimal.

Worse, the label on the window is optimistic. NVIDIA’s RULER benchmark measures “effective context”: of the models claiming 32K or more, only half could actually handle tasks at 32K. Open-source models typically come in under half their claimed length — IBM’s Granite claims 128K and measures around 32K. So when a window is labeled 200K, budgeting half gets you closer to the truth.

Anthropic’s own engineering blog named the phenomenon context rot: attention is a fixed budget, and every additional token thins what each passage gets. Their principle is blunt — find “the smallest possible set of high-signal tokens” for the job. Don’t try to fill the window.

And the official advice for project knowledge bases? Curate your documents. Not wrong, but read it again: you’re deleting knowledge to fit the window. Knowledge that only shrinks doesn’t support long-term use.

Three ways people self-rescue, and how each one fails

After hitting the wall once, the self-rescue attempts mostly fall into three routes.

Route one: split projects, delete documents. Break the big library into small ones holding only what’s current. It works, but the knowledge shatters — cross-project questions go unanswered, the same document lives in three projects, you update one and forget the other two.

Route two: manual retrieval. Search the wiki or the drive yourself, paste the relevant passage to the model. Works instantly. Then you realize you’re back to doing retrieval by hand, which is the exact chore you adopted an agent to eliminate.

Route three: write your own retrieval scripts. Karpathy later championed an LLM wiki approach: have the model compile raw sources into an interlinked wiki, fetch from it when needed. People have actually done this with Claude Code, and there’s a full walkthrough on YouTube.

One detail here gets quoted out of context. Karpathy himself said he expected to need fancy RAG, then found that compiling a wiki was enough. Note the premise: that’s a personal-scale library. A compiled wiki is itself knowledge moved outside the context and fetched on demand — once the volume grows, the compilation hits the same wall. A company sitting on tens of gigabytes of documents doesn’t get the “my library is small” exemption.

The fix fits in one sentence: knowledge outside, fetch on demand

Strip all three routes down and the disease is the same: knowledge lives inside the context, and no retrieval happens when retrieval should.

The fix is almost embarrassingly simple. The context receives “retrieval results,” never “the whole library.” The knowledge base runs alongside as its own service; documents grow freely and cost the model zero tokens. The agent retrieves before answering, sends in only the few passages relevant to the current question, and every passage carries its source. In a company setting that last part is a hard requirement — “the model said so” doesn’t cut it.

Here’s the satisfying part: this is exactly what Anthropic’s engineering team recommends. They call it just-in-time retrieval — the agent holds only lightweight identifiers, paths and links and queries, and loads concrete content through tools when needed. Claude Code’s glob and grep work this way. They also close the escape hatch: context windows of every size face context pollution, and “waiting for bigger windows” is not a solution.

Databricks ran a controlled experiment too: even when the window is large enough to hold the full text, retrieve-then-answer still produces better answers than stuffing everything in. So the real point of RAG isn’t fashion. It’s an architectural decision: move knowledge from RAM to disk.

This is literally what Langhuan does

Enough theory — here’s where it lands. Langhuan is exactly this layer.

Drop documents in and parsing, chunking, and indexing happen automatically. Chinese retrieval runs the pgvector + zhparser hybrid search path. Agents call retrieval over MCP, and the only things entering the context are the matched passages and their source anchors. It ships as a single binary, and since v1.0.0 there’s a zero-config SQLite mode — spin up a local demo without installing a database.

Multiple agents share one library, one permission model, and one evidence chain, instead of every project carrying its own copy. I verified this structure repeatedly in enterprise settings; the piece about redundant knowledge bases tells that story.

Last thing

When your project knowledge base hits context limits, the prompt isn’t the problem. The knowledge is in the wrong place.

Leave the window for reasoning. Keep knowledge outside and fetch it when needed. The more documents you have, the better the math works out.

Further reading

Related · Further reading