← Back to blog
David9 min read

By the third enterprise agent, the knowledge base can't be rebuilt every time

When building several enterprise agents, I found the thing most often rebuilt isn't model capability — it's the ingestion, access control, retrieval, and evidence chain for knowledge.

Enterprise agentsKnowledge baseRAGMCP

When the first enterprise agent project ships, building the knowledge base directly into the application is usually the right call.

There isn’t much data to connect, and the scenario is clear: find a few documents, parse, split, and index them, then hand the retrieval results to the agent. The team can validate business value quickly, and the engineering boundary stays tight enough. The problem doesn’t start with the first agent.

The problem shows up after the second and third scenarios arrive one after another: policy Q&A for employees, product advisory for customers, project assistants for delivery teams — all of them suddenly need to “understand company knowledge.” Each project can wire up a knowledge base quickly; but soon you realize what’s being rebuilt across teams is the same stretch of foundational capability.

What’s duplicated is never just a document

The most visible duplication is data-source ingestion. The same product materials, policies, and project documents get uploaded, parsed, and synced separately by different systems. Then comes duplication that’s harder to see:

  • each team decides for itself how to clean and chunk documents;
  • each team maintains its own vector index and full-text search;
  • each team handles access control across departments, customers, or projects;
  • each team debugs “why didn’t this retrieve anything” on its own;
  • each team handles document updates, retries, and version history on its own.

At first, all of this looks like part of the business implementation. But as agents multiply, the problem shifts from “we rewrote some code” to “the enterprise now has several inconsistent sets of knowledge facts.”

After the same policy updates, one agent already sees the new version while another still quotes the old one; a retrieval result can’t say which document or version it came from; permission rules are scattered across applications, and during troubleshooting it’s hard to tell whether the problem lies in the data, the retrieval, or the agent itself.

What should be shared is the capability to process knowledge

I later reframed this: what an enterprise shares isn’t some document an agent already ingested, but the process of turning documents into usable knowledge.

This layer should at least handle uniformly:

  1. document ingestion, versioning, and asset archiving;
  2. permissions and isolation across different workspaces;
  3. chunking, vector retrieval, and full-text retrieval;
  4. the source anchor and evidence returned by every retrieval;
  5. updates, failure retries, and observable state.

The upper-layer agents keep focusing on what’s genuinely different about them: understanding the task, calling tools, organizing the flow, prompts, and the final answer. They consume knowledge capability rather than each owning a separate knowledge-processing pipeline.

This doesn’t mean everything has to be broken into a platform. With only one scenario, an in-app implementation is still the faster choice. The real tipping point is this: when multiple teams, projects, or customers start repeatedly building the same kind of knowledge-ingestion and retrieval capability, it’s already infrastructure, not just an internal detail of some agent.

Langhuan sits on this boundary

Langhuan doesn’t handle agent orchestration, and it doesn’t generate final answers. It turns documents into retrievable, traceable knowledge and exposes it to different upper-layer apps over REST and MCP.

Concretely, it puts workspace isolation, document versioning, hybrid retrieval, and source anchors on the same knowledge-processing chain. That way, multiple agents can work from the same set of facts; when a recall problem comes up, you can trace it back to a specific document and location, instead of staring at an unexplainable chunk of context.

The capability boundaries and architecture of the project can be checked in the Langhuan repository and the architecture docs.

If your team is moving from one agent toward several, there’s a question worth asking first: should the next knowledge base really be built from scratch inside the next application?

Related · Further reading