Langhuan separates model connections from model capabilities, letting different providers offer embedding, rerank, and other abilities, and runs a rerank pass after hybrid recall. Retrieval is no longer bound to a single model or vendor, and connection config no longer dictates how business code is written.
One provider, many capabilities
The same model service may offer embedding, rerank, or future capabilities at once. This version splits “where to connect” from “what the connection can do”:
- The model catalog shows the capabilities a provider actually offers, not just a connection name.
- At runtime, provider descriptors are built from registered capabilities; missing capabilities fail early with a stable error.
- Connection and model config save a snapshot, so later config changes don’t silently alter the behavior of knowledge bases already created.
Judge once more after hybrid recall
Vector recall and PostgreSQL full-text search are good at quickly widening the candidate set; rerank is responsible for re-judging semantic relevance within that set. The retrieval path can now be expressed as:
vector + full-text → RRF hybrid recall → rerank → parent context
The lifecycle, config, and connection management of the rerank model can all be done in the Web Console, and the search API and MCP tools share the same policy contract.
Easier to locate result differences
The retrieval stage logs structured rerank and recall events, helping developers tell apart “nothing was recalled” from “recalled but dropped by rerank.” For debugging RAG, that’s more useful than seeing only the final answer.