← Back to blog
David7 min read

Hybrid search explained: vector + full-text + RRF fusion

Why pure vector search misses exact keywords, what full-text search adds, how RRF deterministically fuses two ranked lists, and what it costs.

Hybrid searchRAGRetrieval

When building a knowledge base, teams usually reach for pure vector search first, then quickly hit queries that refuse to come back well no matter how the similarity threshold is tuned. The problem isn’t the threshold; it’s that this one retrieval path has things it fundamentally can’t reach.

This article lays out hybrid search: why it exists, what each path contributes, and how RRF combines the two.

What pure vector search misses

Vector search maps text into a vector space and finds “semantically close” content by distance. It’s strong at fuzzy matching: paraphrases, synonyms, and cross-language similarity.

But it’s insensitive to exact terms. Product codes, contract numbers, proper nouns, names, and rare words often have no stable neighbors in that space. Chinese makes it worse: with poor tokenization, a whole sentence becomes one token, and recall for an exact term becomes a coin flip.

The conclusion isn’t “vector search is bad.” It’s that vector search is good at semantic similarity, not term matching. A knowledge base query needs both.

What full-text search adds

Full-text search (PostgreSQL FTS, BM25, and the like) matches by term: it tokenizes documents into an inverted index and scores queries by which terms hit. It covers exactly the gap vector search leaves — exact keywords, codes, and proper nouns. A hit is a hit.

Chinese adds one precondition: you must tokenize first. PostgreSQL’s default full-text parser splits on whitespace, so a Chinese sentence becomes a single token and the search is effectively useless. To make it work for Chinese you need a tokenizer like zhparser, a PostgreSQL extension built on SCWS.

How RRF combines the two paths

Each path returns a ranking; next you have to merge them. RRF (Reciprocal Rank Fusion, from Cormack et al. 2009) turns each result’s rank into a score and sums them.

Concretely it’s 1 / (k + rank), with k usually 60. A result that ranks 1st on the vector path and 5th on the full-text path scores 1/61 + 1/65. You rerank by total score.

RRF’s value is determinism: no training, no hand-tuned weights between paths, and the same input always produces the same output. For a knowledge base that must be reproducible and traceable, that beats a black-box weighting.

The cost

Hybrid search isn’t free:

  • you maintain two indexes (vector and inverted), both updated on write;
  • queries run both paths and then fuse, so latency is slightly higher;
  • the two paths have to align on the same document IDs.

That cost is usually worth it for a knowledge base that needs both semantic recall and exact hits. But if your scenario only needs fuzzy semantic matching, pure vector is simpler.

How Langhuan implements it

Langhuan’s hybrid search is pgvector embeddings plus PostgreSQL full-text search (zhparser tokenization), with two-path recall fused by deterministic RRF and every result carrying a source anchor. The implementation is verifiable in the Langhuan repository and the architecture docs.

If your knowledge base queries keep asking for both “things that mean roughly this” and “this exact term,” you’re probably already at the point where hybrid search is worth it.

Related · Further reading