Hybrid search: embeddings plus keyword, then rerank
— Search, Postgres, AI — 7 min read
Most teams that add "AI search" start the same way. Embed every document, embed the query, return the nearest neighbours. The demo looks great. Then someone searches for a SKU, an error code or a person's surname, and the top result is something vaguely related that doesn't contain the string they typed.
Keyword search has the opposite problem. It finds ERR_CONN_RESET instantly but returns nothing for "the connection keeps dropping".
You usually want both. In Postgres that means full-text search and pgvector side by side, merged with reciprocal rank fusion, with an optional reranker on the short list.
What each retriever is good at
Keyword search (BM25, or Postgres tsvector) scores documents by the query terms they contain. Rare terms count for more. It is precise, explainable and cheap. It knows nothing about meaning, so synonyms, paraphrases and typos fall through.
Vector search compares a query embedding to document embeddings. It handles paraphrase, rough descriptions and cross-language queries well. It is weak at exact tokens: product codes, version numbers, names, anything the embedding model never saw often enough to give a distinct vector. It also always returns something, even when nothing relevant exists.
Their failures don't overlap much, which is why combining them works.
The table
One table holds both representations. Postgres keeps the tsvector in sync through a generated column, so there is nothing extra to maintain on write.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE docs (
id bigint PRIMARY KEY,
title text NOT NULL,
body text NOT NULL,
embedding vector(1536) NOT NULL,
search_tsv tsvector GENERATED ALWAYS AS (
setweight(to_tsvector('english', title), 'A') ||
setweight(to_tsvector('english', body), 'B')
) STORED
);
CREATE INDEX docs_tsv_idx ON docs USING gin (search_tsv);
CREATE INDEX docs_vec_idx ON docs USING hnsw (embedding vector_cosine_ops);Use whatever dimension your embedding model produces. Title terms get weight A so a match in the title beats a match buried in the body.
A note on scoring: Postgres's built-in ts_rank_cd is not BM25. It doesn't use corpus-wide term statistics the same way. For many use cases it is fine. If you need real BM25 inside Postgres, extensions such as ParadeDB's pg_search provide it. The fusion step below only uses rank positions, so you can swap the keyword scorer later without changing anything else.
Merging two lists with reciprocal rank fusion
The two retrievers produce scores on different scales. A cosine distance of 0.21 and a ts_rank_cd of 0.4 can't be compared or added. You could normalise them, but the normalisation shifts with every query.
Reciprocal rank fusion (Cormack, Clarke and Büttcher, 2009) ignores the scores and uses only positions. Each document gets 1 / (k + rank) from every list it appears in, and the contributions are summed. k = 60 is the constant from the paper and works well as a default.
A document that shows up near the top of both lists beats one that tops only a single list. That is usually what you want.
Here is the whole thing as one query. $1 is the raw query text, $2 is the query embedding.
WITH kw AS (
SELECT id, row_number() OVER (ORDER BY rank DESC) AS r
FROM (
SELECT id, ts_rank_cd(search_tsv, q) AS rank
FROM docs, websearch_to_tsquery('english', $1) AS q
WHERE search_tsv @@ q
ORDER BY rank DESC
LIMIT 50
) t
),
vec AS (
SELECT id, row_number() OVER (ORDER BY dist) AS r
FROM (
SELECT id, embedding <=> $2 AS dist
FROM docs
ORDER BY dist
LIMIT 50
) t
)
SELECT id, sum(1.0 / (60 + r)) AS rrf_score
FROM (SELECT * FROM kw UNION ALL SELECT * FROM vec) AS both_lists
GROUP BY id
ORDER BY rrf_score DESC
LIMIT 20;A few details matter:
- Each inner query does its own
ORDER BY ... LIMITso the HNSW index can stop early, and ranking happens on at most 100 rows. websearch_to_tsqueryaccepts user input safely. Quotes and-termwork the way people expect, and stray punctuation doesn't throw a syntax error.- If you filter (tenant, language, status), apply the filter inside both branches. With HNSW, a selective filter can leave you with fewer than 50 vector results. pgvector's filtering notes cover iterative scans and partial indexes for this.
Adding a reranker
Fusion gets the right documents into the candidate list. A reranker puts them in the right order.
Embedding search uses a bi-encoder: query and document are embedded separately and compared with a dot product. That is fast, but the model never sees the two texts together. A cross-encoder reranker takes the query and one candidate as a single input and outputs a relevance score. It is far more accurate and far too slow to run over the whole corpus. Running it over 20 to 50 fused candidates is affordable.
const candidates = await hybridSearch(query, queryEmbedding) // the SQL above
const scores = await reranker.score(
query,
candidates.map(c => c.body),
)
const top = candidates
.map((c, i) => ({ ...c, score: scores[i] }))
.sort((a, b) => b.score - a.score)
.slice(0, 10)reranker stands for whatever you use: a hosted rerank API or a self-hosted cross-encoder from the sentence-transformers family. Measure the latency. A reranker adds a model call to every search, and you may want to skip it for short, exact-looking queries.
The reranker also fixes a weakness of RRF. Fusion knows a document was ranked fourth, but not whether it was a close fourth or a distant one. The reranker reads the actual text.
When each part fails
The last point is the one people forget: a reranker only reorders what retrieval found. If the right document isn't in either top 50, nothing downstream will bring it back. When search quality is bad, check recall of the candidate set first.
Chunking causes a quieter version of the same problem. If you split long documents into chunks for embedding, the title and the product code may end up in a different chunk from the paragraph that matches. I prepend the title and any identifiers to every chunk before embedding and indexing it.
Also, don't panic when the keyword branch comes back empty. For a natural-language question, @@ often matches nothing at all. The vector branch carries the query on its own, and RRF works fine with one list.
How I'd roll it out
- Build a small evaluation set first. Fifty real queries with the document IDs that should come back is enough to start. Without it you're tuning by feel.
- Ship keyword plus vector with RRF. It's one SQL query and no new infrastructure if you're already on Postgres.
- Measure recall@50 of the fused list and precision@10 of the final list on your eval set.
- Add a reranker only if precision@10 is the problem and recall@50 is already good.
- Log queries with zero clicks. They are your next eval cases.
Takeaway
- Keyword search handles exact tokens and vector search handles meaning. Each covers the other's gaps.
- Combine them with reciprocal rank fusion. It uses ranks, not scores, so you never normalise two incompatible scales.
- Put a cross-encoder reranker on the short list when precision matters. It can't fix missing recall.
- If you already run Postgres,
tsvectorand pgvector cover all of this in one database and one query.