Skip to content
Awsaf Alam
GitHubLinkedIn

Hybrid search: embeddings plus keyword, then rerank

— Search, Postgres, AI — 7 min read

Most teams that add "AI search" start the same way. Embed every document, embed the query, return the nearest neighbours. The demo looks great. Then someone searches for a SKU, an error code or a person's surname, and the top result is something vaguely related that doesn't contain the string they typed.

Keyword search has the opposite problem. It finds ERR_CONN_RESET instantly but returns nothing for "the connection keeps dropping".

You usually want both. In Postgres that means full-text search and pgvector side by side, merged with reciprocal rank fusion, with an optional reranker on the short list.

Hybrid search pipeline: the query goes to a keyword retriever (tsvector) and a vector retriever (pgvector) in parallel. Each returns its top 50. Reciprocal rank fusion merges them into one list, a reranker reorders the top 20 to 50, and the final top 10 go to the user or the LLM.

What each retriever is good at

Keyword search (BM25, or Postgres tsvector) scores documents by the query terms they contain. Rare terms count for more. It is precise, explainable and cheap. It knows nothing about meaning, so synonyms, paraphrases and typos fall through.

Vector search compares a query embedding to document embeddings. It handles paraphrase, rough descriptions and cross-language queries well. It is weak at exact tokens: product codes, version numbers, names, anything the embedding model never saw often enough to give a distinct vector. It also always returns something, even when nothing relevant exists.

Their failures don't overlap much, which is why combining them works.

The table

One table holds both representations. Postgres keeps the tsvector in sync through a generated column, so there is nothing extra to maintain on write.

sql
CREATE EXTENSION IF NOT EXISTS vector;
 
CREATE TABLE docs (
  id         bigint PRIMARY KEY,
  title      text NOT NULL,
  body       text NOT NULL,
  embedding  vector(1536) NOT NULL,
  search_tsv tsvector GENERATED ALWAYS AS (
    setweight(to_tsvector('english', title), 'A') ||
    setweight(to_tsvector('english', body),  'B')
  ) STORED
);
 
CREATE INDEX docs_tsv_idx ON docs USING gin (search_tsv);
CREATE INDEX docs_vec_idx ON docs USING hnsw (embedding vector_cosine_ops);

Use whatever dimension your embedding model produces. Title terms get weight A so a match in the title beats a match buried in the body.

A note on scoring: Postgres's built-in ts_rank_cd is not BM25. It doesn't use corpus-wide term statistics the same way. For many use cases it is fine. If you need real BM25 inside Postgres, extensions such as ParadeDB's pg_search provide it. The fusion step below only uses rank positions, so you can swap the keyword scorer later without changing anything else.

Merging two lists with reciprocal rank fusion

The two retrievers produce scores on different scales. A cosine distance of 0.21 and a ts_rank_cd of 0.4 can't be compared or added. You could normalise them, but the normalisation shifts with every query.

Reciprocal rank fusion (Cormack, Clarke and Büttcher, 2009) ignores the scores and uses only positions. Each document gets 1 / (k + rank) from every list it appears in, and the contributions are summed. k = 60 is the constant from the paper and works well as a default.

Two ranked lists merged by reciprocal rank fusion. Doc C is rank 1 in keyword and rank 2 in vector, so it scores 1/61 + 1/62 and comes first. Doc A appears only in the vector list at rank 1 and scores 1/61. Documents found by both retrievers rise to the top.

A document that shows up near the top of both lists beats one that tops only a single list. That is usually what you want.

Here is the whole thing as one query. $1 is the raw query text, $2 is the query embedding.

sql
WITH kw AS (
  SELECT id, row_number() OVER (ORDER BY rank DESC) AS r
  FROM (
    SELECT id, ts_rank_cd(search_tsv, q) AS rank
    FROM docs, websearch_to_tsquery('english', $1) AS q
    WHERE search_tsv @@ q
    ORDER BY rank DESC
    LIMIT 50
  ) t
),
vec AS (
  SELECT id, row_number() OVER (ORDER BY dist) AS r
  FROM (
    SELECT id, embedding <=> $2 AS dist
    FROM docs
    ORDER BY dist
    LIMIT 50
  ) t
)
SELECT id, sum(1.0 / (60 + r)) AS rrf_score
FROM (SELECT * FROM kw UNION ALL SELECT * FROM vec) AS both_lists
GROUP BY id
ORDER BY rrf_score DESC
LIMIT 20;

A few details matter:

  • Each inner query does its own ORDER BY ... LIMIT so the HNSW index can stop early, and ranking happens on at most 100 rows.
  • websearch_to_tsquery accepts user input safely. Quotes and -term work the way people expect, and stray punctuation doesn't throw a syntax error.
  • If you filter (tenant, language, status), apply the filter inside both branches. With HNSW, a selective filter can leave you with fewer than 50 vector results. pgvector's filtering notes cover iterative scans and partial indexes for this.

Adding a reranker

Fusion gets the right documents into the candidate list. A reranker puts them in the right order.

Embedding search uses a bi-encoder: query and document are embedded separately and compared with a dot product. That is fast, but the model never sees the two texts together. A cross-encoder reranker takes the query and one candidate as a single input and outputs a relevance score. It is far more accurate and far too slow to run over the whole corpus. Running it over 20 to 50 fused candidates is affordable.

ts
const candidates = await hybridSearch(query, queryEmbedding) // the SQL above
const scores = await reranker.score(
  query,
  candidates.map(c => c.body),
)
const top = candidates
  .map((c, i) => ({ ...c, score: scores[i] }))
  .sort((a, b) => b.score - a.score)
  .slice(0, 10)

reranker stands for whatever you use: a hosted rerank API or a self-hosted cross-encoder from the sentence-transformers family. Measure the latency. A reranker adds a model call to every search, and you may want to skip it for short, exact-looking queries.

The reranker also fixes a weakness of RRF. Fusion knows a document was ranked fourth, but not whether it was a close fourth or a distant one. The reranker reads the actual text.

When each part fails

Failure modes. Keyword search misses paraphrases and synonyms. Vector search misses exact codes, names and rare tokens, and always returns something. RRF gives rank credit to junk if one retriever's top results are junk. The reranker adds latency and cost, and it cannot recover a document that neither retriever found.

The last point is the one people forget: a reranker only reorders what retrieval found. If the right document isn't in either top 50, nothing downstream will bring it back. When search quality is bad, check recall of the candidate set first.

Chunking causes a quieter version of the same problem. If you split long documents into chunks for embedding, the title and the product code may end up in a different chunk from the paragraph that matches. I prepend the title and any identifiers to every chunk before embedding and indexing it.

Also, don't panic when the keyword branch comes back empty. For a natural-language question, @@ often matches nothing at all. The vector branch carries the query on its own, and RRF works fine with one list.

How I'd roll it out

  1. Build a small evaluation set first. Fifty real queries with the document IDs that should come back is enough to start. Without it you're tuning by feel.
  2. Ship keyword plus vector with RRF. It's one SQL query and no new infrastructure if you're already on Postgres.
  3. Measure recall@50 of the fused list and precision@10 of the final list on your eval set.
  4. Add a reranker only if precision@10 is the problem and recall@50 is already good.
  5. Log queries with zero clicks. They are your next eval cases.

Takeaway

  • Keyword search handles exact tokens and vector search handles meaning. Each covers the other's gaps.
  • Combine them with reciprocal rank fusion. It uses ranks, not scores, so you never normalise two incompatible scales.
  • Put a cross-encoder reranker on the short list when precision matters. It can't fix missing recall.
  • If you already run Postgres, tsvector and pgvector cover all of this in one database and one query.
© 2026 Awsaf Alam