Skip to content
Awsaf Alam
GitHubLinkedIn

Knowledge graphs vs vector search for product data

— AI, Search, Knowledge Graph — 7 min read

Ask a vector index for "a replacement lid for the 750 ml steel bottle" and you'll get lids. Probably good-looking ones. Whether any of them actually fit that bottle is a different question, and embeddings can't answer it.

Product catalogues are full of questions like this. Which parts fit which model. Which products use a material that holds a food-contact certification. Which items in a category come from suppliers in a given region. The answers depend on how things relate to each other, not on how similar two descriptions sound.

Similar is not the same as related

An embedding places each product at a point in space. Products with similar text end up close together. That's useful for "show me things like this" and for messy, vague queries.

But closeness is all you get. Two lids with near-identical descriptions sit next to each other whether or not they fit the same bottle. Compatibility, part-of, made-of, certified-by and supplied-by are facts that exist between records. A vector for one record doesn't hold them.

Left side, vector view: the 750ml bottle sits near three lids (sport lid, straw lid, wide lid) because their descriptions are similar. Right side, graph view: only the sport lid and straw lid have a fits edge to the bottle. The wide lid is close in vector space but doesn't fit.

Vector search struggles with a few kinds of query:

  • Multi-hop questions. "Lids that fit this bottle and are made of a certified material" is a chain of three facts.
  • Exact constraints. "Fits", "is certified" and "is in stock" are true or false. A similarity score of 0.83 doesn't say which.
  • Exclusions. "Not made of BPA-containing plastic" is hard to express as a nearness.
  • Counts. "How many suppliers offer this component?" needs counting, not ranking.

An LLM on top of vector retrieval will answer these questions anyway, from whichever chunks came back. That's how you get confident, wrong compatibility claims.

What a graph gives you

A knowledge graph stores entities (products, materials, certifications, categories, suppliers) as nodes and the relationships between them as typed edges. Questions become traversals.

You don't need a graph database to start. Product data usually already lives in relational tables, and two tables cover a lot:

sql
CREATE TABLE entities (
  id        text PRIMARY KEY,   -- 'p:lid-a', 'm:pp', 'c:food'
  kind      text NOT NULL,      -- product | material | certification | category | supplier
  name      text NOT NULL,
  embedding vector(1536)        -- optional, for the hybrid step below
);
 
CREATE TABLE edges (
  src text NOT NULL REFERENCES entities(id),
  rel text NOT NULL,            -- fits | made_of | certified | in_category | child_of
  dst text NOT NULL REFERENCES entities(id),
  PRIMARY KEY (src, rel, dst)
);
 
CREATE INDEX edges_dst_idx ON edges (dst, rel);

"Lids that fit the bottle and are made of a food-contact certified material" is three joins:

sql
SELECT lid.id, lid.name
FROM edges f
JOIN entities lid ON lid.id = f.src
JOIN edges m ON m.src = lid.id AND m.rel = 'made_of'
JOIN edges c ON c.src = m.dst  AND c.rel = 'certified' AND c.dst = 'c:food'
WHERE f.rel = 'fits' AND f.dst = 'p:bottle-750';

Variable-depth questions, like walking up a category tree, use a recursive CTE:

sql
WITH RECURSIVE up AS (
  SELECT dst AS id, 1 AS depth
  FROM edges WHERE src = 'p:lid-a' AND rel = 'in_category'
  UNION ALL
  SELECT e.dst, up.depth + 1
  FROM edges e JOIN up ON e.src = up.id
  WHERE e.rel = 'child_of' AND up.depth < 10
)
SELECT id, depth FROM up;

The depth < 10 guard stops a bad edge from causing an infinite loop. If traversals get deep, or most of your queries become path queries, a dedicated graph database with a query language like Cypher starts to pay off. Many catalogues never reach that point.

Where GraphRAG fits

"GraphRAG" usually refers to the approach in Microsoft Research's From Local to Global: A Graph RAG Approach to Query-Focused Summarization (Edge et al., 2024). Roughly:

  1. An LLM reads unstructured documents and extracts entities and relationships.
  2. Those form a graph, which is clustered into communities.
  3. The LLM writes a summary for each community.
  4. Queries are answered from the relevant summaries and graph neighbourhoods, not only from raw text chunks.
GraphRAG pipeline: documents go to LLM extraction of entities and relations, which build a graph. The graph is clustered into communities, each community gets an LLM-written summary, and queries are answered from the summaries plus graph context.

It works well for broad questions over large piles of text, such as "what are the main themes across these reports?", which plain chunk retrieval answers badly.

For product data, be careful about step 1. Much of your graph already exists as structured fields: SKU tables, bills of materials, compatibility lists, certification records. Re-extracting those facts from product descriptions with an LLM makes them less reliable. Use LLM extraction for what's only in text, like spec sheets, supplier documents and reviews. Load what's already structured as structured data, and keep track of where every edge came from so you can trust or discard it.

The hybrid I'd build

In practice, users type vague text and expect answers grounded in exact relationships. So use each tool for what it's good at:

  1. Vector (or hybrid keyword + vector) search finds the entry points. "That steel bottle, the 750 one" resolves to p:bottle-750.
  2. The graph expands and filters. Follow fits, made_of and certified edges from those entry nodes, and apply hard constraints like stock, region and certification.
  3. The LLM writes the answer from the returned facts, citing the entities, and says so when the graph has no edge, instead of guessing.
Hybrid flow: the user's question goes to vector search, which finds entry nodes. Graph traversal expands from those nodes along fits, made_of and certified edges and applies filters. The resulting facts go to the LLM, which writes an answer that cites the entities.

Both steps fit in one Postgres query when the embeddings live on the entities table:

sql
WITH entry AS (
  SELECT id FROM entities
  WHERE kind = 'product'
  ORDER BY embedding <=> $1
  LIMIT 3
)
SELECT e.src AS accessory, e.rel, e.dst AS product
FROM entry
JOIN edges e ON e.dst = entry.id AND e.rel = 'fits';

The kind = 'product' filter runs after the HNSW index scan, so if products are a small share of the table you can get fewer rows than the LIMIT. A partial index (CREATE INDEX ... USING hnsw (embedding vector_cosine_ops) WHERE kind = 'product') fixes that.

A useful side effect: when the graph returns nothing, you know the catalogue is missing data. "We have no compatibility data for this bottle" is a far better answer than a guess, and it gives the data team a concrete gap to fill.

When to choose which

Vector search on its own is fine for discovery and "more like this", where an approximate answer hurts nobody. Add a graph when a wrong answer costs something, as with compatibility, compliance and substitutions. GraphRAG-style extraction earns its keep when the relationships are buried in a lot of unstructured text.

If users ask in plain language about things with hard relationships, which describes most product search, I'd start with the hybrid.

Takeaway

  • Embeddings measure how similar two things sound. They don't record how things relate.
  • Product questions often depend on relationships: fits, made of, certified by, supplied by. Store those as explicit edges.
  • Start with an edges table in Postgres. Move to a graph database when path queries dominate.
  • Don't use an LLM to re-extract facts you already have in structured form.
  • Use vector search to find where to start, the graph to decide what's true, and the LLM to explain it.
© 2026 Awsaf Alam