Knowledge graphs vs vector search for product data
— AI, Search, Knowledge Graph — 7 min read
Ask a vector index for "a replacement lid for the 750 ml steel bottle" and you'll get lids. Probably good-looking ones. Whether any of them actually fit that bottle is a different question, and embeddings can't answer it.
Product catalogues are full of questions like this. Which parts fit which model. Which products use a material that holds a food-contact certification. Which items in a category come from suppliers in a given region. The answers depend on how things relate to each other, not on how similar two descriptions sound.
Similar is not the same as related
An embedding places each product at a point in space. Products with similar text end up close together. That's useful for "show me things like this" and for messy, vague queries.
But closeness is all you get. Two lids with near-identical descriptions sit next to each other whether or not they fit the same bottle. Compatibility, part-of, made-of, certified-by and supplied-by are facts that exist between records. A vector for one record doesn't hold them.
Vector search struggles with a few kinds of query:
- Multi-hop questions. "Lids that fit this bottle and are made of a certified material" is a chain of three facts.
- Exact constraints. "Fits", "is certified" and "is in stock" are true or false. A similarity score of 0.83 doesn't say which.
- Exclusions. "Not made of BPA-containing plastic" is hard to express as a nearness.
- Counts. "How many suppliers offer this component?" needs counting, not ranking.
An LLM on top of vector retrieval will answer these questions anyway, from whichever chunks came back. That's how you get confident, wrong compatibility claims.
What a graph gives you
A knowledge graph stores entities (products, materials, certifications, categories, suppliers) as nodes and the relationships between them as typed edges. Questions become traversals.
You don't need a graph database to start. Product data usually already lives in relational tables, and two tables cover a lot:
CREATE TABLE entities (
id text PRIMARY KEY, -- 'p:lid-a', 'm:pp', 'c:food'
kind text NOT NULL, -- product | material | certification | category | supplier
name text NOT NULL,
embedding vector(1536) -- optional, for the hybrid step below
);
CREATE TABLE edges (
src text NOT NULL REFERENCES entities(id),
rel text NOT NULL, -- fits | made_of | certified | in_category | child_of
dst text NOT NULL REFERENCES entities(id),
PRIMARY KEY (src, rel, dst)
);
CREATE INDEX edges_dst_idx ON edges (dst, rel);"Lids that fit the bottle and are made of a food-contact certified material" is three joins:
SELECT lid.id, lid.name
FROM edges f
JOIN entities lid ON lid.id = f.src
JOIN edges m ON m.src = lid.id AND m.rel = 'made_of'
JOIN edges c ON c.src = m.dst AND c.rel = 'certified' AND c.dst = 'c:food'
WHERE f.rel = 'fits' AND f.dst = 'p:bottle-750';Variable-depth questions, like walking up a category tree, use a recursive CTE:
WITH RECURSIVE up AS (
SELECT dst AS id, 1 AS depth
FROM edges WHERE src = 'p:lid-a' AND rel = 'in_category'
UNION ALL
SELECT e.dst, up.depth + 1
FROM edges e JOIN up ON e.src = up.id
WHERE e.rel = 'child_of' AND up.depth < 10
)
SELECT id, depth FROM up;The depth < 10 guard stops a bad edge from causing an infinite loop. If traversals get deep, or most of your queries become path queries, a dedicated graph database with a query language like Cypher starts to pay off. Many catalogues never reach that point.
Where GraphRAG fits
"GraphRAG" usually refers to the approach in Microsoft Research's From Local to Global: A Graph RAG Approach to Query-Focused Summarization (Edge et al., 2024). Roughly:
- An LLM reads unstructured documents and extracts entities and relationships.
- Those form a graph, which is clustered into communities.
- The LLM writes a summary for each community.
- Queries are answered from the relevant summaries and graph neighbourhoods, not only from raw text chunks.
It works well for broad questions over large piles of text, such as "what are the main themes across these reports?", which plain chunk retrieval answers badly.
For product data, be careful about step 1. Much of your graph already exists as structured fields: SKU tables, bills of materials, compatibility lists, certification records. Re-extracting those facts from product descriptions with an LLM makes them less reliable. Use LLM extraction for what's only in text, like spec sheets, supplier documents and reviews. Load what's already structured as structured data, and keep track of where every edge came from so you can trust or discard it.
The hybrid I'd build
In practice, users type vague text and expect answers grounded in exact relationships. So use each tool for what it's good at:
- Vector (or hybrid keyword + vector) search finds the entry points. "That steel bottle, the 750 one" resolves to
p:bottle-750. - The graph expands and filters. Follow
fits,made_ofandcertifiededges from those entry nodes, and apply hard constraints like stock, region and certification. - The LLM writes the answer from the returned facts, citing the entities, and says so when the graph has no edge, instead of guessing.
Both steps fit in one Postgres query when the embeddings live on the entities table:
WITH entry AS (
SELECT id FROM entities
WHERE kind = 'product'
ORDER BY embedding <=> $1
LIMIT 3
)
SELECT e.src AS accessory, e.rel, e.dst AS product
FROM entry
JOIN edges e ON e.dst = entry.id AND e.rel = 'fits';The kind = 'product' filter runs after the HNSW index scan, so if products are a small share of the table you can get fewer rows than the LIMIT. A partial index (CREATE INDEX ... USING hnsw (embedding vector_cosine_ops) WHERE kind = 'product') fixes that.
A useful side effect: when the graph returns nothing, you know the catalogue is missing data. "We have no compatibility data for this bottle" is a far better answer than a guess, and it gives the data team a concrete gap to fill.
When to choose which
Vector search on its own is fine for discovery and "more like this", where an approximate answer hurts nobody. Add a graph when a wrong answer costs something, as with compatibility, compliance and substitutions. GraphRAG-style extraction earns its keep when the relationships are buried in a lot of unstructured text.
If users ask in plain language about things with hard relationships, which describes most product search, I'd start with the hybrid.
Takeaway
- Embeddings measure how similar two things sound. They don't record how things relate.
- Product questions often depend on relationships: fits, made of, certified by, supplied by. Store those as explicit edges.
- Start with an
edgestable in Postgres. Move to a graph database when path queries dominate. - Don't use an LLM to re-extract facts you already have in structured form.
- Use vector search to find where to start, the graph to decide what's true, and the LLM to explain it.