Supplier conversations that can be inspected and improved
A sourcing request needs answers from suppliers before it can become a useful quotation. We rebuilt the agent that asks for them so every conversation can be traced, tested and handed to a person.
R&D Lead at Sourcy · 2026. Agent implementation and evaluation.
The workflow
At Sourcy, a buyer's sourcing request ends in a quotation. Before it can get there, someone has to ask suppliers about price, lead time and material, and get real answers back. An agent does that asking, one conversation per supplier. When those conversations stall or go wrong, the quotation waits.
Problem
The earlier agent was hard to inspect. When a conversation went wrong, it was hard to tell what the agent was trying to find out, what it already knew, and why it said what it said. Each fix was a guess, and a fix for one conversation could quietly break another.
Role
Agent implementation and evaluation. Teammates handled the messaging transport and infrastructure, the integration with the rest of the platform, and product input.
What we rebuilt
We rebuilt the agent around explicit goals, so each conversation knows which facts it still needs. It keeps the conversation history and writes each reply with the whole thread in view. When it reaches something it shouldn't decide alone, it escalates to a person instead of guessing.
Two more pieces made it possible to improve the agent on purpose. End-to-end tracing records every turn: what the agent knew, what it was after, and what it sent. Regression tests replay documented failure cases, so a fix for one conversation gets checked against the ones that already worked.
Press Play to watch four made-up conversations move turn by turn. Two finish, one goes quiet and gets a follow-up, and one hits a blocker and goes to a person.
Simulation
Simulated: a simulated supplier answers through the same channel interface a real one uses, so the turn loop, extraction and follow-ups run on the exact code path that talks to real suppliers.
Turn 0 of 5. 0 of 4 conversations finished.
Supplier A
Not started- Price: ?
- Lead time: ?
- Material: ?
Supplier B
Not started- Price: ?
- Lead time: ?
- Material: ?
Supplier C
Not started- Price: ?
- Lead time: ?
- Material: ?
Supplier D
Not started- Price: ?
- Lead time: ?
- Material: ?
What real conversations showed
Better replies kept suppliers talking, and suppliers then asked more complex questions. That changed the list of failure cases we had to handle before a broader rollout.
One of them is simple to describe. A supplier says no, and the agent keeps pushing for the missing detail. The change: once a supplier declines, the agent thanks them, stops, and marks the conversation as declined.
Status
Rebuilt and tested against real supplier conversations, with failure cases documented.
What to measure next
- How many requests finish without a person taking over
- Operator minutes spent per request
- Time from request to a useful quotation
- Cost per completed request
Need something like this built?
Pick a 30-minute slot and tell me what you're building.