Demo store · est. 2026Kestrel Home

Under the counter

A small agent loop, with every exit written down.

FastAPI + LangGraph backend, Qdrant for retrieval, Gemini as the model, Next.js for this page. The model decides what to look up; code decides anything that has to be exactly right.

1 · The loop

POST /api/chat

Validates input (≤ 2,000 chars, ≤ 20 messages), rate-limits per IP, then streams Server-Sent Events. History lives in the browser; assistant turns carry the ids of products they showed, so “how big is it?” still resolves without server sessions.
↓history trimmed to a 6,000-token budget

agent node

System prompt + history → model with 6 tools bound. If the reply has no tool calls, it is the answer and the loop ends.
↓tool calls · step < 6

tools node

Runs the calls in parallel, 20 s timeout each. A bad argument or a crash becomes an error message the model reads on the next step, so it can correct itself instead of the request failing.
↓back to agent node

give_up node

Reached when the model still wants tools at step 6. Answers the pending calls as “not run”, returns a fixed apology with support contacts, and makes no further model call.

2 · Retrieval

Data. 453 furniture products from Amazon Berkeley Objects (real names, materials, sizes, photos), deduplicated across colour variants. Prices, stock and UK eligibility are generated deterministically from the product id. Vendor lines like “free returns” are dropped so the five policy files are the only source of policy.

Chunking. One document per product (short, chunking would split facts that belong together). Policies split by section heading, each chunk prefixed with “Document > Section” so it stands on its own: 30 chunks.

Two vectors per point. Dense: bge-small-en-v1.5 (384 dimensions, runs locally on CPU via ONNX). Sparse: BM25, with IDF computed by Qdrant. Dense catches “cozy reading chair”; BM25 catches “jute” and “queen”. Both candidate lists are merged with Reciprocal Rank Fusion.

Retrieval eval · 80 shopper queries, no model in the loop

Setuphit@1hit@5MRRms
dense88%98%0.9131
sparse89%99%0.9426.9
hybrid91%100%0.9530.4
dense + category filter88%98%0.9215
sparse + category filter89%99%0.9432.4
hybrid + category filter91%100%0.9542.1

Each query was written by an LLM from one product’s description, without brand or name words. Hit@5 = that product is in the top 5. Other products can also fit a query, so these are lower bounds.

3 · Tools: what the model may not do itself

search_productsHybrid search over 453 products with price, category, country and stock filters.Filters run inside Qdrant, so a $900 sofa never reaches the model when the budget is $500.
get_product_detailsSize in inches and cm, weight, materials, stock, shipping countries.Unit conversion in code, not in the model's head.
search_store_policiesTop 3 policy sections (shipping, VAT, returns, warranty, care, payments).The answer cites the section it came from.
quote_shippingExact shipping cost, delivery time and UK VAT note for a basket.Pricing rules are business logic. The model never adds numbers up.
lookup_orderOrder status, only when order id and email both match.Same reply for wrong email and unknown order, so ids can't be probed.
handoff_to_humanCreates a ticket for complaints, damage, warranty, trade orders.Some decisions should not be automated.

4 · Eval

41 questions in 8 categories, each with expected tools, filters, policy section and answer content. Code checks decide pass or fail, including a rule that every $ or £ amount in an answer must appear in a tool result. An LLM judge adds faithfulness and helpfulness scores. Results are saved with their evidence, so a grader fix is re-scored offline without calling the model again.

5 · Prompt versions

v1 is two sentences. v2 adds grounding, search, order and safety rules. v3 adds one rule per failure the first eval run found: check stock before saying “you can order it”, hand off immediately when asked for a person, explain shipping limits instead of implying “out of stock”.

6 · Not built (yet)

Real auth and a database for orders, a Qdrant server instead of embedded mode, tracing (LangSmith or OpenTelemetry), answer caching for repeated questions, and running the eval in CI on every prompt change.

Listing Studio

Photo → facts → copy → check → repair

1 · Read the photo

One vision call with a Pydantic schema: category (same 12 as the store), type, colours, materials, style, visible features, photo problems. The image is re-encoded to a 768 px JPEG first, which also rejects files that are not images.

2 · Write from facts

A second call writes the listing and three ads using only those facts and the seller's notes. A photo cannot show assembly, durability, comfort or origin, so the prompt keeps those out.

3 · Check, repair

Code checks platform limits, risky ad claims and numbers that came from nowhere. Violations go back to the model, at most twice. A draft that still fails is marked for a person.

4 · Look-alikes

The photo's description becomes a query for the store's hybrid search. No image model on the server: the free host has 512 MB of RAM and the app already uses most of it. Image embeddings were measured offline instead; see the eval report.