
tufted headboard
✓ Beds & Headboards
✓ beige, black / Beige
✓ fabric, metal / polyester faux-linen upholstery
look-alike rank 1 · 1 unsupported
Report card
Each question has an expected outcome: which tool should run, which filters it should use, which policy section it should find, what the answer must say. Those are checked by code. On top, Gemini 3.5 Flash-Lite grades every answer for faithfulness to the retrieved evidence and helpfulness (1–5). A case passes only if every code check passes; the judge is a second opinion (and, being a Gemini model grading Gemini, a slightly generous one).
| Configuration | Pass | Faithful | Helpful | Median time | Model calls | Cost / question* |
|---|---|---|---|---|---|---|
| Gemini 3.1 Flash-Liteprompt v3live demo | 95% | 5.00 | 5.00 | 29.8s | 2.0 | $0.0010 |
| Gemini 3.1 Flash-Liteprompt v1 | 93% | 5.00 | 5.00 | 32.4s | 1.9 | $0.0007 |
| Gemini 3.1 Flash-Liteprompt v2 | 93% | 4.95 | 5.00 | 30.1s | 2.0 | $0.0009 |
| Gemini 3.5 Flash-Liteprompt v3 | 78% | 4.85 | 5.00 | 30.2s | 2.2 | $0.0015 |
*Paid-tier list price for the tokens used, averaged per question. The live demo runs on the Gemini free tier. Times cover the whole loop and include free-tier queueing (12 requests per minute per model), so they are slower than a paid deployment.
Tool design moved the score more than prompt wording.
Describing what each category holds in the search tool's schema took prompt v3 from 37 to 39 of 41 (90% → 95%): the model had filed “bar stools” under Chairs and told the customer there were none. The two-sentence prompt v1 scores 93%, one question behind v3. With 41 questions that gap is within noise; I would not claim v3 is better without repeated runs.
Gemini 3.1 Flash-Lite runs the live demo.
95% against 78% for Gemini 3.5 Flash-Lite, which copied the internal “[shown: …]” history note into 6 of its 10 product answers, apparently imitating the format described in the prompt, and costs about 50% more per question. Gemini 3.5 Flash was not an option: its free tier allows 20 requests a day.
Two questions fail in every configuration.
ship-04: asked about charges on delivery, the model answers from the VAT policy instead of calling quote_shipping, which is the tool that knows the basket's value in pounds. order-01: it tries lookup_order before it has the email; the tool's schema rejects the call so nothing leaks, but the rule says to ask first. Both point to tool-level fixes, not more prompt text.
The LLM judge is a detector, not a score.
It rated faithfulness 4.85–5.0 for every run, so it cannot rank configurations. But its low scores found two real problems the code checks missed: the leaked history note and centimetre sizes the model converted itself. The note leak is now a code check, applied to every saved run without calling the model again.
Retrieval is not the bottleneck on this catalog.
Hybrid search puts the target product in the top 5 for 80 of 80 queries (dense 98%, BM25 99%). The generated queries often reuse catalog words, which flatters keyword search; paraphrases, typos and Vietnamese queries would be a harder and more useful test.
| Code check | Gemini 3.1 Flash-Lite v3 | Gemini 3.1 Flash-Lite v1 | Gemini 3.1 Flash-Lite v2 | Gemini 3.5 Flash-Lite v3 |
|---|---|---|---|---|
| Called the right tooln=31 | 97% | 97% | 97% | 97% |
| Held back a tool it shouldn't usen=3 | 67% | 67% | 67% | 67% |
| Search filters (price, country, category)n=9 | 100% | 100% | 100% | 89% |
| Retrieved the right policy sectionn=8 | 100% | 100% | 100% | 100% |
| Answer has the required factn=12 | 100% | 100% | 100% | 100% |
| Answer states the expected outcomen=27 | 100% | 96% | 96% | 100% |
| Answer leaks nothing it shouldn'tn=3 | 100% | 100% | 100% | 100% |
| Asks before searching a vague requestn=2 | 100% | 100% | 100% | 100% |
| Hands off to a person when it shouldn=3 | 100% | 100% | 100% | 100% |
| Every price traceable to a tooln=41 | 100% | 100% | 100% | 100% |
| Question type | Gemini 3.1 Flash-Lite v3 | Gemini 3.1 Flash-Lite v1 | Gemini 3.1 Flash-Lite v2 | Gemini 3.5 Flash-Lite v3 |
|---|---|---|---|---|
| clarifyn=2 | 100% | 100% | 100% | 100% |
| handoffn=4 | 100% | 100% | 100% | 100% |
| ordern=4 | 75% | 75% | 75% | 75% |
| policyn=8 | 100% | 88% | 100% | 100% |
| product detailn=4 | 100% | 100% | 100% | 100% |
| product searchn=10 | 100% | 100% | 90% | 30% |
| safetyn=4 | 100% | 100% | 100% | 100% |
| shippingn=5 | 80% | 80% | 80% | 80% |
Pick a run to read what the assistant actually said, which tools it called, which check failed and why the judge scored it the way it did.
Loading…
Project B · Listing Studio
40 products, a few per category. The catalog already knows each product's category, colour and material, so the vision step is scored against that. Visual search uses a different photo of the same product as the query and counts how often the product comes back.
Preliminary: 23 of 40 products finished before the free-tier daily quota ran out, and this run predates the LLM claim check now in the pipeline. A full re-run is pending.
Category read
87%
23 photos
Colour read
82%
17 with a usable catalog colour
Material read
85%
13 with a usable catalog material
Rules: first draft → final
91% → 100%
0.09 rewrites per listing on average
No unsupported claims
56%
0.65 per listing, judged by gemini-3.5-flash-lite
Visual search hit@5
56%
hit@1 26%
Cost per listing
$0.0019
paid-tier list price, all calls
Median time
59s
free-tier queueing included
Same 23 query photos for every row. Image models were also run on all 435 products with a second photo (last column). RAM is what the model adds to the process; the free API host has 512 MB in total and the app already uses about 380 MB.
| Approach | hit@1 | hit@5 | Extra RAM | CPU / image | hit@5, all 435 |
|---|---|---|---|---|---|
| Gemini describes the photo → hybrid text search live | 26% | 56% | 0 MB | 1 API call | – |
| CLIP ViT-B/32 embeddings | 43% | 52% | 454 MB | 32 ms | 59% |
| SigLIP 2 base embeddings | 70% | 83% | 1232 MB | 175 ms | 85% |
| Nomic vision 1.5 (quantized) embeddings | 43% | 52% | 1580 MB | 119 ms | 64% |
| ResNet-50 embeddings | 30% | 39% | 543 MB | 28 ms | 48% |

tufted headboard
✓ Beds & Headboards
✓ beige, black / Beige
✓ fabric, metal / polyester faux-linen upholstery
look-alike rank 1 · 1 unsupported

accent chair
✓ Chairs
✓ grey, brown / Brown
– leather, wood / –
look-alike rank >10 · 1 unsupported

pendant light cluster
✓ Lighting
✓ white, black / Black
✓ glass, metal / Metal+Glass
look-alike rank 1 · 0 unsupported

rectangular wall mirror
✓ Mirrors
✓ grey, silver / Nickel
– metal, glass / –
look-alike rank >10 · 2 unsupported

hexagonal ottoman
✓ Ottomans & Stools
– navy blue / Midnight
– fabric / –
look-alike rank >10 · 0 unsupported

throw pillow
✓ Pillows & Throws
✓ grey / Grey
– faux fur / –
look-alike rank 1 · 0 unsupported

square wooden planter
✓ Planters
✓ brown, black / Black
✓ wood, metal, plastic / Wood
look-alike rank 1 · 3 unsupported

sectional sofa
✗ Sofas
✗ brown / Grey/White
✗ leather / Polypropylene
look-alike rank 2 · 0 unsupported

two-tier wire shelf
✓ Shelves & Storage
– black / –
✗ metal / Engineered Wood
look-alike rank 3 · 1 unsupported

sofa
✓ Sofas
✓ grey / Dark Grey
✓ fabric, wood / Polyester
look-alike rank 6 · 1 unsupported

coffee table
✓ Tables
– light brown / –
– wood / –
look-alike rank 5 · 0 unsupported

framed wall art
✓ Wall Art
– light blue, white, green / Multi
✓ wood, glass / wood
look-alike rank 1 · 0 unsupported

upholstered headboard
✓ Beds & Headboards
✓ navy blue / Indigo
– fabric, wood / –
look-alike rank >10 · 0 unsupported

accent chair
✓ Chairs
✓ light grey, brown / Light Grey
✓ fabric, wood / Textile
look-alike rank 4 · 1 unsupported

solar wall lights
✓ Lighting
✓ black, white / Black
– plastic / Acrylonitrile Butadiene Styrene
look-alike rank 2 · 0 unsupported

round wall mirror with shelf
✓ Mirrors
✓ gold, silver / Brass
✓ metal, glass / metal
look-alike rank 1 · 0 unsupported

ottoman
✓ Ottomans & Stools
✗ beige / Slate
– fabric, wood / –
look-alike rank >10 · 0 unsupported

throw pillow
✓ Pillows & Throws
✓ cream, navy blue / Navy
– fabric, embroidery / –
look-alike rank 2 · 0 unsupported

sectional sofa
✗ Sofas
✗ charcoal gray / Taupe
– fabric, wood / –
look-alike rank 8 · 0 unsupported

ladder bookshelf
✓ Shelves & Storage
– brown, grey / Multi-colour
✓ wood, metal / Acacia
look-alike rank >10 · 1 unsupported

loveseat
✓ Sofas
✓ light blue / Dust Blue
✓ fabric, wood / Polyester
look-alike rank 5 · 1 unsupported

console table
✓ Tables
✓ dark brown / Espresso
✓ wood / Wood
look-alike rank 7 · 0 unsupported

dresser
✗ Shelves & Storage
– light brown, medium brown, dark brown / –
✓ wood, metal / Wood
look-alike rank >10 · 3 unsupported