Skip to content

All case studies

Case study · Restaurant commerce platform · October 2026

An AI data assistant for restaurant brands

I built this at a restaurant commerce platform. It answers plain-language questions about a brand's app orders straight from its Snowflake data, and every figure in an answer is filled in from a query result or checked against one.

Several AI models share the work. GPT-6 Luna reads each question while TypeSafe Jev rates it, so a standard question never reaches Claude Opus 5.5. No model sees a table, schema or brand name: the app owns the definitions and guards every query.

It is live on one brand's data, and a second brand is configured and checked. It was built in Claude Code, and each change was reviewed by rounds of independent agents before it shipped.

Standard question, median time
2.4 s
Down from 4.6 s at first release
Analysis question, median time
7.5 s
Down from 42 s at first release
Model cost of a standard question
0.02 cents
Down from about 0.8 cents
Model cost of an analysis question, with a warm cache
2.4 cents
Down from about 15 cents

Tools

  • OpenAI
  • TypeSafe
  • Claude
  • Gemini
  • Vercel AI Gateway
  • Snowflake
  • Postgres
  • Neon

How a question is answered

The cheapest model that can answer does, and the app guards every query
  1. 1. A question arrives

    The brand is set by the app, never by the model or the question. Its daily budget is checked first.

  2. At the same time

    2. GPT-6 Luna reads it

    Via Vercel AI Gateway, at low effort. It answers a standard question by itself. If it cannot be reached, Claude Sonnet 5 reads instead.

    TypeSafe Jev rates it

    Standard, analysis or outside the data, in about 0.33 s. It is sent the question only, never data.

  3. One of three routes

    3a. Standard question

    Luna names the metrics, period and breakdown. The app writes the SQL and lays out the answer. Median 2.4 s.

    3b. Why did it change?

    Luna, as planner, names the change. The app runs totals and six breakdowns together. Claude Opus 5.5 writes it up, once.

    3c. Open analysis

    Claude Opus 5.5 takes the question whole and writes its own queries.

  4. 4. The app guards every query

    SQL from a model must be one SELECT over logical tables: no real table names, no writes. To every query the app adds the brand filter, the date window and the store limits.

    • Snowflake

    The data warehouse. A read-only role on the few views it needs, with a monthly spending cap. The only source of figures in an answer.

    • Postgres
    • Neon

    The app's own database. Results are cached for 10 minutes. Every run is saved with its trace and cost.

  5. 5. The app checks the answer, then shows it

    Figures are placeholders filled from the results. A figure typed by hand is checked against them. A part that fails a check goes back to the model, up to 3 attempts. Each part shows as it passes.

A standard question takes route 3a and never reaches Claude Opus 5.5. A question about why a figure moved takes 3b, where Opus is called once to write. Only a question that needs its own queries takes 3c.

Results

A lot of the work went into making it faster and cheaper. Between the first release on 26 September 2026 and 2 October, a standard answer became about twice as fast and an analysis answer more than five times as fast, with every tested question still passing.

MeasureFirst releaseLatest measured
Tested questions passed44 of 4452 of 52 standard and out-of-scope. 6 of 6 analysis, in each of three runs
Standard question, median time4.6 s2.4 s
Analysis question, median time42 s7.5 s
First part of an analysis answer on screenAbout 31 s6.5 s on average
Model cost of a standard questionAbout 0.8 cents0.02 cents
Model cost of an analysis questionAbout 15 cents2.4 cents with a warm cache

Two reviewing models graded the analysis answers without knowing which model wrote them: Gemini 3.1 Pro gave 4.50 to 4.67 of 5, and GPT-6.1 Sol gave 4.00 to 4.17. Agent reviews before release confirmed 25 faults across two of the changes, and all were fixed.

Design decisions

Most decisions followed a measurement, and the result was measured again afterwards.

DecisionWhyWhat it changed
The app owns the definitions and guards the SQL. The model is a replaceable part.Separating brands should not rest on a prompt.No model sees a table, schema or brand name. 11 of 11 out-of-scope questions were declined.
A standard question is a structured request, not SQL written by the model.Writing SQL and a layout took about 1,500 output tokens and 27 s. A request takes about 440 tokens and 4 to 5 s.First release: 44 of 44 tested questions, with a standard median of 4.6 s.
Figures are placeholders filled from results, and a row is picked by its label.In 6 of 29 analysis answers checked by hand, the figure came from the wrong row and every check passed.Picking a row by position is refused once the model has seen the rows.
GPT-6 Luna reads and TypeSafe Jev routes, in place of Claude Sonnet 5.Luna read 54 of 54 standard questions at 3.0 s for 0.02 cents, against 4.3 s and 0.83 cents. Alone it kept 2 of 6 analysis questions, so a router was added.Standard median about 4.3 s to 2.4 s. Model cost about 0.8 to 0.02 cents.
For "why did this change", the app runs the breakdowns and the model only writes.A model wrote the same five queries each time: about 3,700 tokens and half a minute.Analysis median 27 s to 7.5 s. Model cost 9.2 to 2.4 cents.
The page no longer wakes the warehouse, and the warehouse moved to Gen1.On the live site, 4.4 of the 7.2 cents of warehouse cost per answer was the page waking it.Gen1 bills 1.0 credits an hour against 1.35, and ran no slower.
Wrong answers are fixed in code, not in the prompt.Six answers were wrong yet passed every check: the figures were real, but they answered a different question.On held-out sets of 64 questions: share questions 63 right, totals 48 to 62 right.
A planner answers what one request can. Opus is called only for what only it does.Most questions sent to Opus came down to one request: a cent or more on Opus, a fiftieth of a cent on Luna."Why were sales down yesterday?" took 13.9 to 15.5 s before and 11.5 to 12.3 s after. Reviewer grades did not fall.
A brand is a config file, and the live brand's output is frozen by a test.Adding a second brand must not change the live one.The live brand's prompts, tools and SQL are identical bar one deliberate line. 15 of 15 questions for the second brand were answered in 2.4 to 3.6 s.

What these figures do not show

They come from evaluation runs on a laptop against the live models and warehouse, and nothing was timed on the live site after 1 October 2026.

A question asked on its own costs more, about 5.3 cents for a standard one, because Snowflake bills a 60-second minimum each time the warehouse wakes.

More case studies

Case study · Restaurant commerce platform

Ordering from a restaurant inside ChatGPT

In three days, a planning document became a working ChatGPT app where a guest orders from a restaurant in plain language. In testing, an AI model built 20 of 20 orders correctly and never invented a menu item.

From a planning document to a working app
3 days
Test orders built correctly
20 of 20
Target: 95% or more
Invented items or options
0
Target: 0
Median tool response time
0.46 s
Target: 2 to 5 s

Set your team up properly.

A 20-minute call to see if it's a fit. If it isn't, I'll tell you what I'd do instead.

Book a 20-minute fit call

Know a team that needs this?

Book a 20-minute fit call