ReStock

Keeps a restaurant’s purchasing plan valid as sales, stock and suppliers change. The model chooses what to check, never a number.

Context
NUS-ISS Show Me Your Agents Hackathon, team of four
My part
The agent layer: Coordinator, specialists, tools, contracts, guardrails, live-model integration and evaluation. I also built the pipeline that produced the submission video, and submitted the entry
Stack
Python, FastAPI, Pydantic, PostgreSQL, Next.js, Claude via the organiser gateway, Playwright
Year
2026
ReStock purchase recommendation, version 2 pending approval: 2.000 kg of vegetables, additional purchase only, S$11 of new purchase cash

The problem

A restaurant’s purchasing plan is a bet on the next service: expected sales, the stock on the shelves, and what each approved supplier can deliver and when. It is correct only while those hold. A promotion lifts demand, a closing count comes in low, a supplier marks an ingredient unavailable, a delivery arrives short, and yesterday’s plan is quietly wrong.

The hard part is not producing one recommendation. It is knowing, each time something changes, whether the plan still stands and, if not, what the smallest correct change is. Some changes arrive after money is already committed, where planning again from zero would buy the same vegetables twice. Deciding what to look at next suits a language model. Working out how many kilograms to buy, or whether a plan can be approved, does not.

How it works

ReStock keeps a versioned purchasing plan for one synthetic restaurant: five dishes, eight ingredients, three approved suppliers and 24 frozen supplier offers. When a fact changes, the Backend records it, moves the state revision on and queues an assessment. A separate worker claims it and freezes its inputs (operational time, knowledge cutoff and the captured revision) before anything reasons about them.

  • Coordinator and specialistsThe Coordinator chooses which of Demand, Inventory and Procurement a change needs, and checks that every result belongs to its task and the captured revision. Each specialist has an exact tool allowlist and returns a typed decision: call one of its tools, or finish with a next step. None can call another, write to the database or approve.
  • Deterministic enginePython kernels for forecasts, recipe requirements, FEFO stock projections, supplier allocation, optimisation and independent validation. Every number in a plan comes from here.
  • Backend and managerFastAPI and PostgreSQL own the facts, recheck the revision before publishing, and store each plan as an immutable version pending approval, with an append-only audit trail. The manager approves one exact version, and approval never places an order or adds stock.

From the submission video

Architecture slide: Manager UI, FastAPI and PostgreSQL, a queued frozen assessment, the worker and Coordinator, the Demand, Inventory and Procurement specialists, shared deterministic tools, Backend validation and exact-version manager approval
The live worker calls the organiser’s model gateway. OpenClaw, planned early on as the runtime, ended up as an isolated smoke test, not part of the application.

Checking only what changed

The Coordinator doesn’t run every specialist on every event. Each trigger has one first stop, and follow-ups are bounded and de-duplicated within a round. Before any model is involved, deterministic materiality runs. If a supplier change doesn’t touch anything the plan relies on, the run ends KEEP_CURRENT_PLAN with no specialist and no tool call. A test asserts exactly that: the tool log is empty and the existing version stays pending.

  • Sales update, promotion created or changed

    First specialist Demand

  • Closing count or correction, waste, short, late or cancelled delivery

    First specialist Inventory

  • Supplier availability, price or status change

    First specialist Procurement

  • Scheduled full planning

    First specialist All three

  • Manager instruction

    First specialist Classified into allowed routes only

TriggerFirst specialist
Sales update, promotion created or changedDemand
Closing count or correction, waste, short, late or cancelled deliveryInventory
Supplier availability, price or status changeProcurement
Scheduled full planningAll three
Manager instructionClassified into allowed routes only

Materiality, before any model call

Sales change assessment table under threshold policy SALES_MATERIALITY_V1: chicken rice expected 55, observed 0, material change yes
The live sales run from the recording night. Chicken rice was expected at 55 portions and observed at 0, which the threshold policy marks material.

What the specialists chose

Coordinator routing: sales updated, three specialists, twelve tool calls, zero retries
The same run: three specialists, the twelve deterministic tools they called, no retries.

Three outcomes and named failures

An assessment can succeed without authorising anything. Execution status and business outcome are separate fields: KEEP_CURRENT_PLAN keeps the plan, REVISE_PLAN publishes a validated version for approval, and ESCALATE records why automation stopped. The contracts refuse contradictory combinations, so a revision without a candidate, or a search-limit detail on anything but an incomplete calculation, can’t even be constructed.

# agent_contracts.py: contradictory outcomes can't exist
if self.escalation_detail is not None and (
    self.escalation_reason
    is not EscalationReason.CALCULATION_INCOMPLETE
):
    raise ValueError(
        "SEARCH_LIMIT_REACHED is only valid with "
        "CALCULATION_INCOMPLETE"
    )
if self.outcome is AgentOutcome.REVISE_PLAN:
    if self.candidate_result_ref is None:
        raise ValueError(
            "REVISE_PLAN requires a candidate result reference"
        )

# procurement_specialist.py: a tool raises only what it knows
reason = reasons.get(
    error.error.code, EscalationReason.TOOL_FAILURE
)
allowed_tools = TRUSTED_DOMAIN_ERRORS.get(error.error.code)
if allowed_tools is not None and (
    request.tool not in allowed_tools
):
    reason = EscalationReason.TOOL_FAILURE
services/api/src/agent_contracts.py and procurement_specialist.py (abridged)
  • CALCULATION_INCOMPLETEThe engine couldn’t finish. With SEARCH_LIMIT_REACHED, the bounded search ran out before proving a result, so its best candidate stays diagnostic and is never actionable.
  • NO_FEASIBLE_SUPPLIERThe search finished and proved no approved supplier can cover the need. Only the feasibility, allocation and optimisation tools are trusted to say it.
  • TOOL_FAILUREA tool, the model gateway, or an identity or revision check failed. It says nothing about stock either way.
  • MISSING_REQUIRED_DATAUnknown is not zero. Missing sales intervals or policy stop the run instead of being filled in.

KEEP_CURRENT_PLAN

Assessment outcome KEEP_CURRENT_PLAN: no additional purchase recommended
Post-purchase check: the purchases already arranged covered the need.

REVISE_PLAN

Assessment outcome REVISE_PLAN: a new recommendation is ready to review
The live normal assessment from the recording night, published as an engine-calculated version pending approval.

ESCALATE · CALCULATION_INCOMPLETE

Assessment outcome ESCALATE, CALCULATION_INCOMPLETE: manager review needed, no new purchase authorized
Material sales after the approved order window. The engine returned UNSUPPORTED_ISSUE_OPENING and no late order was invented.

ESCALATE · TOOL_FAILURE

Assessment outcome ESCALATE, TOOL_FAILURE: a calculation or tool failed, not evidence that no purchase is needed
One of two live attempts that night where the model returned no valid decision. It failed closed.

Buying only the gap

The case that matters most in a real kitchen. A 10 kg vegetable order has delivered 6 kg, the other 4 kg is delayed to the next day, and the manager has arranged a 4 kg emergency delivery elsewhere. Those purchases are facts, not suggestions, and planning again from zero would buy them twice.

Chun Yang’s Backend records them as fixed commitments, and Aniq’s deterministic post-purchase engine calculates only what is still uncovered. With the 4 kg rescue due in time, it keeps the plan. When the supplier cuts the rescue to 2 kg, the short delivery queues a reassessment and the engine’s answer is 2.000 kg more, nothing else.

My part was the route through the worker and Coordinator, built so no model is called: the worker stores the engine’s result, and the Coordinator binds it to the run and records it. Both recorded reassessments show zero specialist calls.

10 kg needed − 6 kg usable − 2 kg rescue in time = 2 kg. New cash: 2 kg × S$2 + S$3 delivery + S$4 emergency fee = S$11, labelled as new-purchase cash, not total economic cost.

Version 2, pending approval

Purchase plan version 2, pending approval: 2.000 kg vegetables from Market Supply; existing purchases of 10 kg from Fresh Foods and 2 kg from Market Supply listed as already arranged; new purchase cash total S$11
The recorded recommendation, with the original and rescue purchases listed as already arranged and the S$11 itemised. The starting plan was set up with controlled reasoning; both reassessments after it ran through the real worker.

How it was tested

On the recorded commit the backend PostgreSQL suite passed 1,341 tests with Ruff and Pyright clean, and the frontend lint and production build passed. What matters more is what they pin down: tool allowlists are exact and forbidden tools never execute; injected text in an objective or a supplier note can’t add a route, a permission or a business value; malformed model output gets one repair and then fails closed; stale evidence and stale approvals are rejected; and audit history can’t be edited.

I also built a provider-free harness that runs the same scenarios through three adapters: a static plan, a rule-based router, and ReStock’s agent path with scripted reasoning. On seven development scenarios, the agent path did not win:

  • Static plan

    Outcome correctness 0.714Routing accuracy 0.143Failed runs 0 / 7

  • Rule-based router

    Outcome correctness 0.714Routing accuracy 0.714Failed runs 0 / 7

  • ReStock, scripted reasoning

    Outcome correctness 0.714Routing accuracy 0.571Failed runs 0 / 7

AdapterOutcome correctnessRouting accuracyFailed runs
Static plan0.7140.1430 / 7
Rule-based router0.7140.7140 / 7
ReStock, scripted reasoning0.7140.5710 / 7

It missed one replan and made one unnecessary one, and the rule router routed better. That is consistent with how ReStock is built: each trigger’s first stop is already a fixed table, and the model is kept for choosing investigations inside a specialist. Seven development scenarios with scripted reasoning, measured on 19 September; not a live-model score.

Guardrail regressions, as shown in the video

Slide from the submission video: nine guardrail regression tests passed on commit eba56bf, including prompt injection, exact tool allowlists and malformed gateway output failing closed
Selected tests run with pytest on eba56bf at 21:14 on 27 September; the slide is generated from that log. Tests, not a live attack or a success rate.

Who built what

Four of us, one layer each. This comes from commits, pull requests and each person’s handover documents, not surviving lines alone: large squashed commits and code reworked by others make raw counts misleading.

  • Me · agent and AII set up the repository and froze the agent-facing contracts on day two. Then the Coordinator and its runtime, the three specialists and their tools, the control plane that joins them to the Backend, the organiser-gateway adapter and its strict parser, the manager-evidence projection, and the provider-free evaluation harness. In the final week I integrated the live model gateway, moved materiality ahead of coordination, wired post-purchase contingency through the worker, and fixed how incomplete-calculation reasons reach the manager. Separately, I produced the submission video and submitted the entry.
  • Chun Yang · backendThe authoritative service everything else answers to. Auth and sessions; closing counts, sales batches, deliveries and receipts; the frozen run inputs every assessment is checked against; trigger coalescing, supersession and exact-version approval, including the stale-version rejection; the queued worker loop; fixed commitments and the Backend contracts for the first synthetic contingency case; manager reads for stored calculations; and much of the shared integration contract.
  • Aniq · ML and decision engineEvery number in a plan: seasonal and promotion forecasting, recipe requirements, FEFO projections and multi-day coverage, the sales materiality policy, procurement search and independent validation, contingency procurement and the post-purchase engine behind the 2.000 kg result, a seven-day simulator, and bounded economic search and waste accounting. 864 numerical tests at its publication checkpoint.
  • Ethan · frontendThe manager’s whole experience: the Next.js workspace, recommendation and assessment views that keep counts, estimates, commitments and proposed supply distinct, exact-version approval with stale-version handling, queued-assessment refresh, the first-login tour, and browser test suites at desktop, tablet and phone widths. Also the manager feature inventory the other layers were checked against, and the 27 September scope review that mapped our evidence to the judging rubric.

The submission video

The brief asked for a video, the repository, two PDFs and a live deployment. I produced the video and submitted the entry. The video is not an edited screen recording: it came out of a small pipeline I built so that it could be reproduced from a runbook, recorded against the real system, and honest about what each scene proves.

  • Scripted capturePlaywright drives the real app at 1080p against fresh databases, requesting live assessments on camera. Chrome’s screencast writes every frame with a timestamp, and a shot timeline records where each scene starts, ends and waits.
  • Captions and evidence labelsAn overlay injected into the page draws the caption, the section and an evidence label on every scene: live model run, deterministic worker on a controlled baseline, automated test, provider-free evaluation, local prototype or verified hosted run.
  • One timelineA Python script cuts the real waiting, orders the shots into the story and encodes the cut. The narration script, timestamps and subtitles are written from the same timeline, so they can’t drift from the picture.
  • Narration48 ElevenLabs clips, one per shot, placed at their start times and loudness-matched, with no time-stretching or pitch change.
  • Trimming the first cutThe first narrated cut ran 12:42. A second pass measured each shot’s last on-screen action from the frame timestamps and removed only static holds, bringing it to 10:12 without cutting anything that happens.
  • Honest editingThe guardrail slide is generated from a pytest log, not typed. The gateway returned no valid decision twice that night; those attempts failed closed and were cut from the story, and the runbook records it. The hosted segment is a live assessment on the team’s Lightsail deployment.
  • SubmissionI reran the full backend suite on the recorded commit, published the video, added the video, PDFs and deployment links to the repository, and submitted the entry.

Hosted segment, with its labels

Video frame from the hosted segment: caption reads the hosted run succeeded with REVISE_PLAN and waits for exact-version approval; labels read Verified hosted run and Synthetic restaurant data
The caption and evidence labels are burned into the recorded frames, so they stay with every scene.

Decisions

  • A decision has nowhere to put a number

    A specialist’s decision is typed: call one of its allowlisted tools with references to existing evidence, or finish with a next step. The schema has nowhere to put a price, a quantity or an MOQ, and a test confirms a decision carrying them is rejected. Every figure a manager sees was calculated by the engine and validated before the Backend published it.

  • Materiality is decided before the Coordinator runs

    Whether a sales swing or a supplier change matters is arithmetic against a versioned threshold policy, so it runs deterministically and is stored with the run. The Coordinator reads the result. A model never gets to call a change harmless, and unknown never counts as safe.

  • When the answer is fully determined, don’t call a model

    Post-purchase contingency has one correct answer given the commitments and the policy. Sending it through specialists would only add latency and another way to fail, so the worker publishes the engine’s result through the same completion contract and audit trail.

  • A tool can only report what it can know

    NO_FEASIBLE_SUPPLIER from the optimiser is a finding about the world; the same code from a context read is a bug. Each domain error is trusted only from the tools able to establish it, and anything else becomes TOOL_FAILURE, so a manager never mistakes a failure for a fact about stock.

  • Bounded, and strict about output

    Two rounds, six specialist calls and one retry; Procurement gets six tool calls. The gateway must return exactly one JSON object bound to the right run and task. Malformed output gets one repair, and the parser never picks an object out of ambiguous text. Against a live gateway that sometimes returned nothing usable, this is what kept failures boring.

Results

  • 2.000 kg

    The only purchase recommended after the rescue delivery fell from 4 kg to 2 kg. Nothing already bought was bought again

  • 409

    PLAN_VERSION_STALE, when approval was attempted on a version that newer facts had invalidated

  • 1,341

    Backend PostgreSQL tests passing on the recorded commit, with Ruff and Pyright clean

  • Normal assessment

    Result REVISE_PLAN · ENGINE · PENDING_APPROVALEvidence Live model gateway, local; again on the hosted deployment

  • Material sales after the order window

    Result ESCALATE · CALCULATION_INCOMPLETE, no new planEvidence Live model gateway, local

  • Rescue delivery cut from 4 kg to 2 kg

    Result +2.000 kg vegetables, S$11 new cashEvidence Deterministic engine through the real worker, local; controlled baseline

  • Approving the invalidated version

    Result HTTP 409 PLAN_VERSION_STALEEvidence Local API

  • Unrelated supplier change

    Result KEEP_CURRENT_PLAN, zero tool callsEvidence PostgreSQL test

ScenarioResultEvidence
Normal assessmentREVISE_PLAN · ENGINE · PENDING_APPROVALLive model gateway, local; again on the hosted deployment
Material sales after the order windowESCALATE · CALCULATION_INCOMPLETE, no new planLive model gateway, local
Rescue delivery cut from 4 kg to 2 kg+2.000 kg vegetables, S$11 new cashDeterministic engine through the real worker, local; controlled baseline
Approving the invalidated versionHTTP 409 PLAN_VERSION_STALELocal API
Unrelated supplier changeKEEP_CURRENT_PLAN, zero tool callsPostgreSQL test

All data is synthetic: one restaurant, one service day, a bounded cash slice. These results show the decision loop and its safeguards working. They are not savings, waste reduction or live-model reliability, none of which were measured.

Limits

  • Synthetic data only. There was no restaurant pilot, so no savings, waste, stockout or cost outcome was measured.
  • The connected normal flow is a one-day cash slice (CASH_SLICE_V1). Waste accounting and a bounded 21-day economic search exist as numerical modules but are not wired into a manager workflow.
  • The live model path is not dependable yet: on the recording night two of three normal attempts returned no valid decision and failed closed. Token use, cost and latency were never measured.
  • The hosted deployment verified login, manager reads and one supported live assessment, nothing more. Exact-version approval, the stale-version 409 and safe escalation were verified on local runs and in tests only. Recovery after a restart was not verified anywhere.
  • The contingency case starts from a plan set up with controlled reasoning; only the reassessments after it ran through the real worker.
  • The routing comparison is seven development scenarios with scripted reasoning. It says nothing about how the live model routes.
  • REQUEST_HUMAN_APPROVAL is typed, but no persisted human-review workflow sits behind it. There is no POS integration or supplier checkout: approval never places an order.