ReRoute
When a late vessel breaks a transshipment plan, recover it without letting the model pick the container.
- Context
- PSA Code Sprint 2.0, team project
- My part
- Built all of it: backend, optimizer, agent runtime and console. The team made the slides
- Stack
- FastAPI, Pydantic v2, SQLModel, OR-Tools CP-SAT, OpenAI, React 19, Vite
- Year
- 2026

The problem
A mainline vessel arrives 195 minutes late. Its boxes were booked onto onward services with fixed cut-offs, so one delay becomes 24 separate deadline problems at once. The delay has already happened: the question is what to do with each box.
Thirteen containers would make their connection if expedited, and the yard has eight slots, split again by handling group, reefer plugs and dangerous-goods limits. Readiness is a spread, not a time: a box whose median clears the cut-off but whose p90 does not is a worse bet than one with a tight band.
Some levers belong to someone else. The terminal runs the yard, not the onward carrier’s schedule, so buying time means asking another company and waiting. And some calls must not be automated at all.
How it works
A console over a durable backend carries one disruption through seven chapters. Four actors hold separate authority, and the split lives in the code, not in a prompt.
- Agent (OpenAI model)Picks the next tool from the set currently permitted, sequences recovery, and waits when evidence or authority is missing.
- Optimizer (CP-SAT)Proves feasible allocations under capacity, handling-group, reefer and DG limits. The model never picks containers.
- Operator and carrierApprove outbound requests and counter-offers, and settle trade-offs no allocation wins outright. The carrier owns its schedule.
- Deterministic policyWorkflow guards, approval bindings, safety disposition and escalation.
Coordinate

Protect

The allocator
- Scenario worlds50 seeded worlds with correlated delay at the vessel, handling-group and container level, drawn as antithetic pairs.
- CoefficientsFor each box, how many worlds flip from missed to preserved if it takes a slot.
- CP-SAT solveMaximise expected preserved connections under the hard limits, with solver parameters pinned so runs reproduce bit for bit.
- Every optimumA second model, objective fixed at the proven optimum, collects every optimal allocation.
- Strict dominanceOne allocation is chosen only if it beats the others on every measure. Otherwise the trade-off goes to a person.
On the canonical fixture the median-based baseline and the scenario-aware allocator share only four of eight slots: the same eight slots, rebalanced toward the service with tighter bands, keep more connections.
Decisions
The agent only sees tools it is allowed to use right now
The tool registry is rebuilt from typed state on every turn and the runtime rejects anything outside it. While an unhandled yard assessment exists, all three carrier tools disappear. Holding a feeder, changing a carrier schedule, overriding a DG rule or setting yard capacity have no implementation at all, and a test asserts it.
The solver allocates, the model sequences
Scarce slots are a constraint problem with a provable optimum, so CP-SAT owns them. The model decides what happens next and when to wait, which is the part that needs judgement over changing evidence.
Approval is bound to the exact payload
An operator approval carries a fingerprint of the persisted request; any other payload is rejected with no state change. A carrier counter-offer needs a second, separate approval. Carrier silence writes nothing, because recording it as a no invents information the terminal does not have.
The safety model can only report
Its output schema is a result, an explanation and a verbatim excerpt. There is no field for a DG class or a disposition, so it cannot express a verdict. Deterministic policy turns the contradiction into an escalation, and any provider error, invalid output or stale evidence fails closed.
State lives in the database, not the context window
Agent runs, wait states, approvals and decisions are persisted; decisions are superseded, never edited; audit events are append-only with typed actors. Both model interfaces have offline stand-ins, so the whole demo and test suite run without an API key through the same code path.
Results
+4.12%
Expected preserved connections against a median-based greedy baseline, same eight slots
2,500
Synthetic evaluation worlds from 50 holdout seeds, frozen before evaluation
0
Capacity violations or unsafe allocations, for either allocator
Expected preserved connections
Median greedy 12.0136Scenario-aware CP-SAT 12.5088
Preserved across 2,500 worlds
Median greedy 30,034Scenario-aware CP-SAT 31,272
Expedite slots used
Median greedy 8 / 8Scenario-aware CP-SAT 8 / 8
| Metric | Median greedy | Scenario-aware CP-SAT |
|---|---|---|
| Expected preserved connections | 12.0136 | 12.5088 |
| Preserved across 2,500 worlds | 30,034 | 31,272 |
| Expedite slots used | 8 / 8 | 8 / 8 |
A synthetic benchmark: a reproducible claim about the allocator on this fixture under a stated distribution, not a measured gain in PSA throughput. Every vessel, container and timing is invented. In a live run the model chose the same tool sequence as the offline replay, 10 of 10 calls. The backend suite has 541 passing tests and the frontend 104; they assert the authority boundaries, not just the happy path.
Limits
- Synthetic feeds, not production terminal, yard or schedule integrations.
- A representative carrier workflow modelled on the public DCSA timing pattern, not a live carrier connection.
- The berth-time lever is an inference; real terminals may negotiate the cargo cut-off instead.
- SQLite fits one terminal’s exception queue; several terminals would want PostgreSQL.
- No authentication, rate limiting or multi-tenancy: a single-operator demonstration.

