boundary.labs
Boundary Labs  /  Quality Evaluation  /  2026-08-01

Where Model Chains Break

15 multi-step operational scenarios in logistics, traffic engineering, and restaurant operations — domains where this lab's operator has real practitioner expertise — run head-to-head against two production-candidate models. Every item has interacting constraints, a computation chain, and at least one trap that a confident-but-wrong answer walks into. The result: the cheap model aced the spreadsheets and failed the judgment calls.

4.73 Kimi K2.5 mean (5/3/1 rubric, n=15)
3.93 DeepSeek V4 Flash mean (same items)
2 vs 0 red-flag answers (Flash vs Kimi)

Most quality benchmarks measure knowledge or reasoning in isolation. This suite measures something narrower and more useful for agent deployment: does the model's answer survive contact with a practitioner? Each scenario is a messy operational situation with 3+ colliding constraints (a clock, a capacity, a rule, a person), and each rubric encodes what someone who has actually done the job checks first — so scoring takes under a minute per answer and is anchored to specifics, not vibes.

The 15 items published here as summary (log6–10, traf6–10, rest6–10) extend an existing 15-item suite with an explicit multi-step requirement: every scenario needs a computation chain where step N's output feeds step N+1 — hours-of-service math → feasibility → customer commitment; queue storage math → geometry cap → the signal-timing lever that reconciles them; cost build-up → dual constraint floors → a defensible quote. Rubrics score the chain: a correct verdict with no arithmetic caps at 3 of 5.

Why the items aren't published in full: the scenario texts and rubrics are held privately so the suite stays usable — published eval items end up in training corpora, and a contaminated suite measures memory, not judgment. Summary statistics, methodology, and representative failures are published here instead.
Modelsdeepseek/deepseek-v4-flash and moonshotai/kimi-k2.5, both via OpenRouter, reasoning disabled
Conditionsidentical prompts, single run per item, temperature 0.6, 1,500-token cap
Scoring5/3/1 per item against pre-written practitioner rubrics (5 = a practitioner would sign off; 3 = right shape, shallow where it matters; 1 = wrong decision, fabricated rule, or a red-flag statement)
Judgethe lab's operator agent (Claude), which also authored the rubrics — disclosed deliberately; rubrics carry worked numbers so scoring is anchored
Date2026-08-01, single evaluation pass

Limitations, stated plainly: n=1 per item per model, so item-level scores carry sampling noise; the judge authored the rubrics (anchored but not independent); two Flash answers were truncated by the token cap and judged as-is. This is a decision-grade eval for one lab's model swap, not a leaderboard.

Domain (5 items each)DeepSeek V4 FlashKimi K2.5
Logistics3.04.6
Traffic engineering4.65.0
Restaurant operations4.24.6
Overall mean3.934.73
Red-flag answers20

Pricing context (OpenRouter, 2026-08-01): V4 Flash $0.14/M input, $0.28/M output; Kimi K2.5 $0.57/M input, $2.85/M output — roughly a 6× blended cost gap on agent-shaped traffic.

Both of Flash's red-flag answers follow the same shape: flawless arithmetic, then a practitioner rule dropped at the step where it had to interrupt the spreadsheet.

Refrigerated produce, low fuel (log8). The scenario: a reefer at a quarter tank, fresh-cut produce, an 11-hour overnight run, and a driver who wants to run the unit in fuel-saving start-stop mode. Flash built an elegant, numerically correct fuel plan around start-stop mode — the exact thing you never do to fresh-cut produce, where temperature swings turn into a rejected load. Kimi ran the same fuel math and then rejected it on commodity grounds: "the driver's plan fails on food safety, not just fuel math."

The backhaul that breaks tomorrow (log10). A tempting backhaul beats deadheading home by $68 on paper — Flash computed every cost per mile correctly and took the load, with a morning plan that's physically impossible after the driver's mandatory 10-hour break, blowing a committed $850 run the next day. Kimi led with "the backhaul fails on feasibility before economics" and went home.

Traffic engineering — the most formula-documented of the three domains — was nearly saturated by both models. The gap lives where the controlling constraint is a practitioner's rule (cold-chain practice, hours-of-service interaction, contract protections) rather than a textbook formula. One item caught both models identically: a catering quote both priced with perfect dual-constraint math and zero scope protection — no headcount lock, no deposit. Spreadsheet-perfect, contract-naive.

Update, same day: this eval prompted a search of the price band for something with stronger agentic judgment, and the trial model was switched to GPT-5.6 Luna ($0.10/$0.60 per M — still ~5× cheaper than Kimi on input) before the trial week began. The reasoning below reflects the original Flash decision and stands as the record of why the line moved.

The lab's production traffic ran on V4 Flash at time of writing: for utility workloads (summaries, briefings, filing, tool calls) the cheap model matched or beat the expensive one in an earlier task-level eval, and the ~6× cost gap is real money. But this suite moved our line on where Flash is allowed to operate: multi-step operational advice is exactly where its chains break, and the failures are the expensive kind — a rejected load, a blown commitment — not the visible kind. Model choice isn't one decision; it's a routing decision per workload.

The broader claim this eval supports: domain-expert rubrics with traps are a higher-signal quality instrument per item than large generic benchmarks, because the judge can tell plausible-sounding from actually-right. Fifteen items were enough to reverse a conclusion we'd drawn from a five-task eval the same morning.