Real data. The AI Model Benchmark's 6 real _infra_dnf runs (Kimi Code / Kimi K3 stream cut, OpenCode / DeepSeek prompt never delivered, Claude Code / GLM tmux name bug ×2, Grok Build 403 refusal, Claude Code / GLM "stream ended without receiving any events"):
- The reviewer's own diagnosis is used only as the expected label and is never sent to the model.
- Folder names carry the diagnosis, so the model sees only the neutral cell id.
| configuration | labels right | auto-routed (gate 0.90) | wrong auto-routes |
|---|
| configurationrubric v1 (original 9 classes) | labels right4/6 | auto-routed (gate 0.90)3/6 | wrong auto-routes1 (Kimi stream cut → model_gave_up, conf 0.96) |
configurationrubric v2 (+ provider_stream_cut) | labels right5/6 | auto-routed (gate 0.90)4/6 | wrong auto-routes1 (Kimi again → model_gave_up, conf 0.94) |
configurationv2 + usage_record_missing fact, always present (false on most rows) | labels right4/6 | auto-routed (gate 0.90)3/6 | wrong auto-routes0. Adding a false flag flipped both GLM tmux rows to ok (conf 0.64/0.73, human-routed) |
| configurationv2 + fact sent only when true + fact-conflict guard (shipped), 3 runs | labels right6/6, 6/6, 6/6 | auto-routed (gate 0.90)3/6 | wrong auto-routes0 |
| configurationsame, nonce robustness test (5 runs, random irrelevant field) | labels right29/30 | auto-routed (gate 0.90)15/30 | wrong auto-routes0. The one flip was the Kimi row (conf 0.63–0.71), which is always human-routed |
| configurationmock, same 6 runs | labels right5/6 | auto-routed (gate 0.90)0/6 | wrong auto-routes0 |
Hand-written fixtures (8 runs, rubric v2): 8/8 correct, 8/8 auto-routed. In 3 nonce runs every confidence was ≥0.996, with no flips.
Failures, stated plainly:
- The Kimi case is hard. The endpoint reported 0 tokens, but the model had reasoned for 17 minutes. Without the harness's own context count as a fact, openjev confidently called it "the model gave up".
- The fix was a code fact plus a guard.
- Both GLM tmux rows are right but only at 0.70–0.74 confidence, so they go to a human.
- On this data, openjev's confidence separates the easy cases (≥0.99) from the hard ones (0.63–0.74) well. The auto gate never fired on a wrong label once the guard was in.
Full details, raw outputs and the mock baseline: RESULTS.md and results/.