Jev Labs

Failure Triage

Model failure or broken infrastructure?

Who decides what

  1. CodePull facts from the run log: tokens, was the prompt delivered, was usage missing, artifacts, HTTP codes, about 25 filtered lines
  2. Typed modelOne request, three typed questions: failure_class (Choice), is_infra and retry_helps (Nouls)
  3. CodePolicy: 0.90 confidence gate; facts beat judgments
  4. ResultAuto-route, or send to a human
codetyped modelhuman
Details · results, methods & full write-up
← All projects

Project 01

Failure Triage

why did this agent run fail?

6/6real infra failures diagnosedREAL JEV
openjev (Codiv free hosted) - NOT JEVREAL JEV (OpenCode Zen free, jev-1.13-free)

Failure labels and routing to human review

The AI Model Benchmark runs coding agents across many harnesses.

Who decides what

  1. CodePull facts from the run log: tokens, was the prompt delivered, was usage missing, artifacts, HTTP codes, about 25 filtered lines
  2. Typed modelOne request, three typed questions: failure_class (Choice), is_infra and retry_helps (Nouls)
  3. CodePolicy: 0.90 confidence gate; facts beat judgments
  4. ResultAuto-route, or send to a human
codetyped modelhuman

Headline result

REAL JEV (OpenCode Zen free, jev-1.13-free)

label accuracy: 6/6 (1.000); n=6; 95% interval [0.610, 1.000]; baseline always majority expected class: 0.500; zen_mr_infra_dnf.json

openjev (Codiv free hosted) - NOT JEV

held-out authored exact failure class: 57/60 (0.950); n=60; 95% interval [0.863, 0.983]; baseline always majority failure class: 0.100; challenge_v1_codiv.json

openjev vs real Jev

REAL JEV (OpenCode Zen free, jev-1.13-free)openjev (Codiv free hosted) - NOT JEV
Score: diagnosed correctly
real Jev6/6
openjev57/60
Requests
real Jev32
openjev114
Input tokens
real Jev45,958
openjev125,299
Median latency lower is better
real Jev431 ms
openjev456 ms
Cost at list price the openjev tokens would be $0.0053
real Jev$0
openjev$0

Requests, input tokens and median latency from the project's call log as recorded in RESULTS_SUMMARY (development runs included); both routes were free tiers. free tiers change, unverified

Details

The problem

The AI Model Benchmark runs coding agents across many harnesses. When a run ends without a deliverable, someone has to decide if it was the model (it gave up or delivered garbage) or the infrastructure (the provider cut the stream, credits ran out, tmux lost the session, the prompt was never delivered). Getting this wrong puts a model DNF on the public grid for a harness bug. failure-triage labels each run's ending from its log and routes confident labels automatically. Unsure ones go to a human.

How Jev is used

One request per run, three typed questions on the same state (code-extracted facts plus about 25 filtered log lines):

idtypeasks
failure_classChoice (10 in rubric v2)ok, model_gave_up, model_wrong_deliverable, provider_credit_or_quota, rate_limited, auth_error, harness_crash, hung_no_progress, provider_stream_cut (new in v2), not_in_this_list. Each option carries {what, examples, not_for}
is_infraNoulis the harness, provider or network at fault, rather than the model?
retry_helpsNoulwould an unchanged re-run succeed now?

Why decomposing helps. Code states the facts (tokens, whether the model got the prompt, whether usage was missing, artifact count, HTTP codes). The model makes one closed-set judgment. Code then applies the policy:

  • the 0.90 confidence gate;
  • "facts beat judgments": a model_* label is sent to a human when the model never got the prompt, or when the endpoint reported no usage.

Only the class question can be wrong. Everything around it can be checked. Pattern sources: tax-doc-classifier (closed-set Choice plus an escape option), Canny (facts can overrule), TypeSafe's jaggedness page (filter the log first).

Live results: openjev (Codiv free hosted) - NOT JEV, failures included

Real data. The AI Model Benchmark's 6 real _infra_dnf runs (Kimi Code / Kimi K3 stream cut, OpenCode / DeepSeek prompt never delivered, Claude Code / GLM tmux name bug ×2, Grok Build 403 refusal, Claude Code / GLM "stream ended without receiving any events"):

  • The reviewer's own diagnosis is used only as the expected label and is never sent to the model.
  • Folder names carry the diagnosis, so the model sees only the neutral cell id.
configurationlabels rightauto-routed (gate 0.90)wrong auto-routes
rubric v1 (original 9 classes)4/63/61 (Kimi stream cut → model_gave_up, conf 0.96)
rubric v2 (+ provider_stream_cut)5/64/61 (Kimi again → model_gave_up, conf 0.94)
v2 + usage_record_missing fact, always present (false on most rows)4/63/60. Adding a false flag flipped both GLM tmux rows to ok (conf 0.64/0.73, human-routed)
v2 + fact sent only when true + fact-conflict guard (shipped), 3 runs6/6, 6/6, 6/63/60
same, nonce robustness test (5 runs, random irrelevant field)29/3015/300. The one flip was the Kimi row (conf 0.63–0.71), which is always human-routed
mock, same 6 runs5/60/60

Hand-written fixtures (8 runs, rubric v2): 8/8 correct, 8/8 auto-routed. In 3 nonce runs every confidence was ≥0.996, with no flips.

Failures, stated plainly:

  • The Kimi case is hard. The endpoint reported 0 tokens, but the model had reasoned for 17 minutes. Without the harness's own context count as a fact, openjev confidently called it "the model gave up".
  • The fix was a code fact plus a guard.
  • Both GLM tmux rows are right but only at 0.70–0.74 confidence, so they go to a human.
  • On this data, openjev's confidence separates the easy cases (≥0.99) from the hard ones (0.63–0.74) well. The auto gate never fired on a wrong label once the guard was in.

Full details, raw outputs and the mock baseline: RESULTS.md and results/.

Real Jev run

Real Jev = jev-1.13-free on OpenCode Zen's limited-time free tier, run 2026-09-24 13:19–13:21 UTC with JEV_PROVIDER=zen. It was $0, and no response carried a cost field. Each result is labelled REAL JEV (OpenCode Zen free, jev-1.13-free). The free tier rate-limited us (HTTP 429 FreeUsageLimitError) after about 288 calls in about 2 minutes. Later calls were paced at ≤4 per minute.

openjev (NOT JEV)REAL JEV
real infra DNFs, rubric v14/64/6 (Kimi → model_gave_up 0.94, sent to a human by the fact guard; PS5 stream cut → not_in_this_list 0.24)
real infra DNFs, rubric v2 + guard (shipped)6/6 ×3 runs, 3/6 auto, 0 wrong auto6/6, 3/6 auto, 0 wrong auto
nonce test29/30 over 5 runs12/12 over 2 runs
confidence on the hard rows (Kimi, GLM tmux ×2)0.63–0.740.39–0.78, lower on the GLM rows
hand-written fixtures8/8, all auto8/8, all auto (min conf 0.98)
tokens per run / median latency~1,150 / 360–640 ms~1,480 / 410–455 ms

Real Jev lands on the same labels, and is less sure on the genuinely ambiguous tmux rows (as low as 0.39). That is what a calibrated model should do, and those rows go to a human either way. The v2 class and the fact guard were needed on both models.

Raw outputs: results/zen_*.

Cost, tokens, latency
  • Real DNF set: about 1,150 input tokens per run (6,925 for 6).
  • Hand-written fixtures: about 1,030 per run.
  • Latency: median 360–640 ms per request across runs.
  • Cost: $0 on Codiv. At TypeSafe's list price ($0.042 per 1M input tokens) that is about $0.00005 per run, or roughly $0.05 per 1,000 runs.
  • Everything in this upgrade: 106 requests, about 118k tokens.
What real Jev would change

(Written before the real-Jev run. The section above shows what it did.)

  • Confidence may behave differently. Real Jev is trained for calibration (RLCD). The 0.90 gate needs re-measuring on it before anything is auto-labelled. Run the nonce test again first (see the Latent Space notes).
  • Pin the version. Pin jev-1.13.x in the thresholds. The talk says there is no LTS, so re-tune when the version string changes.
  • Keep the guards. The v2 class and the usage-missing fact guard are model-independent and should stay whatever the model.
  • Next data. Ask the reviewer for a larger labelled export (about 50 finished runs, not only DNFs) to measure false infra labels on successful runs.
Advice folded in
  • Latent Space (Diogo Almeida): "verify everything" is one of the use cases he names. Confidence decides escalation, and the nonce test is his robustness check. We ran it: 29/30 stable, and the flips only happen below the gate.
  • building-with-jev (dbreunig): its advice is "keep the answer space stable once code depends on it", so v1 and v2 are versioned and reported separately. It also says "a Noul has no confidence; distance from 0.5 plays that role", which is why is_infra and retry_helps are reported as raw probabilities.
  • jev-browser-agent (smartdio): escalates below 0.6 to a bigger model. Here, below 0.90 goes to a human reviewer. A bigger-LLM tier in between is a cheap addition if the human queue grows. Sources: ../../research/latent-space-diogo-almeida.md, /dbreunig/building-with-jev-skill, /smartdio/jev-browser-agent, /jyje/pilot-typesafeai-jev.
Patterns it borrows
  • tax-doc-classifier: a big closed-set Choice whose criteria are {what, examples, not_for}, plus a not_in_this_list escape and a confidence gate (0.90 here; tax-doc uses 0.95).
  • Jaggedness page: filter first. Code keeps only signal lines and the tail of the log, and never asks Jev to read numbers. Token counts, HTTP codes, exit status and minutes idle are code facts.
  • Canny rule: facts beat judgments. If model_received_prompt is false, a model_* label is overruled and the run goes to a human.
Raw result files
REAL JEV (OpenCode Zen free, jev-1.13-free)results/zen_mr_infra_dnf.txt
mode=live [REAL JEV (OpenCode Zen free, jev-1.13-free)]  rubric=v2  runs=6  calls=6  median_latency_ms=418.9  input_tokens=8872  est_usd=0.000000
brown_pelican_detailed__kimi-code__kimi-k3__ provider_stream_cut        conf=0.67 infra=0.36 retry=0.35 -> human_review  ✓
macbook_pro_16__opencode__deepseek-flash__at harness_crash              conf=0.92 infra=0.72 retry=0.23 -> auto  ✓
pelican_bicycle__claude-code__glm-5.3__attem harness_crash              conf=0.61 infra=0.75 retry=0.24 -> human_review  ✓
pelican_bicycle__claude-code__glm-5.3__attem harness_crash              conf=0.39 infra=0.73 retry=0.21 -> human_review  ✓
pelican_bicycle_animated__grok-build__grok-4 provider_credit_or_quota   conf=0.99 infra=0.97 retry=0.06 -> auto  ✓
ps5_controller__claude-code__glm-5.3__attemp provider_stream_cut        conf=1.00 infra=0.83 retry=0.48 -> auto  ✓
agreement with expected labels: 6/6   auto-routed: 3/6
REAL JEV (OpenCode Zen free, jev-1.13-free)results/zen_fixtures.txt
mode=live [REAL JEV (OpenCode Zen free, jev-1.13-free)]  rubric=v2  runs=8  calls=8  median_latency_ms=444.5  input_tokens=10864  est_usd=0.000000
animated_pelican__grok-build__attempt1       provider_credit_or_quota   conf=1.00 infra=0.96 retry=0.06 -> auto  ✓
pelican_bicycle__claude-code__glm__attempt2  harness_crash              conf=1.00 infra=0.88 retry=0.30 -> auto  ✓
pelican_bicycle__claude-code__glm__attempt1  harness_crash              conf=1.00 infra=0.65 retry=0.32 -> auto  ✓
isometric_town__opencode__kimi               ok                         conf=1.00 infra=0.05 retry=0.65 -> auto  ✓
fjord__raw-api__model-x__low                 model_gave_up              conf=1.00 infra=0.08 retry=0.21 -> auto  ✓
w3_svg__codex__model-y                       hung_no_progress           conf=1.00 infra=0.34 retry=0.31 -> auto  ✓
w3_svg__opencode__model-z                    rate_limited               conf=1.00 infra=0.94 retry=0.41 -> auto  ✓
brown_pelican__claude-code__model-w          model_wrong_deliverable    conf=0.98 infra=0.14 retry=0.27 -> auto  ✓
agreement with expected labels: 8/8   auto-routed: 8/8
Run it
cd samples/failure-triage && ./run.sh
# = python3 build_mr_fixtures.py <ai_model_benchmark>/results/_infra_dnf fixtures/mr_infra_dnf.jsonl (only if that folder exists)
#   python3 triage.py fixtures/mr_infra_dnf.jsonl   (6 REAL AI Model Benchmark infra DNFs, scrubbed)
#   python3 triage.py fixtures/runs.jsonl           (8 hand-written fixtures)