Jev Labs

Run Verdict

Check the claim against the evidence.

Who decides what

  1. CodeFact checks: deliverable exists, parses, rendered after the last write, run status. A failed fact is FAIL with no model call
  2. Typed modelFour atomic judgments: addresses_prompt and claim_support (Scores), self_reported_gaps and claims_done (Nouls)
  3. CodeWeights and bands (rubric v2), confidence gate, overclaim rule, “admits a gap → NEEDS_REVIEW”
  4. ResultA review band, or human review
codetyped modelhuman
Details · results, methods & full write-up
← All projects

Project 02

Run Verdict

a typed pre-screen for run review runs

5/6runs pre-screened correctlyREAL JEV
openjev (Codiv free hosted) - NOT JEVREAL JEV (OpenCode Zen free, jev-1.13-free)

Whether completion claims match recorded evidence

Reviewers score every finished run by hand. The most common trap is overclaiming: the final message says "verified, looks great" when nothing was rendered after the last write.

Who decides what

  1. CodeFact checks: deliverable exists, parses, rendered after the last write, run status. A failed fact is FAIL with no model call
  2. Typed modelFour atomic judgments: addresses_prompt and claim_support (Scores), self_reported_gaps and claims_done (Nouls)
  3. CodeWeights and bands (rubric v2), confidence gate, overclaim rule, “admits a gap → NEEDS_REVIEW”
  4. ResultA review band, or human review
codetyped modelhuman

Headline result

REAL JEV (OpenCode Zen free, jev-1.13-free)

5/6 band agreement. The confidence branch ran before the overclaim rule and sent the overclaiming run to human review.

openjev (Codiv free hosted) - NOT JEV

Historical openjev result withdrawn: the row-level artifact was not retained. No checkable comparison is published here.

Details

The problem

Reviewers score every finished run by hand. The most common trap is overclaiming: the final message says "verified, looks great" when nothing was rendered after the last write. run-verdict gives each run a band (FAIL / LIKELY_WEAK / NEEDS_REVIEW / PRE_PASS) and a 0–100 composite, so the queue can be ordered and overclaims flagged. It never replaces the rubric judge. Only code facts can FAIL a run.

How Jev is used

One request per run that passed the fact checks:

idtypeasks
addresses_promptScore (4 levels)how completely the final message describes delivering what the prompt asked
claim_supportScore (4 levels)how well the code-checked evidence supports the success claims
self_reported_gapsNouldoes the message admit something is missing or unverified?
claims_doneNouldoes the message claim the task is complete?

Why decomposing helps. The work is split three ways:

  • Code checks the facts: deliverable exists, parses, rendered after the last write, run status. If a fact fails, the run is FAIL with no model call.
  • The model gives four atomic judgments.
  • Code does the rest: weights and bands (killmyidea-style composite scoring, rubric v2), a confidence gate, and two named rules: overclaim, and (new in v2) "admits a gap → NEEDS_REVIEW".

A bad band can be traced to one question.

Real Jev run

REAL JEV (OpenCode Zen free, jev-1.13-free). Source: samples/run-verdict/results/zen_fixtures.json.

The overclaiming run received NEEDS_REVIEW with composite 28. The code policy checked addresses_prompt confidence 0.51 before the overclaim rule. Its claim_support score was 0.01 at confidence 0.99; the model clearly rejected claim support. This result reflects policy ordering, not uncertainty about whether the claim was supported.

RunExpected bandRecorded bandComposite
pelican__model-a__honestPRE_PASSPRE_PASS94
pelican__model-b__overclaimsLIKELY_WEAKNEEDS_REVIEW28
pelican__model-c__admits_gapNEEDS_REVIEWNEEDS_REVIEW65
pelican__model-d__missing_fileFAILFAIL—
pelican__model-e__dnfFAILFAIL—
pelican__model-f__broken_xmlFAILFAIL—
Cost, tokens, latency
  • Fixtures: 1,627 input tokens over 3 requests (about 540 per judged run). Median latency 280–350 ms.
  • DNF overclaim check: 2,195 tokens over 3 requests, median 445 ms.
  • Cost: $0. At list price that is about $0.00002 per judged run.
What real Jev would change

(Written before the real-Jev run. The section above shows what it did.)

  • Re-measure the gates. Real Jev's calibration may make the CONF_GATE (0.6) branch useful again. The historical openjev confidence claim was withdrawn because the row-level artifact was not retained.
  • Keep the v2 rule anyway. "Admits a gap" is a policy rule, so it stays whatever the model.
  • Real runs next. The next honest test is 50+ finished AI Model Benchmark runs with reviewer scores, comparing the band against the human score.
Advice folded in
  • Latent Space: "add the question, the threshold and the test case" (about 1:07) and "verify everything" are the two moves this sample makes. The failing run stays in fixtures/runs.jsonl as a regression case.
  • building-with-jev (dbreunig): the rule "the final decision is wrong while each answer is right → change the policy in code, leave the questions alone" is exactly the v1→v2 fix. Sources: ../../research/latent-space-diogo-almeida.md, /dbreunig/building-with-jev-skill, /smartdio/jev-browser-agent, /jyje/pilot-typesafeai-jev.
Patterns it borrows
  • Canny / our evidence-gate: facts to code, judgments to Jev, only facts decide. In production the facts come from experiments/evidence-gate.
  • killmyidea: atomic Score questions whose levels are written as concrete situations. Weights and bands live in code, and RUBRIC_VERSION is bumped on any change.
  • TypeSafe confidence routing: if any Score's confidence is under 0.6, the run goes to NEEDS_REVIEW whatever its composite.
Raw result files
REAL JEV (OpenCode Zen free, jev-1.13-free)results/zen_fixtures.txt
mode=live [REAL JEV (OpenCode Zen free, jev-1.13-free)]  rubric=v2  runs=6  jev_calls=3  input_tokens=2352  median_latency_ms=525.9  est_usd=0.000000
pelican__model-a__honest                 PRE_PASS      composite=  94  claims supported  ✓
pelican__model-b__overclaims             NEEDS_REVIEW  composite=  28  low confidence  ✗ expected LIKELY_WEAK
pelican__model-c__admits_gap             NEEDS_REVIEW  composite=  65  model admits a gap  ✓
pelican__model-d__missing_file           FAIL          composite=None  deliverable missing  ✓
pelican__model-e__dnf                    FAIL          composite=None  run status DNF  ✓
pelican__model-f__broken_xml             FAIL          composite=None  deliverable does not parse  ✓
agreement: 5/6
REAL JEV (OpenCode Zen free, jev-1.13-free)results/zen_mr_infra_dnf.txt
mode=live [REAL JEV (OpenCode Zen free, jev-1.13-free)]  rubric=v2  runs=6  jev_calls=3  input_tokens=2952  median_latency_ms=419.3  est_usd=0.000000
brown_pelican_detailed__kimi-code__kimi- FAIL          composite=None  infra, not a model result: run status infra_dnf  claims_done=0.01  ✓
macbook_pro_16__opencode__deepseek-flash FAIL          composite=None  infra, not a model result: run status infra_dnf_prompt_not_delivered  ✓
pelican_bicycle__claude-code__glm-5.3__a FAIL          composite=None  infra, not a model result: run status infra_dnf  ✓
pelican_bicycle__claude-code__glm-5.3__a FAIL          composite=None  infra, not a model result: run status infra_dnf  ✓
pelican_bicycle_animated__grok-build__gr FAIL          composite=None  infra, not a model result: run status infra_dnf  claims_done=0.01  ✓
ps5_controller__claude-code__glm-5.3__at FAIL          composite=None  infra, not a model result: run status infra_dnf  claims_done=0.02  ✓
agreement: 6/6
Run it
cd samples/run-verdict && ./run.sh
# = python3 verdict.py fixtures/runs.jsonl                        (6 hand-written runs)
#   python3 verdict.py fixtures/mr_infra_dnf.jsonl --judge-failed (6 REAL AI Model Benchmark infra DNFs, scrubbed)