Jev Labs

Patch Verdict

did the coding agent's patch really fix the issue?

Details · results, methods & full write-up
← All projects

Project 14

Patch Verdict

did the coding agent's patch really fix the issue?

Details

How Jev is used

Code first: it parses the final patch (files, test files, new scratch scripts, lines changed), finds the last edit and the last test command in the run, and reads that command's exit code. The model then answers typed questions about the same state.

ArmCode decidesModel questions
A: code rulereject if nothing changed outside tests and new scratch scripts; accept if a source change was followed by a passing test run; otherwise unknownnone
B: one callnothingNoul: will the hidden tests pass with the final patch?
C: stagedA's reject rule, without a callNoul addresses_issue, Noul verified_after_last_edit, Choice failure_mode, Score risk; p = mean(addresses_issue, verified_after_last_edit, P(likely_fixed))

The state holds the issue, the clipped final patch, the last test command, its exit code and output tail, the agent's final message (its own claim, which can be wrong) and the code facts. Labels, run and issue IDs, the repository name and the licence never reach the model; a test checks this for every case.

Auto-accept and auto-reject thresholds are chosen on the dev split only, as the widest bands with at least 90% precision, then applied unchanged to the test split and the Zen subset (score.py).

Cost, tokens, latency

Each case records requests, reported input and output tokens, HTTP latency and elapsed time for both arms. Both free routes cost $0; the list-price estimate uses $0.042 per million input tokens with output free (research/JEV_API.md), and is labelled as an estimate.

Results

Live runs on both free routes, reported separately. Every number below is generated from the artifacts in results/ by samples/_shared/report_index.py; threshold transfer uses the same function as score.py.

  • In the first complete test repeat (v2_codiv_full_r1.json), with openjev (Codiv free hosted) - NOT JEV on the 150-run test split, no arm separated resolved from unresolved runs better than chance; every AUROC interval includes 0.5 (code rule 0.543 [0.468, 0.627], one call 0.589 [0.491, 0.679], staged 0.522 [0.426, 0.619]).
  • The staged pipeline did not beat one call. Its AUROC is lower, and on correctness at the 0.5 cut McNemar finds 6 runs only one call got right and 8 runs only staged got right (p = 0.7905).
  • Both model arms' Brier scores (0.4324 and 0.3643) are worse than always answering 0.5 (0.25), so their probabilities are badly calibrated.
  • Staged led on the dev split (0.643 [0.502, 0.788] against 0.548 [0.392, 0.689] for one call), and that lead did not replicate on the test split.
  • With thresholds frozen on dev, one call auto-decides 22 of 150 runs for openjev (Codiv free hosted) - NOT JEV on the test split, 16 of them correctly (73%), below the 90% precision the band was chosen for.
  • With thresholds frozen on dev, staged makes no automatic decisions, because no band reached 90% precision on dev.
  • With thresholds frozen on dev, one call auto-decides 77 of 77 runs for REAL JEV (OpenCode Zen free, jev-1.13-free) on the test split, 27 of them correctly (35%), below the 90% precision the band was chosen for.
  • With REAL JEV (OpenCode Zen free, jev-1.13-free) on 77 completed runs from the test split, one call AUROC is 0.664 [0.524, 0.782] and staged is 0.541 [0.405, 0.682]; the arms disagree on 3 runs at the 0.5 cut (McNemar p = 1.0).
  • REAL JEV (OpenCode Zen free, jev-1.13-free) stopped on rate_limit. These are the completed paired runs, not a full-test estimate; retained provider wording: “Rate limit exceeded. Please try again later.”.
  • On the same 77 runs, openjev (Codiv free hosted) - NOT JEV one call scored 0.609 [0.465, 0.737] and 0.590 [0.454, 0.715] and 0.616 [0.476, 0.743] across its test repeats; REAL JEV (OpenCode Zen free, jev-1.13-free) scored 0.664 [0.524, 0.782].
  • Intervals that exclude 0.5: one call on the test split (v2_codiv_full_r2.json); one call on the test split (v2_zen_full_r1.json); one call on the Zen subset (v2_zen_zen_r1.json); staged on the dev split (v2_codiv_dev_r1.json).
  • All 3 complete openjev (Codiv free hosted) - NOT JEV test repeats are included (v2_codiv_full_r1.json, v2_codiv_full_r2.json, v2_codiv_full_r3.json). The AUROC range is 0.042 for one call and 0.002 for staged. Against v2_codiv_full_r1.json, the largest verdict change at the 0.5 cut is 4 one-call and 6 staged runs.

openjev (Codiv free hosted) - NOT JEV, test split, 150 runs (75 resolved), v2_codiv_full_r1.json. McNemar, one call against staged at the 0.5 cut: 6 runs only one call got right, 8 runs only staged got right, p = 0.7905.

ArmAUROC [95% CI]CorrectBrierECERequests ok / attemptedInput / output tokensHTTP p50 / p95 ms
A: code rule0.543 [0.468, 0.627]45/83 decidedn/an/a0 / 00 / 0n/a
B: one call0.589 [0.491, 0.679]83/150 at 0.50.43240.4284150 / 150308,881 / 0620 / 1194
C: staged0.522 [0.426, 0.619]85/150 at 0.50.36430.3212149 / 150333,036 / 0796 / 1344

openjev (Codiv free hosted) - NOT JEV, test split, 150 runs (75 resolved), v2_codiv_full_r2.json. McNemar, one call against staged at the 0.5 cut: 3 runs only one call got right, 9 runs only staged got right, p = 0.146.

ArmAUROC [95% CI]CorrectBrierECERequests ok / attemptedInput / output tokensHTTP p50 / p95 ms
A: code rule0.543 [0.468, 0.627]45/83 decidedn/an/a0 / 00 / 0n/a
B: one call0.631 [0.535, 0.718]79/150 at 0.50.44210.4508150 / 151308,881 / 0512 / 1099
C: staged0.520 [0.424, 0.612]85/150 at 0.50.36560.3216149 / 149333,036 / 0586 / 1458

openjev (Codiv free hosted) - NOT JEV, test split, 150 runs (75 resolved), v2_codiv_full_r3.json. McNemar, one call against staged at the 0.5 cut: 5 runs only one call got right, 8 runs only staged got right, p = 0.5811.

ArmAUROC [95% CI]CorrectBrierECERequests ok / attemptedInput / output tokensHTTP p50 / p95 ms
A: code rule0.543 [0.468, 0.627]45/83 decidedn/an/a0 / 00 / 0n/a
B: one call0.592 [0.495, 0.683]81/150 at 0.50.43970.4482150 / 150308,881 / 0634 / 4200
C: staged0.522 [0.426, 0.617]84/150 at 0.50.36790.3227149 / 150333,036 / 0788 / 3969

REAL JEV (OpenCode Zen free, jev-1.13-free), test split, 77 runs (50 resolved), v2_zen_full_r1.json. McNemar, one call against staged at the 0.5 cut: 2 runs only one call got right, 1 run only staged got right, p = 1.0.

ArmAUROC [95% CI]CorrectBrierECERequests ok / attemptedInput / output tokensHTTP p50 / p95 ms
A: code rule0.457 [0.345, 0.572]29/47 decidedn/an/a0 / 00 / 0n/a
B: one call0.664 [0.524, 0.782]51/77 at 0.50.21980.082777 / 81159,013 / 1,540618 / 2630
C: staged0.541 [0.405, 0.682]50/77 at 0.50.24410.116477 / 77175,106 / 9,660492 / 2601

REAL JEV (OpenCode Zen free, jev-1.13-free), Zen subset, 22 runs (11 resolved), v2_zen_zen_r1.json. McNemar, one call against staged at the 0.5 cut: no runs only one call got right, no runs only staged got right, p = 1.0.

ArmAUROC [95% CI]CorrectBrierECERequests ok / attemptedInput / output tokensHTTP p50 / p95 ms
A: code rule0.455 [0.260, 0.643]7/15 decidedn/an/a0 / 00 / 0n/a
B: one call0.773 [0.551, 0.952]13/22 at 0.50.21160.210522 / 2245,500 / 440586 / 816
C: staged0.661 [0.417, 0.893]13/22 at 0.50.26150.205622 / 2250,098 / 2,760467 / 836

openjev (Codiv free hosted) - NOT JEV, dev split, 60 runs (30 resolved), v2_codiv_dev_r1.json. McNemar, one call against staged at the 0.5 cut: 1 run only one call got right, 1 run only staged got right, p = 1.0.

ArmAUROC [95% CI]CorrectBrierECERequests ok / attemptedInput / output tokensHTTP p50 / p95 ms
A: code rule0.383 [0.268, 0.504]14/35 decidedn/an/a0 / 00 / 0n/a
B: one call0.548 [0.392, 0.689]33/60 at 0.50.43090.439960 / 60128,701 / 0795 / 1522
C: staged0.643 [0.502, 0.788]33/60 at 0.50.34410.331160 / 60139,261 / 0942 / 1350
Dev-frozen thresholds applied to other passes
Scored artifactArmDev band: accept at or above / reject at or belowAuto-accepted (errors)Auto-rejected (errors)Correct among auto decisionsShare auto-decided
v2_codiv_full_r1.jsonB: one callnone / 0.9190 (0)22 (6)16/220.147
v2_codiv_full_r1.jsonC: stagednone / none0 (0)0 (0)n/a0.0
v2_codiv_full_r2.jsonB: one callnone / 0.9190 (0)21 (6)15/210.14
v2_codiv_full_r2.jsonC: stagednone / none0 (0)0 (0)n/a0.0
v2_codiv_full_r3.jsonB: one callnone / 0.9190 (0)19 (5)14/190.127
v2_codiv_full_r3.jsonC: stagednone / none0 (0)0 (0)n/a0.0
v2_zen_full_r1.jsonB: one callnone / 0.9190 (0)77 (50)27/771.0
v2_zen_full_r1.jsonC: stagednone / none0 (0)0 (0)n/a0.0
v2_zen_zen_r1.jsonB: one callnone / 0.9190 (0)22 (11)11/221.0
v2_zen_zen_r1.jsonC: stagednone / none0 (0)0 (0)n/a0.0
Real Jev and openjev on the same runs
Route (artifact)RunsOne call AUROC [95% CI]Staged AUROC [95% CI]One call correct at 0.5Staged correct at 0.5McNemar p
REAL JEV (OpenCode Zen free, jev-1.13-free) (v2_zen_full_r1.json)770.664 [0.524, 0.782]0.541 [0.405, 0.682]51/7750/771.0
openjev (Codiv free hosted) - NOT JEV (v2_codiv_full_r1.json)770.609 [0.465, 0.737]0.556 [0.413, 0.702]51/7751/771.0
openjev (Codiv free hosted) - NOT JEV (v2_codiv_full_r2.json)770.590 [0.454, 0.715]0.544 [0.405, 0.689]50/7752/770.7266
openjev (Codiv free hosted) - NOT JEV (v2_codiv_full_r3.json)770.616 [0.476, 0.743]0.537 [0.400, 0.681]50/7751/771.0
Repeat variation (openjev, test split)
Repeat (artifact)One call AUROC [95% CI]Staged AUROC [95% CI]One call correct at 0.5Staged correct at 0.5Verdicts flipped against v2_codiv_full_r1.json: one call / staged
v2_codiv_full_r1.json0.589 [0.491, 0.679]0.522 [0.426, 0.619]83/15085/150reference
v2_codiv_full_r2.json0.631 [0.535, 0.718]0.520 [0.424, 0.612]79/15085/1504 / 6
v2_codiv_full_r3.json0.592 [0.495, 0.683]0.522 [0.426, 0.617]81/15084/1502 / 3
Requests, tokens, latency, time and cost
ArtifactRouteRequests ok / attemptedInput / output tokensHTTP p50 / p95 msSummed call time, sLast process wall, sResumed afterList price per 1,000 runs
v2_codiv_full_r1.jsonopenjev (Codiv free hosted) - NOT JEV299 / 300641,917 / 0701 / 1342236.2173.346http_error$0.1797
v2_codiv_full_r2.jsonopenjev (Codiv free hosted) - NOT JEV299 / 300641,917 / 0561 / 1207191.562.416spend guard: JEV_MAX_CALLS=200 reached$0.1797
v2_codiv_full_r3.jsonopenjev (Codiv free hosted) - NOT JEV299 / 300641,917 / 0729 / 4200365.3179.86http_error$0.1797
v2_zen_full_r1.jsonREAL JEV (OpenCode Zen free, jev-1.13-free)154 / 158334,119 / 11,200583 / 26012514.015.874rate_limit$0.1822
v2_zen_zen_r1.jsonREAL JEV (OpenCode Zen free, jev-1.13-free)44 / 4495,598 / 3,200536 / 826689.0880.621no$0.1825
v2_codiv_dev_r1.jsonopenjev (Codiv free hosted) - NOT JEV120 / 120267,962 / 0864 / 1522108.3109.456no$0.1876

Summed call time counts every attempt, including HTTP time and, on Zen, the 15-second lock cooldown; the campaign's pacing between calls falls outside it, which is why the Zen process wall is longer. The last-process wall covers only the final process of a resumed pass. Both free routes cost $0; the last column prices the reported input tokens at Jev's list price, output free. Attempt counts include recorded local guards, which send no HTTP request and use no provider tokens; usage of HTTP failures may be unknown.

Zen full-test stop history

Each row is a distinct saved stop. Archived copies of the current stop appear once. Counts are cumulative at that stop and are not added to request totals or counted as repeats. HTTP status and provider wording appear only when recorded; a failure class does not supply them. Attempt starts are runner timestamps, before any client pacing or lock waits.

Route / artifactAttempt started (UTC)Stopped at (UTC)Completed paired runsRequests ok / attemptedSaved HTTP statusFailure classProvider wording
REAL JEV (OpenCode Zen free, jev-1.13-free) / interrupted/v2_zen_full_r1_1.jsonunavailable (not recorded)2026-09-24 20:50:31 UTC77154 / 155unavailablerate_limitprovider wording unavailable (not recorded)
REAL JEV (OpenCode Zen free, jev-1.13-free) / interrupted/v2_zen_full_r1_2.json2026-09-24 21:09:40 UTC2026-09-24 21:09:55 UTC77154 / 156429rate_limitretained provider wording: “Rate limit exceeded. Please try again later.”
REAL JEV (OpenCode Zen free, jev-1.13-free) / interrupted/v2_zen_full_r1_3.json2026-09-24 21:25:11 UTC2026-09-24 21:25:26 UTC77154 / 157429rate_limitretained provider wording: “Rate limit exceeded. Please try again later.”
REAL JEV (OpenCode Zen free, jev-1.13-free) / v2_zen_full_r1.json2026-09-24 21:55:38 UTC2026-09-24 21:55:53 UTC77154 / 158429rate_limitretained provider wording: “Rate limit exceeded. Please try again later.”
  • openjev (Codiv free hosted) - NOT JEV, interrupted/v2_codiv_full_r1_1.json stopped after 41 of 150 test runs on http_error: a stop recorded by the runner. This snapshot is not counted as another repeat.
  • openjev (Codiv free hosted) - NOT JEV, interrupted/v2_codiv_full_r2_1.json stopped after 100 of 150 test runs on spend guard: JEV_MAX_CALLS=200 reached: the client's local call cap at the time, not a provider fault. This snapshot is not counted as another repeat.
  • openjev (Codiv free hosted) - NOT JEV, interrupted/v2_codiv_full_r3_1.json stopped after 33 of 150 test runs on http_error: a stop recorded by the runner. This snapshot is not counted as another repeat.
Code rule (no model)

From the frozen fixtures; the offline artifacts results/v2_code_*_r1.json hold the same counts.

SubsetRuns (resolved)Accepts (resolved among them)Rejects (unresolved among them)UnknownDecided correct
test split150 (75)82 (44)1 (1)6745/83
dev split60 (30)35 (14)0 (0)2514/35
Zen subset22 (11)15 (7)0 (0)77/15

38 of the code rule's 82 test-split accepts are unresolved: the agent's own last test run passed and the hidden tests did not. That is the gap the model arms had to close.

The code gate rejects 1 of the 150 test runs, so the staged arm calls the model on 149 of them.

RESULTS.md logs every pass.

What would change the result
  • Facts the hidden tests depend on. The model sees the issue, the final patch, the last test run and the agent's claim. It never sees the hidden tests, the files they cover, or whether the agent's own tests exercise the reported behaviour. Code can extract those facts from the task data, and the typed questions can then judge them.
  • A larger real-Jev sample. Real Jev's one-call result on the Zen subset is the most promising number here, and the subset is too small to call. A paced pass over the full test split, spread across several campaigns under the shared lock, would settle whether it holds.
  • Thresholds measured on the route that uses them. The openjev dev band does not carry over to real Jev, whose probabilities sit on a different scale. Real-Jev thresholds need a real-Jev dev pass.
  • Calibration before cut-offs. On the test split both openjev arms are calibrated worse than a constant guess. Recalibrating their probabilities on dev, for example by isotonic regression in code, has to come before any accept or reject band can be trusted.
Data

build_fixtures.py fetched only the rows it needed through the Hugging Face datasets-server API, in pages of four runs, and kept no raw download. It samples pages in a seeded random order across the whole dataset, keeps one run per issue and at most four issues per repository, looks up each repository's licence in the task dataset, scrubs emails, URLs and home paths, and drops any case that still trips the checkpoint gate. fixtures/MANIFEST.json records the dataset revisions, each row's provenance and licence, the selection settings, and the sha256 of fixtures/cases.jsonl.

The splits are balanced by label, so the test split's resolved share is not the agent's real resolve rate; precision at natural prevalence has to be recomputed from the per-case scores.

Limits
  • Published tests of Jev judging whole agent trajectories were weak, so this project gives the model code-extracted facts, the final patch and the last test run instead of raw history. It may still fail; the pre-registration says how that would be reported.
  • One agent scaffold (OpenHands) and one source benchmark; other agents may leave different traces.
  • Patches are clipped at 3,500 characters and test output to its last 1,200; long fixes lose context.
  • Hidden-test labels can be flaky, and scratch-file detection is a filename heuristic.
  • No frontier-model arm: none is free. An optional same-weights chat baseline (Codiv diffusiongemma-26b) waits until someone confirms the existing key covers that endpoint.
Verification
python3 -m unittest discover -s tests -p 'test_patch_verdict.py'
scripts/check.sh --allow-site-findings      # from the repository root: all tests, the gate and every project's mock run
Raw result files
CODE ONLY - NO MODELresults/v2_code_dev_r1.json
{
 "schema_version": 2,
 "project": "patch-verdict",
 "subset": "dev",
 "repeat": 1,
 "limit": null,
 "timestamp": "2026-09-24T18:54:11Z",
 "input_hashes": {
  "samples/patch-verdict/fixtures/cases.jsonl": "efa48dde7d27f9d5818f94591b87a8859496096b2ddcdc21a546c9c894975a2d"
 },
 "source_hashes": {
  "samples/_shared/jev.py": "0ab1237a037b005379d3b6de37834a7c552435671d111be5c1cc41e96d0636ac",
  "samples/patch-verdict/../_shared/community.py": "4cd057b6bdf56c2f0ceddd624b6d7098b26c2186bac562bb0aba3f9d03c8b8ee",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "label": "code only (no model)",
 "mode": "code",
 "provider": null,
 "rows": [],
 "summary": {
  "dataset": {
   "subset": "dev",
   "cases": 60,
   "resolved": 30,
   "limit": null,
   "plan": {
    "cases": 60,
    "one_call_requests": 60,
    "staged_requests": 60,
    "total_requests": 120,
    "zen_minutes_at_spacing": 40.0
   }
  },
  "code_rule": {
   "accept": 35,
   "reject": 0,
   "unknown": 25,
   "decided_correct": "14/35"
  }
 },
 "policy": {
  "version": "patch-verdict-v1",
  "code_gate_rejects": "no patch, or no change outside tests and new scratch scripts",
  "code_rule_accepts": "source change + last test run after the last edit + exit code 0",
  "one_call_score": "p = Noul(passes)",
  "staged_score": "p = mean(Noul(addresses_issue), Noul(verified_after_last_edit), P(failure_mode=likely_fixed))",
  "default_threshold": 0.5,
  "thresholds": "auto-accept/auto-reject cut-offs are chosen on the dev split (openjev) and then frozen",
  "frozen_before_live_runs": true
 },
 "status": "complete",
 "calls": [],
 "metrics": {
  "requests": 0,
  "completed_requests": 0,
  "input_tokens": 0,
  "output_tokens": 0,
  "median_latency_ms": null,
  "wall_s": 0,
  "cost_usd": 0.0,
  "cost_basis": "free route; missing provider cost fields are not receipts",
  "reported_cost_usd": null
 },
 "wall_s": 0.048
}

2172 of 2172 bytes shown. Complete safe artifact

CODE ONLY - NO MODELresults/v2_code_full_r1.json
{
 "schema_version": 2,
 "project": "patch-verdict",
 "subset": "full",
 "repeat": 1,
 "limit": null,
 "timestamp": "2026-09-24T18:54:11Z",
 "input_hashes": {
  "samples/patch-verdict/fixtures/cases.jsonl": "efa48dde7d27f9d5818f94591b87a8859496096b2ddcdc21a546c9c894975a2d"
 },
 "source_hashes": {
  "samples/_shared/jev.py": "0ab1237a037b005379d3b6de37834a7c552435671d111be5c1cc41e96d0636ac",
  "samples/patch-verdict/../_shared/community.py": "4cd057b6bdf56c2f0ceddd624b6d7098b26c2186bac562bb0aba3f9d03c8b8ee",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "label": "code only (no model)",
 "mode": "code",
 "provider": null,
 "rows": [],
 "summary": {
  "dataset": {
   "subset": "full",
   "cases": 150,
   "resolved": 75,
   "limit": null,
   "plan": {
    "cases": 150,
    "one_call_requests": 150,
    "staged_requests": 149,
    "total_requests": 299,
    "zen_minutes_at_spacing": 99.7
   }
  },
  "code_rule": {
   "accept": 82,
   "reject": 1,
   "unknown": 67,
   "decided_correct": "45/83"
  }
 },
 "policy": {
  "version": "patch-verdict-v1",
  "code_gate_rejects": "no patch, or no change outside tests and new scratch scripts",
  "code_rule_accepts": "source change + last test run after the last edit + exit code 0",
  "one_call_score": "p = Noul(passes)",
  "staged_score": "p = mean(Noul(addresses_issue), Noul(verified_after_last_edit), P(failure_mode=likely_fixed))",
  "default_threshold": 0.5,
  "thresholds": "auto-accept/auto-reject cut-offs are chosen on the dev split (openjev) and then frozen",
  "frozen_before_live_runs": true
 },
 "status": "complete",
 "calls": [],
 "metrics": {
  "requests": 0,
  "completed_requests": 0,
  "input_tokens": 0,
  "output_tokens": 0,
  "median_latency_ms": null,
  "wall_s": 0,
  "cost_usd": 0.0,
  "cost_basis": "free route; missing provider cost fields are not receipts",
  "reported_cost_usd": null
 },
 "wall_s": 0.045
}

2178 of 2178 bytes shown. Complete safe artifact

CODE ONLY - NO MODELresults/v2_code_zen_r1.json
{
 "schema_version": 2,
 "project": "patch-verdict",
 "subset": "zen",
 "repeat": 1,
 "limit": null,
 "timestamp": "2026-09-24T18:54:11Z",
 "input_hashes": {
  "samples/patch-verdict/fixtures/cases.jsonl": "efa48dde7d27f9d5818f94591b87a8859496096b2ddcdc21a546c9c894975a2d"
 },
 "source_hashes": {
  "samples/_shared/jev.py": "0ab1237a037b005379d3b6de37834a7c552435671d111be5c1cc41e96d0636ac",
  "samples/patch-verdict/../_shared/community.py": "4cd057b6bdf56c2f0ceddd624b6d7098b26c2186bac562bb0aba3f9d03c8b8ee",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "label": "code only (no model)",
 "mode": "code",
 "provider": null,
 "rows": [],
 "summary": {
  "dataset": {
   "subset": "zen",
   "cases": 22,
   "resolved": 11,
   "limit": null,
   "plan": {
    "cases": 22,
    "one_call_requests": 22,
    "staged_requests": 22,
    "total_requests": 44,
    "zen_minutes_at_spacing": 14.7
   }
  },
  "code_rule": {
   "accept": 15,
   "reject": 0,
   "unknown": 7,
   "decided_correct": "7/15"
  }
 },
 "policy": {
  "version": "patch-verdict-v1",
  "code_gate_rejects": "no patch, or no change outside tests and new scratch scripts",
  "code_rule_accepts": "source change + last test run after the last edit + exit code 0",
  "one_call_score": "p = Noul(passes)",
  "staged_score": "p = mean(Noul(addresses_issue), Noul(verified_after_last_edit), P(failure_mode=likely_fixed))",
  "default_threshold": 0.5,
  "thresholds": "auto-accept/auto-reject cut-offs are chosen on the dev split (openjev) and then frozen",
  "frozen_before_live_runs": true
 },
 "status": "complete",
 "calls": [],
 "metrics": {
  "requests": 0,
  "completed_requests": 0,
  "input_tokens": 0,
  "output_tokens": 0,
  "median_latency_ms": null,
  "wall_s": 0,
  "cost_usd": 0.0,
  "cost_basis": "free route; missing provider cost fields are not receipts",
  "reported_cost_usd": null
 },
 "wall_s": 0.045
}

2169 of 2169 bytes shown. Complete safe artifact

openjev (Codiv free hosted) - NOT JEVresults/v2_codiv_dev_r1.json
{
 "schema_version": 2,
 "project": "patch-verdict",
 "subset": "dev",
 "repeat": 1,
 "limit": null,
 "timestamp": "2026-09-24T18:58:42Z",
 "input_hashes": {
  "samples/patch-verdict/fixtures/cases.jsonl": "efa48dde7d27f9d5818f94591b87a8859496096b2ddcdc21a546c9c894975a2d"
 },
 "source_hashes": {
  "samples/_shared/jev.py": "0ab1237a037b005379d3b6de37834a7c552435671d111be5c1cc41e96d0636ac",
  "samples/patch-verdict/../_shared/community.py": "4cd057b6bdf56c2f0ceddd624b6d7098b26c2186bac562bb0aba3f9d03c8b8ee",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "label": "openjev (Codiv free hosted) - NOT JEV",
 "mode": "live",
 "provider": "codiv",
 "rows": [
  {
   "id": "pv_0048736ad6",
   "resolved": false,
   "code": "unknown",
   "gate_rejected": false,
   "p_one_call": 0.6691030936487606,
   "p_staged": 0.25339775597585107,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2475,
    "output_tokens": 0,
    "median_latency_ms": 1521.6,
    "wall_s": 1.522,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "unclear",
   "risk": 0.7650568654607963,
   "metrics_staged": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2651,
    "output_tokens": 0,
    "median_latency_ms": 1121.6,
    "wall_s": 1.122,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 1
  },
  {
   "id": "pv_01b5176583",
   "resolved": true,
   "code": "unknown",
   "gate_rejected": false,
   "p_one_call": 0.9998849252414705,
   "p_staged": 0.8450752239064183,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2318,
    "output_tokens": 0,
    "median_latency_ms": 1166.2,
    "wall_s": 1.166,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "likely_fixed",
   "risk": 0.7837989122790144,
   "metrics_staged": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2494,
    "output_tokens": 0,
    "median_latency_ms": 1602.3,
    "wall_s": 1.603,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 1
  },
  {
   "id": "pv_020b31a16f",
   "resolved": true,
   "code": "unknown",
   "gate_rejected": false,
   "p_one_call": 0.9999939914883107,
   "p_staged": 0.9999049917126216,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 1661,
    "output_tokens": 0,

… truncated, 3000 of 192437 bytes shown. Complete safe artifact

openjev (Codiv free hosted) - NOT JEVresults/v2_codiv_full_r1.json
{
 "schema_version": 2,
 "project": "patch-verdict",
 "subset": "full",
 "repeat": 1,
 "limit": null,
 "timestamp": "2026-09-24T19:27:00Z",
 "input_hashes": {
  "samples/patch-verdict/fixtures/cases.jsonl": "efa48dde7d27f9d5818f94591b87a8859496096b2ddcdc21a546c9c894975a2d"
 },
 "source_hashes": {
  "samples/_shared/jev.py": "bd3cdc5ce9c7414e609d5d43db4d38941393951677493abf544153472a41fa9d",
  "samples/patch-verdict/../_shared/community.py": "96f6d81afffe8914293fda4b963b6a2e7a2a02bd8c240f01af4c20c3ac1ecc95",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "label": "openjev (Codiv free hosted) - NOT JEV",
 "mode": "live",
 "provider": "codiv",
 "call_cap": 314,
 "resumed_after": "http_error",
 "original_source_hashes": {
  "samples/_shared/jev.py": "0ab1237a037b005379d3b6de37834a7c552435671d111be5c1cc41e96d0636ac",
  "samples/patch-verdict/../_shared/community.py": "4cd057b6bdf56c2f0ceddd624b6d7098b26c2186bac562bb0aba3f9d03c8b8ee",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "rows": [
  {
   "id": "pv_322dd01fa9",
   "resolved": false,
   "code": "accept",
   "gate_rejected": false,
   "p_one_call": 0.9995195782261261,
   "p_staged": 0.9997266743538299,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2315,
    "output_tokens": 0,
    "median_latency_ms": 554.4,
    "wall_s": 0.555,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "likely_fixed",
   "risk": 0.27756626134947526,
   "metrics_staged": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2491,
    "output_tokens": 0,
    "median_latency_ms": 1341.8,
    "wall_s": 1.342,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 1
  },
  {
   "id": "pv_333f5a6224",
   "resolved": false,
   "code": "unknown",
   "gate_rejected": false,
   "p_one_call": 0.9995553529768693,
   "p_staged": 0.9998508825546204,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2348,
    "output_tokens": 0,
    "median_latency_ms": 580.6,
    "wall_s": 0.581,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "likely_fixed",
   "risk": 

… truncated, 2985 of 475785 bytes shown. Complete safe artifact

openjev (Codiv free hosted) - NOT JEVresults/v2_codiv_full_r2.json
{
 "schema_version": 2,
 "project": "patch-verdict",
 "subset": "full",
 "repeat": 2,
 "limit": null,
 "timestamp": "2026-09-24T19:58:57Z",
 "input_hashes": {
  "samples/patch-verdict/fixtures/cases.jsonl": "efa48dde7d27f9d5818f94591b87a8859496096b2ddcdc21a546c9c894975a2d"
 },
 "source_hashes": {
  "samples/_shared/jev.py": "bd3cdc5ce9c7414e609d5d43db4d38941393951677493abf544153472a41fa9d",
  "samples/patch-verdict/../_shared/community.py": "51898960baadf029853581bec51ebdffdfd19cac056b4e113906741899331e5a",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "label": "openjev (Codiv free hosted) - NOT JEV",
 "mode": "live",
 "provider": "codiv",
 "call_cap": 314,
 "resumed_after": "spend guard: JEV_MAX_CALLS=200 reached",
 "original_source_hashes": {
  "samples/_shared/jev.py": "0ab1237a037b005379d3b6de37834a7c552435671d111be5c1cc41e96d0636ac",
  "samples/patch-verdict/../_shared/community.py": "4cd057b6bdf56c2f0ceddd624b6d7098b26c2186bac562bb0aba3f9d03c8b8ee",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "rows": [
  {
   "id": "pv_322dd01fa9",
   "resolved": false,
   "code": "accept",
   "gate_rejected": false,
   "p_one_call": 0.9996258709459511,
   "p_staged": 0.9998046511960953,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2315,
    "output_tokens": 0,
    "median_latency_ms": 1122.6,
    "wall_s": 1.123,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "likely_fixed",
   "risk": 0.32537287431059886,
   "metrics_staged": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2491,
    "output_tokens": 0,
    "median_latency_ms": 677.9,
    "wall_s": 0.678,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 2
  },
  {
   "id": "pv_333f5a6224",
   "resolved": false,
   "code": "unknown",
   "gate_rejected": false,
   "p_one_call": 0.9998123799039462,
   "p_staged": 0.999938112054518,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2348,
    "output_tokens": 0,
    "median_latency_ms": 1248.9,
    "wall_s": 1.249,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "likely_fixed"

… truncated, 3000 of 475632 bytes shown. Complete safe artifact

openjev (Codiv free hosted) - NOT JEVresults/v2_codiv_full_r3.json
{
 "schema_version": 2,
 "project": "patch-verdict",
 "subset": "full",
 "repeat": 3,
 "limit": null,
 "timestamp": "2026-09-24T19:33:38Z",
 "input_hashes": {
  "samples/patch-verdict/fixtures/cases.jsonl": "efa48dde7d27f9d5818f94591b87a8859496096b2ddcdc21a546c9c894975a2d"
 },
 "source_hashes": {
  "samples/_shared/jev.py": "bd3cdc5ce9c7414e609d5d43db4d38941393951677493abf544153472a41fa9d",
  "samples/patch-verdict/../_shared/community.py": "96f6d81afffe8914293fda4b963b6a2e7a2a02bd8c240f01af4c20c3ac1ecc95",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "label": "openjev (Codiv free hosted) - NOT JEV",
 "mode": "live",
 "provider": "codiv",
 "call_cap": 314,
 "resumed_after": "http_error",
 "original_source_hashes": {
  "samples/_shared/jev.py": "0ab1237a037b005379d3b6de37834a7c552435671d111be5c1cc41e96d0636ac",
  "samples/patch-verdict/../_shared/community.py": "4cd057b6bdf56c2f0ceddd624b6d7098b26c2186bac562bb0aba3f9d03c8b8ee",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "rows": [
  {
   "id": "pv_322dd01fa9",
   "resolved": false,
   "code": "accept",
   "gate_rejected": false,
   "p_one_call": 0.9994297483479814,
   "p_staged": 0.9990120477742114,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2315,
    "output_tokens": 0,
    "median_latency_ms": 4325.8,
    "wall_s": 4.326,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "likely_fixed",
   "risk": 0.31930004749855034,
   "metrics_staged": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2491,
    "output_tokens": 0,
    "median_latency_ms": 4759.1,
    "wall_s": 4.76,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 3
  },
  {
   "id": "pv_333f5a6224",
   "resolved": false,
   "code": "unknown",
   "gate_rejected": false,
   "p_one_call": 0.9997645829312177,
   "p_staged": 0.9999409954867439,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2348,
    "output_tokens": 0,
    "median_latency_ms": 2906.4,
    "wall_s": 2.907,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "likely_fixed",
   "risk": 

… truncated, 2986 of 475882 bytes shown. Complete safe artifact

REAL JEV (OpenCode Zen free, jev-1.13-free)results/v2_zen_full_r1.json
{
 "schema_version": 2,
 "project": "patch-verdict",
 "subset": "full",
 "repeat": 1,
 "limit": null,
 "campaign": "m3-pv-zen-full",
 "timestamp": "2026-09-24T21:55:37Z",
 "input_hashes": {
  "samples/patch-verdict/fixtures/cases.jsonl": "efa48dde7d27f9d5818f94591b87a8859496096b2ddcdc21a546c9c894975a2d"
 },
 "source_hashes": {
  "samples/_shared/jev.py": "a20ab00f05c9fc838194f68b33f88b64bca3a0dd7e1ddf78d4c347bf9cdf3322",
  "samples/patch-verdict/../_shared/community.py": "79d28fe055b35dada9eeb70122da25ae3487e901c0d0992ba1270207ad553968",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "label": "REAL JEV (OpenCode Zen free, jev-1.13-free)",
 "mode": "live",
 "provider": "zen",
 "call_cap": 314,
 "resumed_after": "rate_limit",
 "original_source_hashes": {
  "samples/_shared/jev.py": "a20ab00f05c9fc838194f68b33f88b64bca3a0dd7e1ddf78d4c347bf9cdf3322",
  "samples/patch-verdict/../_shared/community.py": "79d28fe055b35dada9eeb70122da25ae3487e901c0d0992ba1270207ad553968",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "status": "partial",
 "rows": [
  {
   "id": "pv_322dd01fa9",
   "resolved": false,
   "code": "accept",
   "gate_rejected": false,
   "p_one_call": 0.53,
   "p_staged": 0.53,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2266,
    "output_tokens": 20,
    "median_latency_ms": 620.7,
    "wall_s": 15.622,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "likely_fixed",
   "risk": 2.13,
   "metrics_staged": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2475,
    "output_tokens": 125,
    "median_latency_ms": 488.2,
    "wall_s": 15.49,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 1
  },
  {
   "id": "pv_333f5a6224",
   "resolved": false,
   "code": "unknown",
   "gate_rejected": false,
   "p_one_call": 0.56,
   "p_staged": 0.6233333333333334,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2247,
    "output_tokens": 20,
    "median_latency_ms": 618.8,
    "wall_s": 15.619,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "likely_fixed",
   "risk": 2.16,
   

… truncated, 2999 of 225844 bytes shown. Complete safe artifact

REAL JEV (OpenCode Zen free, jev-1.13-free)results/v2_zen_zen_r1.json
{
 "schema_version": 2,
 "project": "patch-verdict",
 "subset": "zen",
 "repeat": 1,
 "limit": null,
 "timestamp": "2026-09-24T19:01:42Z",
 "input_hashes": {
  "samples/patch-verdict/fixtures/cases.jsonl": "efa48dde7d27f9d5818f94591b87a8859496096b2ddcdc21a546c9c894975a2d"
 },
 "source_hashes": {
  "samples/_shared/jev.py": "0ab1237a037b005379d3b6de37834a7c552435671d111be5c1cc41e96d0636ac",
  "samples/patch-verdict/../_shared/community.py": "4cd057b6bdf56c2f0ceddd624b6d7098b26c2186bac562bb0aba3f9d03c8b8ee",
  "samples/patch-verdict/build_fixtures.py": "a2d631aeb30be78a045af0b976b8e42dbd8bedc7a8dc360f7a69a25579bec4ea",
  "samples/patch-verdict/patch_verdict.py": "367bdb7bc22076e6e7440653674e14546618c90943eb414088ab7b452d15853d",
  "samples/patch-verdict/score.py": "59231c58f7c0965ab60c2c7d055be6fb64211f427ece31873a811a80e720a44c"
 },
 "label": "REAL JEV (OpenCode Zen free, jev-1.13-free)",
 "mode": "live",
 "provider": "zen",
 "rows": [
  {
   "id": "pv_322dd01fa9",
   "resolved": false,
   "code": "accept",
   "gate_rejected": false,
   "p_one_call": 0.54,
   "p_staged": 0.5599999999999999,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2266,
    "output_tokens": 20,
    "median_latency_ms": 628.7,
    "wall_s": 15.629,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "likely_fixed",
   "risk": 2.04,
   "metrics_staged": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2475,
    "output_tokens": 125,
    "median_latency_ms": 835.6,
    "wall_s": 15.836,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 1
  },
  {
   "id": "pv_333f5a6224",
   "resolved": false,
   "code": "unknown",
   "gate_rejected": false,
   "p_one_call": 0.59,
   "p_staged": 0.6166666666666667,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2247,
    "output_tokens": 20,
    "median_latency_ms": 480.1,
    "wall_s": 15.481,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "failure_mode": "likely_fixed",
   "risk": 2.08,
   "metrics_staged": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 2456,
    "output_tokens": 125,
    "median_latency_ms": 415.9,
    "wall_s": 15.417,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 1
  },
  {
   "id": "pv_3430efee5f",
   "resolved": true,
   "code": "unknown",
   "gate_rejected": false,
   "p_one_call": 0.78,
   "p_staged": 0.7566666666666667,
   "metrics_one_call": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 1700,
    "output_tokens": 20,
    "median_latency_ms": 608.3,
    "wall_s": 15.609,

… truncated, 2999 of 66150 bytes shown. Complete safe artifact

Run it
./run.sh --plan --subset zen                          # exact request counts; no client, no network
./run.sh --offline --repeat-start 4                   # fixtures + code rule; 0 model calls
./run.sh --mock --limit 20 --repeat-start 4           # MOCK - NOT JEV plumbing check; never a result
JEV_PROVIDER=codiv JEV_KEY_FILE="$CODIV_KEY_FILE" JEV_MAX_USD=0 ./run.sh --subset dev --repeat-start 1
JEV_PROVIDER=codiv JEV_KEY_FILE="$CODIV_KEY_FILE" JEV_MAX_USD=0 ./run.sh --subset full --repeat-start 1
python3 score.py results/v2_codiv_dev_r1.json results/v2_codiv_full_r1.json
JEV_PROVIDER=zen JEV_KEY_FILE="$ZEN_KEY_FILE" JEV_ZEN_LOCK="$ZEN_LOCK_FILE" JEV_MAX_USD=0 JEV_MIN_INTERVAL_S=20 \
  ./run.sh --subset zen --repeat-start 1 --campaign pv-zen-1