Jev Labs

Find the bug.

Review a diff. See the line, risk and verdict.

REAL JEV Tied Across Both Strategies

Same detailed questions in both arms. The routes used different-sized authored sets.

REAL JEV (OpenCode Zen free, jev-1.13-free)

Eight-case subset: both arms tied.

Staged Review

0/4false alarms4/4 bugs caughtC = clear · A = alarm

One Typed Call

0/4false alarms4/4 bugs caughtC = clear · A = alarm
Review evidence artifact

476.7 ms median HTTP time · 20 successful requests · $0 expected free-route cost

openjev (Codiv free hosted) - NOT JEV

Original authored diffs: staging traded one missed bug for one fewer false alarm.

Staged Review

0/9false alarms10/11 bugs caughtC = clear · A = alarm

One Typed Call

1/9false alarms11/11 bugs caughtC = clear · A = alarm
Review evidence artifact
Limits, latency & cost

One pass on authored Python diffs. These are two prompting strategies for each model, not a comparison with a frontier LLM. On openjev, staging used more requests and tokens.

Provider cost fields were absent. HTTP latency excludes lock waits and cooldowns; this is not a controlled speed comparison. Method: samples/review-triage/README.md.

Details · results, methods & full write-up
← All projects

Project 10

Review Triage

find planted bugs without alarming on clean changes

4/4bugs caught in both review armsREAL JEV
openjev (Codiv free hosted) - NOT JEVREAL JEV (OpenCode Zen free, jev-1.13-free)

Review must find behavior bugs, locate and explain them, and stay quiet on safe changes. We score line and mechanism accuracy alongside detection, including bugs discarded before detailed inspection.

Headline result

REAL JEV (OpenCode Zen free, jev-1.13-free)

8-case subset: staged recall 4/4, false alarms 0/4; one call 4/4, false alarms 0/4

openjev (Codiv free hosted) - NOT JEV

original 20: staged recall 10/11, false alarms 0/9; one call 11/11, false alarms 1/9

Details

The problem

Review must find behavior bugs, locate and explain them, and stay quiet on safe changes. We score line and mechanism accuracy alongside detection, including bugs discarded before detailed inspection.

Both arms receive identical old/new source, diff, and behavior contracts. Prompts exclude labels, case IDs, cohort, evaluator notes, gold lines, and expected mechanisms.

ArmTyped questionsCode decision
StagedRisk Score and area Choice, then detailed questions if risk ≥1.5Below the risk threshold, stop without a flag
One callThe same detailed questions immediatelyApply the same final flag rule

Detailed questions ask source-line Choice, mechanism Choice, concrete-failure Noul, and severity Score. Flags require Noul ≥0.5, severity ≥2, and mechanism other than none. We froze prompts and thresholds before hosted evaluation.

Identical detailed requests exclude earlier-stage answers, isolating the gate's effect. Line choices use exact old:N and new:N coordinates, including removed guards, unchanged context, and none. Repeated text at the wrong coordinate earns no credit. Localisation and mechanism credit require a flagged planted bug.

Cost, tokens, latency

Each completed case records requests, input/output tokens, latency, elapsed time, and cost, distinguishing provider values, estimates, and unknown usage. Route names alone supply no price evidence. Screening can save tokens on a safe diff when it skips pricier detailed review, but cannot beat the one-call request count; token and wall-time measurements capture that tradeoff.

What real Jev would change

Real Jev's typed distributions corrected the matched-case errors under frozen thresholds. Calibration and the remaining 20 cases remain untested by Jev; Codiv results cannot fill those gaps. Frozen prompts, identical detailed requests, and the preselected subset make the comparison reproducible; deployment claims need larger independent Python and non-Python diffs. Future rate-limit stops leave unrun cases pending: operator live run, excluded from review scores.

Results

First pass, 2026-09-24: openjev (Codiv free hosted) - NOT JEV, model openjev-0.1. Both arms completed all 28 cases; cohorts remain separate.

Cohort / armBug recallFalse alarmsCorrect lineRequests: completed / attemptedInput / output tokensMedian latency
original / staged10/110/99/1131 / 3319,203 / 0422.9 ms
original / one call11/111/910/1120 / 2017,687 / 0421.4 ms
challenge / staged3/40/43/413 / 137,611 / 0568.1 ms
challenge / one call3/40/43/48 / 86,585 / 0502.1 ms

In the first pass, staging avoided the timezone false alarm but lost an extra original bug at the severity threshold. Both arms missed the optional-value challenge there and chose the wrong optional-name line/mechanism. The risk gate missed no bugs. Staging added 2,542 input tokens (10.5%) and requests, without establishing a repeatable advantage.

Expected cost: $0 on the required free route; provider cost was unreported. Tokens cover successful requests; usage for the two failures is unknown.

RESULTS.md details failures, p95 latency, two connection-error resumes, and the complete artifact. It excludes MOCK - NOT JEV.

Three Codiv passes on the same fixtures

openjev (Codiv free hosted) - NOT JEV. The first pass remains the headline; each repeat appears separately. Cells show bug recall · false alarms · correct line:

RepeatOriginal stagedOriginal one callChallenge stagedChallenge one call
110/11 · 0/9 · 9/1111/11 · 1/9 · 10/113/4 · 0/4 · 3/43/4 · 0/4 · 3/4
211/11 · 0/9 · 10/1111/11 · 1/9 · 10/113/4 · 0/4 · 3/43/4 · 0/4 · 2/4
311/11 · 0/9 · 10/1110/11 · 1/9 · 9/113/4 · 0/4 · 3/43/4 · 0/4 · 3/4

b08_wrongvar crossed the severity threshold in opposite arms. One-call localisation dropped in repeat 2 on x03_repeated_line_bug. Both arms kept the wrong optional-name line; staged got its mechanism right in repeat 3. The optional-value miss and one-call timezone false alarm persisted. The risk gate missed no bugs.

RepeatCompleted / attemptedInput / output tokensSuccessful HTTP total, sCumulative attempt time, sRecorded process wall, sWall-time scope
172 / 7451,086 / 0195.134279.818108.660last resumed process only
272 / 7251,086 / 042.093141.698142.482complete uninterrupted process
372 / 7451,086 / 038.326263.81050.694last resumed process only

Each pass makes fresh calls on the same 28 authored diffs. Resumes restore completed responses only within that pass; the datasets are not independent. Cost remains unreported ($0 expected on the free route); failed-request usage is unknown. RESULTS.md gives per-arm resources, unstable scores, and resumes. Thresholds and labels stayed fixed.

Preselected eight-case route comparison

REAL JEV (OpenCode Zen free, jev-1.13-free) returned jev-1.13-free on its eight-case subset. We compare it with the same four bugs and four clean cases from the first Codiv pass.

Route / armBug recallFalse alarmsCorrect lineCorrect mechanism
openjev (Codiv free hosted) - NOT JEV / staged4/40/43/43/4
openjev (Codiv free hosted) - NOT JEV / one call4/41/43/43/4
REAL JEV (OpenCode Zen free, jev-1.13-free) / staged4/40/44/44/4
REAL JEV (OpenCode Zen free, jev-1.13-free) / one call4/40/44/44/4

Real Jev made no detection, localisation, mechanism, or false-alarm errors, correcting the optional-name line and avoiding the timezone false alarm. This small subset cannot establish success across all 28 cases or real repositories.

Route / armCompleted / attemptedInput / output tokensMedian / p95 HTTP latency, msCumulative attempt time, s
openjev (Codiv free hosted) - NOT JEV / staged12 / 137,312 / 0520.7 / 50,440.992.156
openjev (Codiv free hosted) - NOT JEV / one call8 / 87,059 / 0395.0 / 14,218.217.432
REAL JEV (OpenCode Zen free, jev-1.13-free) / staged12 / 1210,253 / 1,566466.4 / 2,066.0291.620
REAL JEV (OpenCode Zen free, jev-1.13-free) / one call8 / 89,414 / 1,965493.9 / 560.9186.085

Zen completed 20 requests / 20 attempts, reporting 19,667 input and 3,531 output tokens. HTTP took 11.993 seconds; execution took 513.773 seconds, including lock waits, pacing, and cooldowns. Neither provider reported cost; $0 is the expected free-route cost. Providers count tokens differently, and these timings offer no controlled speed comparison. RESULTS.md details accounting and provenance.

The problem and Jev/openjev method

Review must find behavior bugs, locate and explain them, and stay quiet on safe changes. We score line and mechanism accuracy alongside detection, including bugs discarded before detailed inspection.

Both arms receive identical old/new source, diff, and behavior contracts. Prompts exclude labels, case IDs, cohort, evaluator notes, gold lines, and expected mechanisms.

ArmTyped questionsCode decision
StagedRisk Score and area Choice, then detailed questions if risk ≥1.5Below the risk threshold, stop without a flag
One callThe same detailed questions immediatelyApply the same final flag rule

Detailed questions ask source-line Choice, mechanism Choice, concrete-failure Noul, and severity Score. Flags require Noul ≥0.5, severity ≥2, and mechanism other than none. We froze prompts and thresholds before hosted evaluation.

Identical detailed requests exclude earlier-stage answers, isolating the gate's effect. Line choices use exact old:N and new:N coordinates, including removed guards, unchanged context, and none. Repeated text at the wrong coordinate earns no credit. Localisation and mechanism credit require a flagged planted bug.

Why the decomposition might help

The risk stage may skip safe changes; line and mechanism answers expose vague detection. Any gain must offset an extra gate request because the detailed questions already fit one call. We measure recall, false alarms, early-gate misses, localisation, mechanisms, requests, tokens, latency, and cost. Code handles arithmetic and routing.

Data and changes from the offline pass

We retained the original 20 diffs in fixtures/diffs.jsonl and supplied missing context: optional full_name, positional-only usage for the private parameter rename, and intended retry caps and expiry skew. Other fixtures also state supported inputs. Reviewers now receive the contracts used for grading.

The original one-Noul baseline changed the task as well as its decomposition and could not score localisation. We replaced it with the staged arm's detailed request. Exact coordinates replaced added-line substring matching, admitting removed guards and context while preventing credit for repeated text or unflagged cases.

We score the eight new cases in fixtures/challenge.jsonl separately:

PairPlanted failureClean comparison
Deleted guardNegative totals silently accepted after validation is removedEquivalent integer-bound guard
Optional valuedict.get default does not replace explicit NoneFallback remains valid for absent, empty, and None values
Security changeor grants access across organizationsEquivalent negated conjunction

Deterministic tests exercise every defect and clean contract. The instrumented counter checks lock protection without reproducing a scheduler-dependent race on this interpreter. These small, authored, synthetic pairs provide no independent production sample. We did not tune thresholds on the results.

Verification
nice -n 10 python3 -m unittest discover -s tests -p test_review_triage.py -v

See VALIDATION.md for the RED/GREEN evidence and the limits of these checks.

Raw result files
openjev (Codiv free hosted) - NOT JEVresults/v2_codiv_full_r1.json
{
 "schema_version": 2,
 "project": "review-triage",
 "subset": "full",
 "repeat": 1,
 "timestamp": "2026-09-24T16:03:31Z",
 "input_hashes": {
  "samples/review-triage/fixtures/challenge.jsonl": "c44b7004f0069a5e99104d6df0adf48c0bf2e8454ca414313e958eb28edf9836",
  "samples/review-triage/fixtures/diffs.jsonl": "bf238e5943e7d134694f6b594c7f770bc44c355cff1d407b453777df46f742e4"
 },
 "source_hashes": {
  "samples/_shared/community.py": "88155b9fcb87dbc2af50a59a02cdcc85a87e9404debaa6e6e3a96744c104d2cd",
  "samples/_shared/jev.py": "6bd98ea32afdccc886148ee92aaa1213a4ed2c9d2c797739450a1a1856d16a06",
  "samples/review-triage/triage.py": "eec07c74b82554e3cb9514095ca2b21b141f5509fe985fd54d10846794b7cb2a"
 },
 "label": "openjev (Codiv free hosted) - NOT JEV",
 "mode": "live",
 "provider": "codiv",
 "resumed_after": "connection_error",
 "original_source_hashes": {
  "samples/_shared/community.py": "8e2eb5e4c8d57d6226d38bf0e218d3ba379f4a207ee3bbfad1e1670b8df50482",
  "samples/_shared/jev.py": "1e936cb0f83f45baa62e1a2417682e0630e52bf0143d1bb09608a05fcb6dbcc9",
  "samples/review-triage/triage.py": "eec07c74b82554e3cb9514095ca2b21b141f5509fe985fd54d10846794b7cb2a"
 },
 "rows": [
  {
   "risk": 3.9640427783161796,
   "area": "logic",
   "stage2": true,
   "flag": true,
   "line": "new:3",
   "mechanism": "off_by_one",
   "failure": 0.9999800626322678,
   "severity": 2.0534221771613783,
   "id": "b01_offbyone",
   "cohort": "original",
   "pipeline": "staged",
   "planted": true,
   "gold_lines": [
    "new:3"
   ],
   "gold_mechanism": "off_by_one",
   "line_ok": true,
   "mechanism_ok": true,
   "metrics": {
    "requests": 2,
    "completed_requests": 2,
    "input_tokens": 1359,
    "output_tokens": 0,
    "median_latency_ms": 552.0,
    "wall_s": 1.104,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 1
  },
  {
   "risk": null,
   "area": null,
   "stage2": true,
   "flag": true,
   "line": "new:3",
   "mechanism": "off_by_one",
   "failure": 0.9999722419133104,
   "severity": 2.047179784829845,
   "id": "b01_offbyone",
   "cohort": "original",
   "pipeline": "single",
   "planted": true,
   "gold_lines": [
    "new:3"
   ],
   "gold_mechanism": "off_by_one",
   "line_ok": true,
   "mechanism_ok": true,
   "metrics": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 890,
    "output_tokens": 0,
    "median_latency_ms": 378.6,
    "wall_s": 0.379,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 1
  },
  {
   "risk": 3.9950653331890247,
   "area": "logic",
   "stage2": true,
   "flag": true,
   "line": "new:2",
   "mechanism": "inverted_condition",
   "failure": 0.9999544670142974,
   "severity": 3.7378411405181673,
   "id": "b02_inverted",
   "cohort": "original",
   "pipeline": "staged",
   "planted": true,
   "gold_lines": [
    "new:2"
   

… truncated, 3000 of 175099 bytes shown. Complete safe artifact

openjev (Codiv free hosted) - NOT JEVresults/v2_codiv_full_r2.json
{
 "schema_version": 2,
 "project": "review-triage",
 "subset": "full",
 "repeat": 2,
 "timestamp": "2026-09-24T16:25:30Z",
 "input_hashes": {
  "samples/review-triage/fixtures/challenge.jsonl": "c44b7004f0069a5e99104d6df0adf48c0bf2e8454ca414313e958eb28edf9836",
  "samples/review-triage/fixtures/diffs.jsonl": "bf238e5943e7d134694f6b594c7f770bc44c355cff1d407b453777df46f742e4"
 },
 "source_hashes": {
  "samples/_shared/community.py": "88155b9fcb87dbc2af50a59a02cdcc85a87e9404debaa6e6e3a96744c104d2cd",
  "samples/_shared/jev.py": "6bd98ea32afdccc886148ee92aaa1213a4ed2c9d2c797739450a1a1856d16a06",
  "samples/review-triage/triage.py": "eec07c74b82554e3cb9514095ca2b21b141f5509fe985fd54d10846794b7cb2a"
 },
 "label": "openjev (Codiv free hosted) - NOT JEV",
 "mode": "live",
 "provider": "codiv",
 "rows": [
  {
   "risk": 3.9715573881452726,
   "area": "logic",
   "stage2": true,
   "flag": true,
   "line": "new:3",
   "mechanism": "off_by_one",
   "failure": 0.9999428880257052,
   "severity": 2.0477315246346186,
   "id": "b01_offbyone",
   "cohort": "original",
   "pipeline": "staged",
   "planted": true,
   "gold_lines": [
    "new:3"
   ],
   "gold_mechanism": "off_by_one",
   "line_ok": true,
   "mechanism_ok": true,
   "metrics": {
    "requests": 2,
    "completed_requests": 2,
    "input_tokens": 1359,
    "output_tokens": 0,
    "median_latency_ms": 1244.0,
    "wall_s": 3.087,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 2
  },
  {
   "risk": null,
   "area": null,
   "stage2": true,
   "flag": true,
   "line": "new:3",
   "mechanism": "off_by_one",
   "failure": 0.9999665622271429,
   "severity": 2.0432567476475487,
   "id": "b01_offbyone",
   "cohort": "original",
   "pipeline": "single",
   "planted": true,
   "gold_lines": [
    "new:3"
   ],
   "gold_mechanism": "off_by_one",
   "line_ok": true,
   "mechanism_ok": true,
   "metrics": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 890,
    "output_tokens": 0,
    "median_latency_ms": 678.9,
    "wall_s": 1.59,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 2
  },
  {
   "risk": 3.994360790039272,
   "area": "logic",
   "stage2": true,
   "flag": true,
   "line": "new:2",
   "mechanism": "inverted_condition",
   "failure": 0.9999320402589706,
   "severity": 3.716427042102948,
   "id": "b02_inverted",
   "cohort": "original",
   "pipeline": "staged",
   "planted": true,
   "gold_lines": [
    "new:2"
   ],
   "gold_mechanism": "inverted_condition",
   "line_ok": true,
   "mechanism_ok": true,
   "metrics": {
    "requests": 2,
    "completed_requests": 2,
    "input_tokens": 1312,
    "output_tokens": 0,
    "median_latency_ms": 501.0,
    "wall_s": 3.939,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    

… truncated, 2988 of 173933 bytes shown. Complete safe artifact

openjev (Codiv free hosted) - NOT JEVresults/v2_codiv_full_r3.json
{
 "schema_version": 2,
 "project": "review-triage",
 "subset": "full",
 "repeat": 3,
 "timestamp": "2026-09-24T16:33:43Z",
 "input_hashes": {
  "samples/review-triage/fixtures/challenge.jsonl": "c44b7004f0069a5e99104d6df0adf48c0bf2e8454ca414313e958eb28edf9836",
  "samples/review-triage/fixtures/diffs.jsonl": "bf238e5943e7d134694f6b594c7f770bc44c355cff1d407b453777df46f742e4"
 },
 "source_hashes": {
  "samples/_shared/community.py": "88155b9fcb87dbc2af50a59a02cdcc85a87e9404debaa6e6e3a96744c104d2cd",
  "samples/_shared/jev.py": "5a0a37f5097a0ba9c7ccfc4a75990c7d825c809e48f2e407341526b6b8bbcc28",
  "samples/review-triage/triage.py": "eec07c74b82554e3cb9514095ca2b21b141f5509fe985fd54d10846794b7cb2a"
 },
 "label": "openjev (Codiv free hosted) - NOT JEV",
 "mode": "live",
 "provider": "codiv",
 "resumed_after": "connection_error",
 "original_source_hashes": {
  "samples/_shared/community.py": "88155b9fcb87dbc2af50a59a02cdcc85a87e9404debaa6e6e3a96744c104d2cd",
  "samples/_shared/jev.py": "5a0a37f5097a0ba9c7ccfc4a75990c7d825c809e48f2e407341526b6b8bbcc28",
  "samples/review-triage/triage.py": "eec07c74b82554e3cb9514095ca2b21b141f5509fe985fd54d10846794b7cb2a"
 },
 "rows": [
  {
   "risk": 3.961079523839213,
   "area": "logic",
   "stage2": true,
   "flag": true,
   "line": "new:3",
   "mechanism": "off_by_one",
   "failure": 0.9999640377260781,
   "severity": 2.0530663249791123,
   "id": "b01_offbyone",
   "cohort": "original",
   "pipeline": "staged",
   "planted": true,
   "gold_lines": [
    "new:3"
   ],
   "gold_mechanism": "off_by_one",
   "line_ok": true,
   "mechanism_ok": true,
   "metrics": {
    "requests": 2,
    "completed_requests": 2,
    "input_tokens": 1359,
    "output_tokens": 0,
    "median_latency_ms": 454.9,
    "wall_s": 4.145,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 3
  },
  {
   "risk": null,
   "area": null,
   "stage2": true,
   "flag": true,
   "line": "new:3",
   "mechanism": "off_by_one",
   "failure": 0.9999869985389485,
   "severity": 2.050946068976244,
   "id": "b01_offbyone",
   "cohort": "original",
   "pipeline": "single",
   "planted": true,
   "gold_lines": [
    "new:3"
   ],
   "gold_mechanism": "off_by_one",
   "line_ok": true,
   "mechanism_ok": true,
   "metrics": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 890,
    "output_tokens": 0,
    "median_latency_ms": 509.2,
    "wall_s": 1.997,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 3
  },
  {
   "risk": 3.995061673865197,
   "area": "logic",
   "stage2": true,
   "flag": true,
   "line": "new:2",
   "mechanism": "inverted_condition",
   "failure": 0.9999624638652352,
   "severity": 3.787647200262909,
   "id": "b02_inverted",
   "cohort": "original",
   "pipeline": "staged",
   "planted": true,
   "gold_lines": [
    "new:2"
   ],

… truncated, 2999 of 175130 bytes shown. Complete safe artifact

REAL JEV (OpenCode Zen free, jev-1.13-free)results/v2_zen_zen_r1.json
{
 "schema_version": 2,
 "project": "review-triage",
 "subset": "zen",
 "repeat": 1,
 "timestamp": "2026-09-24T16:09:05Z",
 "input_hashes": {
  "samples/review-triage/fixtures/challenge.jsonl": "c44b7004f0069a5e99104d6df0adf48c0bf2e8454ca414313e958eb28edf9836",
  "samples/review-triage/fixtures/diffs.jsonl": "bf238e5943e7d134694f6b594c7f770bc44c355cff1d407b453777df46f742e4"
 },
 "source_hashes": {
  "samples/_shared/community.py": "88155b9fcb87dbc2af50a59a02cdcc85a87e9404debaa6e6e3a96744c104d2cd",
  "samples/_shared/jev.py": "6bd98ea32afdccc886148ee92aaa1213a4ed2c9d2c797739450a1a1856d16a06",
  "samples/review-triage/triage.py": "eec07c74b82554e3cb9514095ca2b21b141f5509fe985fd54d10846794b7cb2a"
 },
 "label": "REAL JEV (OpenCode Zen free, jev-1.13-free)",
 "mode": "live",
 "provider": "zen",
 "rows": [
  {
   "risk": 3.68,
   "area": "logic",
   "stage2": true,
   "flag": true,
   "line": "new:3",
   "mechanism": "off_by_one",
   "failure": 0.7,
   "severity": 2.06,
   "id": "b01_offbyone",
   "cohort": "original",
   "pipeline": "staged",
   "planted": true,
   "gold_lines": [
    "new:3"
   ],
   "gold_mechanism": "off_by_one",
   "line_ok": true,
   "mechanism_ok": true,
   "metrics": {
    "requests": 2,
    "completed_requests": 2,
    "input_tokens": 1863,
    "output_tokens": 314,
    "median_latency_ms": 1257.8,
    "wall_s": 58.315,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 1
  },
  {
   "risk": null,
   "area": null,
   "stage2": true,
   "flag": true,
   "line": "new:3",
   "mechanism": "off_by_one",
   "failure": 0.72,
   "severity": 2.1,
   "id": "b01_offbyone",
   "cohort": "original",
   "pipeline": "single",
   "planted": true,
   "gold_lines": [
    "new:3"
   ],
   "gold_mechanism": "off_by_one",
   "line_ok": true,
   "mechanism_ok": true,
   "metrics": {
    "requests": 1,
    "completed_requests": 1,
    "input_tokens": 1178,
    "output_tokens": 245,
    "median_latency_ms": 560.9,
    "wall_s": 31.104,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 1
  },
  {
   "risk": 3.63,
   "area": "security",
   "stage2": true,
   "flag": true,
   "line": "new:3",
   "mechanism": "injection",
   "failure": 0.79,
   "severity": 3.87,
   "id": "b03_sqlinject",
   "cohort": "original",
   "pipeline": "staged",
   "planted": true,
   "gold_lines": [
    "new:3"
   ],
   "gold_mechanism": "injection",
   "line_ok": true,
   "mechanism_ok": true,
   "metrics": {
    "requests": 2,
    "completed_requests": 2,
    "input_tokens": 1835,
    "output_tokens": 287,
    "median_latency_ms": 457.6,
    "wall_s": 62.033,
    "cost_usd": 0.0,
    "cost_basis": "free route; missing provider cost fields are not receipts",
    "reported_cost_usd": null
   },
   "repeat": 1
  },
  {
   "risk": null,
   "area": null,
   "stage2": true,
   "flag": true

… truncated, 3000 of 41489 bytes shown. Complete safe artifact

Run it
JEV_PROVIDER=codiv JEV_KEY_FILE="$CODIV_KEY_FILE" JEV_MAX_USD=0 ./samples/review-triage/run.sh --repeat-start 4
./samples/review-triage/run.sh --offline --repeat-start 4
./samples/review-triage/run.sh --mock --repeat-start 4
JEV_PROVIDER=zen JEV_KEY_FILE="$ZEN_KEY_FILE" JEV_MAX_USD=0 JEV_MIN_INTERVAL_S=20 ./samples/review-triage/run.sh --subset zen --repeat-start 4