Jev Labs

Compaction Check

Keep the evidence when context shrinks.

Headline result

REAL JEV (OpenCode Zen free, jev-1.13-free)

model: 0.2% estimated tokens saved; evidence 17/17; reader 17/17

openjev (Codiv free hosted) - NOT JEV

model: 8.6% estimated tokens saved; evidence 14/17; reader 11/17

Corrected offline baseline

Source: samples/compaction-check/results/v2_code_full_r2.json.

Code-only context policies, corrected pass
MethodEstimated tokensSavedEvidence retained
full11,2680.0%17/17
last68,95420.5%11/17
trunc1209,65714.3%15/17

v2_code_full_r1.json: archived / non-reproducible (its pinned input was never committed). It does not support current baseline claims.

Details · results, methods & full write-up
← All projects

Project 11

Compaction Check

drop stale tool output instead of summarising

0.2%model: estimated tokens saved; evidence 17/17; r…REAL JEV
openjev (Codiv free hosted) - NOT JEVREAL JEV (OpenCode Zen free, jev-1.13-free)

Tool output competes with useful context in long agent runs. Removing an old file listing can save space; removing the only failed-check result can hide what happened. Summarisation adds another failure mode by rewriting facts.

Headline result

REAL JEV (OpenCode Zen free, jev-1.13-free)

model: 0.2% estimated tokens saved; evidence 17/17; reader 17/17

openjev (Codiv free hosted) - NOT JEV

model: 8.6% estimated tokens saved; evidence 14/17; reader 11/17

Details

The problem

Tool output competes with useful context in long agent runs. Removing an old file listing can save space; removing the only failed-check result can hide what happened. Summarisation adds another failure mode by rewriting facts. This experiment preserves every user message, assistant message, and tool call, and changes only tool-result text.

We check whether readers can still retrieve recorded errors, file sizes, image measurements, and validation outcomes. Correct answers without surviving evidence are counted separately as unsupported guesses.

How Jev is used

One typed Choice per visible tool result asks keep or drop, batched into one request per excerpt. The selector sees all committed text, including surrounding messages, but receives no evaluation questions, gold answers, or evidence annotations.

Python keeps selected output verbatim and replaces dropped output with [tool output dropped by compaction]. Separate batched Choice requests ask the 17 retrieval questions under each policy. Readers must use the excerpt and choose insufficient_evidence when support is missing. Python applies policies, counts, checks evidence, and scores answers.

Typed decisions expose what the selector drops; code checks what survives. The reader then tests whether it can find those facts, including cases where saving space makes the run harder to review.

What real Jev would change

The completed REAL JEV (OpenCode Zen free, jev-1.13-free) pass used the same frozen questions and policies and produced a different tradeoff: all 17 answers and supporting facts survived model selection, while context savings fell from Codiv's 8.6% to 0.2%. The added selector usage exceeded the one-read saving. The cross-route comparison uses the first pass from each route. The additional Codiv replays do not establish calibration or transfer to longer runs, and Zen still has only one pass.

Code preserves verbatim text, calls, and messages and checks evidence on either route. Measuring unabridged runs would require new data, a privacy review, and a separate context-budget study.

Results

Measured code baselines, version compaction-v2; versioned artifact:

MethodContext tokens, estimatedSavedTool evidencePermanent controlsAll evidence
Full context11,2680.0%11/116/617/17
Last six results8,95420.5%5/116/611/17
First 120 chars of older results9,65714.3%9/116/615/17

The last-six policy loses the missing cairosvg error and recorded ds7.jpg size from the Opus excerpt; it also loses the dsh search refusal, unreadable image, measured aspect ratio, and pre-replacement file size. Truncation loses the two Opus facts that occur beyond the first 120 characters. The Codex excerpt retains all five annotated answers under both code policies.

The first openjev (Codiv free hosted) - NOT JEV pass completed all 15 calls, with returned model openjev-0.1:

MethodReader correctEvidence retainedContext savedReported input tokens, including selection
full15/1717/170.0%21,528
last610/1711/1720.5%16,024
trunc12015/1715/1714.3%17,755
model keep/drop11/1714/178.6%62,875

The model selector dropped three supporting facts and the reader lost four answers it had previously answered correctly. Trunc120 matched full-context aggregate accuracy but included one correct unsupported guess and lost one previously correct answer. Even the full reader missed the visible continuation summary and confused the pre-replacement SVG size with a later size. Follow-up replays below measure variation beyond this first pass.

Total provider-reported usage was 118,182 input tokens and 0 output tokens; median HTTP latency 1,430.1 ms; full wall time 32.776 s. The approved free-route accounting cost was $0, while provider monetary cost fields were absent. RESULTS.md has strata, individual failures, abstentions, selector overhead, and per-method timings; the versioned artifact records every typed answer.

Three openjev (Codiv free hosted) - NOT JEV repeats used the same excerpts. Entries follow repeat 1, repeat 2, repeat 3, each scored on the same 17 questions.

MethodReader correct out of 17 (r1, r2, r3)Evidence retained out of 17 (r1, r2, r3)Context saved (r1, r2, r3)
full15, 15, 1517, 17, 170.0%, 0.0%, 0.0%
last610, 10, 1111, 11, 1120.5%, 20.5%, 20.5%
trunc12015, 15, 1515, 15, 1514.3%, 14.3%, 14.3%
model keep/drop11, 12, 1514, 14, 168.6%, 6.5%, 5.8%

Model repeat 3 retained and recovered dsh's search-error and final-size facts, while every repeat still dropped the Opus final-size evidence and answered that question incorrectly. Its improved quality came with lower savings. Selector input usage stayed at 43,663 tokens per pass; selector plus reader input usage was 62,875, 63,506, and 63,552, versus 21,528 for full-context readers each time. All three passes used 45 requests, 355,854 reported input tokens, and 100.354 s summed wall time, on the $0 free-route accounting basis with no monetary receipt.

The repeats reuse transcript data; treating them as 51 independent questions would overstate the sample. Fixture, evaluator, runner, and question hashes are unchanged. A legacy non-strict user-agent URL update changed the client file hash while leaving strict evaluation unchanged. RESULTS.md records failures, usage, and the disk/imported-client provenance limitation.

The REAL JEV (OpenCode Zen free, jev-1.13-free) pass also completed all 15 calls on the same three excerpts and 17 questions, returning jev-1.13-free:

MethodReader correctEvidence retainedContext savedReported input tokens, including selectionReported output tokens
full17/1717/170.0%20,9911,022
last612/1711/1720.5%15,8681,039
trunc12016/1715/1714.3%17,4421,023
model keep/drop17/1717/170.2%46,7563,189

Real Jev's full and model readers answered all 17 correctly, compared with 15 and 11 for Codiv. Its model selector preserved all 17 facts but saved only 24 estimated tokens (0.2%). Opus grew by 57 estimated tokens because dropping short image markers inserted a longer stub; Codex saved 81, and dsh kept every result. Under last6 and trunc120, the reader answered the missing cairosvg fact without annotated support, so their factual scores exceed their retained-evidence counts.

Total provider-reported usage was 101,057 input and 6,273 output tokens; median HTTP latency 530.0 ms; full CLI wall time 300.070 s, including serial pacing, locking, and cooldowns. The approved free-route accounting cost was $0 with no provider monetary receipt. The Zen artifact and RESULTS.md preserve per-question outcomes and selector overhead. Zen has one pass; the three Codiv repeats above measure variation only on these same excerpts.

The [design notes](registered protocol) record the Codex plan-mode Q&A.

Data and changes from the offline pass

The three excerpts with questions in fixtures/sessions.jsonl remain Claude Code + Opus 5.5, Codex CLI + GPT-6 Astra, and dsh + DeepSeek Flash. The fourth excerpt has no tool results or compaction questions. Evaluation reads no raw transcripts and never runs the historical importer, build_excerpts.py.

The original import kept the first 18 and last 12 tool steps of longer runs, clipped content, replaced images with markers, and scrubbed sensitive text. Original tool IDs and omitted history are unavailable. Savings cover these shortened excerpts only; complete runs were not measured.

The first offline pass is preserved in results/code_code.json as historical v1 output. The current version changes:

  • Every result is enumerated. The old matcher silently omitted parallel results. All results now remain, with eight ambiguous associations marked unknown. Interleaved calls stay grouped until all results return; original IDs remain unknown.
  • The selector now sees all committed text instead of 240-character previews and chooses only keep/drop. Truncation remains a separate code baseline.
  • The image-count question now retrieves the recorded size of ds7.jpg. Search and build-file questions ask what commands request or reference, since call text cannot establish success. Final-size questions use the final directory listings.
  • Evidence requires exact supporting spans in their annotated events, including every part of multi-part evidence. Eleven tool-dependent questions are scored separately from six permanent-text controls.
  • Readers can abstain with insufficient_evidence. Reading full context establishes each reader's baseline for lost answers. Calls are uncached across methods and repeats.
  • Context size uses ceil(characters / 4); provider usage and selector overhead are separate. The old serialized-string estimate makes v1 token totals unsuitable for direct comparison.
Metrics, cost, tokens, latency

Results record context characters and token estimates, retention by question and stratum, reader accuracy, abstentions, unsupported correct answers, and answers lost against full context. Answerability-aware scoring expects the gold answer when evidence survives and insufficient_evidence otherwise.

Provider usage, request counts, HTTP latency, wall time, and cost provenance are separate. HTTP latency includes the network round trip and excludes locking and cooldown; backend inference time is not isolated. Call and run elapsed times include their waiting and pacing. Missing usage stays unknown. Model-policy usage includes selector overhead even when reader context shrinks. Estimated reuse break-even is ceil(selector request-and-answer token estimate / per-read context token estimate saved), or null without savings. It approximates token accounting and guarantees neither cost nor quality.

All code-only policies cost $0 and use no model tokens. The measured local baseline pass took 0.013 seconds. Mock usage estimates do not enter measured hosted tables.

Limits and verification

This small, authored benchmark has one annotation pass and measures recorded-excerpt replay. Live-agent quality and behavior after compaction remain untested. Scrubbing and clipping remove information. Images are unavailable, and uncertain call associations remain unknown. Annotations make support auditable but may omit valid paraphrases or inferences.

python3 -m unittest discover -s tests -p test_compaction_check.py from the repository root checks parallel/orphan results, immutable compaction, exact evidence spans, abstention scoring, selector isolation, fixture validity, and the 0/15-call budgets.

Raw result files
CODE ONLY - NO MODELresults/code_code.json
{
 "label": "code only (no model)",
 "mode": "code",
 "provider": null,
 "ts": "2026-09-24T15:19:18Z",
 "wall_s": 0.0,
 "calls": [],
 "rows": [
  {
   "session": "claude-code__opus-5-5",
   "method": "full",
   "tokens": 4133,
   "tokens_full": 4133,
   "decisions": {
    "keep": 0,
    "truncate": 0,
    "drop": 0
   },
   "retained": 6,
   "questions": 6,
   "qa_correct": null
  },
  {
   "session": "claude-code__opus-5-5",
   "method": "last6",
   "tokens": 3089,
   "tokens_full": 4133,
   "decisions": {
    "keep": 0,
    "truncate": 0,
    "drop": 23
   },
   "retained": 4,
   "questions": 6,
   "qa_correct": null
  },
  {
   "session": "claude-code__opus-5-5",
   "method": "trunc120",
   "tokens": 3239,
   "tokens_full": 4133,
   "decisions": {
    "keep": 0,
    "truncate": 23,
    "drop": 0
   },
   "retained": 4,
   "questions": 6,
   "qa_correct": null
  },
  {
   "session": "codex__gpt-6-astra",
   "method": "full",
   "tokens": 2368,
   "tokens_full": 2368,
   "decisions": {
    "keep": 0,
    "truncate": 0,
    "drop": 0
   },
   "retained": 5,
   "questions": 5,
   "qa_correct": null
  },
  {
   "session": "codex__gpt-6-astra",
   "method": "last6",
   "tokens": 2162,
   "tokens_full": 2368,
   "decisions": {
    "keep": 0,
    "truncate": 0,
    "drop": 4
   },
   "retained": 5,
   "questions": 5,
   "qa_correct": null
  },
  {
   "session": "codex__gpt-6-astra",
   "method": "trunc120",
   "tokens": 2262,
   "tokens_full": 2368,
   "decisions": {
    "keep": 0,
    "truncate": 4,
    "drop": 0
   },
   "retained": 5,
   "questions": 5,
   "qa_correct": null
  },
  {
   "session": "dsh__deepseek-flash",
   "method": "full",
   "tokens": 5114,
   "tokens_full": 5114,
   "decisions": {
    "keep": 0,
    "truncate": 0,
    "drop": 0
   },
   "retained": 6,
   "questions": 6,
   "qa_correct": null
  },
  {
   "session": "dsh__deepseek-flash",
   "method": "last6",
   "tokens": 4022,
   "tokens_full": 5114,
   "decisions": {
    "keep": 0,
    "truncate": 0,
    "drop": 21
   },
   "retained": 2,
   "questions": 6,
   "qa_correct": null
  },
  {
   "session": "dsh__deepseek-flash",
   "method": "trunc120",
   "tokens": 4448,
   "tokens_full": 5114,
   "decisions": {
    "keep": 0,
    "truncate": 21,
    "drop": 0
   },
   "retained": 6,
   "questions": 6,
   "qa_correct": null
  }
 ],
 "per_question": [
  {
   "session": "claude-code__opus-5-5",
   "method": "full",
   "q": "cc1",
   "retained": true,
   "answer": null,
   "gold": "no",
   "correct": null
  },
  {
   "session": "claude-code__opus-5-5",
   "method": "full",
   "q": "cc2",
   "retained": true,
   "answer": null,
   "gold": "commons",
   "correct": null
  },
  {
   "session": "claude-code__opus-5-5",
   "method": "full",
   "q": "cc3",
   "retained": true,
   "answer": null,
   "gold": "7",
   "correct": null
  },
  {
   "session": "claude-code__opus-5-5",
   "method": "full",
   "q": "cc4",
   "retained": true,
   "answer": null,
   "gold": "95586",
   "correct": null

… truncated, 2998 of 10683 bytes shown. Complete safe artifact

CODE ONLY - NO MODELresults/v2_code_full_r2.json
{
 "schema_version": 2,
 "project": "compaction-check",
 "subset": "full",
 "repeat": 2,
 "timestamp": "2026-09-24T16:03:03Z",
 "input_hashes": {
  "samples/compaction-check/fixtures/questions.json": "0d348f75b58e9df36b14babdfe9ae57f192a70d2452d757385839706712f00d1",
  "samples/compaction-check/fixtures/sessions.jsonl": "6b6e01855f16358971ca6734810ebec8d3a7b2240d9a8c05dbfd7c3f0ca115e6"
 },
 "source_hashes": {
  "samples/_shared/jev.py": "1e936cb0f83f45baa62e1a2417682e0630e52bf0143d1bb09608a05fcb6dbcc9",
  "samples/compaction-check/../_shared/community.py": "138e743ef9d0382da0d823cbe18ec9cf4ca3018122f10da809ff05d883322016",
  "samples/compaction-check/build_excerpts.py": "28a991f6ddcd8324b2336543bfde6f3d3ba379dd78c45a876d24d93d54d1e981",
  "samples/compaction-check/compact.py": "80df701cac548bf404ad02794635baab5bb7cf92b2b9e7acd92c263b6c9fd5ec"
 },
 "label": "code only (no model)",
 "mode": "code",
 "provider": null,
 "rows": [
  {
   "session": "claude-code__opus-5-5",
   "case_id": "claude-code__opus-5-5",
   "repeat": 2,
   "method": "full",
   "context_chars": 16157,
   "tokens_estimate": 4040,
   "tokens_full_estimate": 4040,
   "tokens_saved_estimate": 0,
   "saved_fraction": 0.0,
   "decisions": {
    "keep": 30,
    "truncate": 0,
    "drop": 0
   },
   "result_decisions": {
    "1": "keep",
    "3": "keep",
    "5": "keep",
    "7": "keep",
    "9": "keep",
    "12": "keep",
    "14": "keep",
    "16": "keep",
    "18": "keep",
    "21": "keep",
    "23": "keep",
    "25": "keep",
    "27": "keep",
    "30": "keep",
    "31": "keep",
    "33": "keep",
    "35": "keep",
    "37": "keep",
    "41": "keep",
    "43": "keep",
    "45": "keep",
    "47": "keep",
    "49": "keep",
    "51": "keep",
    "53": "keep",
    "56": "keep",
    "58": "keep",
    "60": "keep",
    "62": "keep",
    "64": "keep"
   },
   "unknown_result_associations": 2,
   "questions": 6,
   "answered": 0,
   "retained": 6,
   "correct": null,
   "abstained": null,
   "answerability_correct": null,
   "unsupported_correct_guess": null,
   "lost_answer": null,
   "qa_available": false,
   "strata": {
    "tool_dependent": {
     "questions": 4,
     "answered": 0,
     "retained": 4,
     "correct": null,
     "abstained": null,
     "answerability_correct": null,
     "unsupported_correct_guess": null,
     "lost_answer": null
    },
    "permanent_control": {
     "questions": 2,
     "answered": 0,
     "retained": 2,
     "correct": null,
     "abstained": null,
     "answerability_correct": null,
     "unsupported_correct_guess": null,
     "lost_answer": null
    }
   },
   "per_question": [
    {
     "question": "cc1",
     "kind": "tool_dependent",
     "retained": true,
     "gold": "no",
     "answer": null,
     "correct": null,
     "abstained": null,
     "answerability_correct": null,
     "unsupported_correct_guess": null,
     "lost_answer": null
    },
    {
     "question": "cc2",
     "kind": "permanent_control",
     "retained": true,
     "gold": 

… truncated, 2998 of 37088 bytes shown. Complete safe artifact

CODE ONLY - NO MODELresults/v2_code_full_r3.json
{
 "schema_version": 2,
 "project": "compaction-check",
 "subset": "full",
 "repeat": 3,
 "limit": null,
 "campaign": null,
 "timestamp": "2026-09-25T11:44:40Z",
 "input_hashes": {
  "samples/compaction-check/fixtures/questions.json": "0d348f75b58e9df36b14babdfe9ae57f192a70d2452d757385839706712f00d1",
  "samples/compaction-check/fixtures/sessions.jsonl": "6b6e01855f16358971ca6734810ebec8d3a7b2240d9a8c05dbfd7c3f0ca115e6"
 },
 "source_hashes": {
  "samples/_shared/community.py": "79d28fe055b35dada9eeb70122da25ae3487e901c0d0992ba1270207ad553968",
  "samples/_shared/jev.py": "a20ab00f05c9fc838194f68b33f88b64bca3a0dd7e1ddf78d4c347bf9cdf3322",
  "samples/compaction-check/build_excerpts.py": "28a991f6ddcd8324b2336543bfde6f3d3ba379dd78c45a876d24d93d54d1e981",
  "samples/compaction-check/compact.py": "80df701cac548bf404ad02794635baab5bb7cf92b2b9e7acd92c263b6c9fd5ec"
 },
 "label": "code only (no model)",
 "mode": "code",
 "provider": null,
 "call_cap": null,
 "rows": [
  {
   "session": "claude-code__opus-5-5",
   "case_id": "claude-code__opus-5-5",
   "repeat": 3,
   "method": "full",
   "context_chars": 16157,
   "tokens_estimate": 4040,
   "tokens_full_estimate": 4040,
   "tokens_saved_estimate": 0,
   "saved_fraction": 0.0,
   "decisions": {
    "keep": 30,
    "truncate": 0,
    "drop": 0
   },
   "result_decisions": {
    "1": "keep",
    "3": "keep",
    "5": "keep",
    "7": "keep",
    "9": "keep",
    "12": "keep",
    "14": "keep",
    "16": "keep",
    "18": "keep",
    "21": "keep",
    "23": "keep",
    "25": "keep",
    "27": "keep",
    "30": "keep",
    "31": "keep",
    "33": "keep",
    "35": "keep",
    "37": "keep",
    "41": "keep",
    "43": "keep",
    "45": "keep",
    "47": "keep",
    "49": "keep",
    "51": "keep",
    "53": "keep",
    "56": "keep",
    "58": "keep",
    "60": "keep",
    "62": "keep",
    "64": "keep"
   },
   "unknown_result_associations": 2,
   "questions": 6,
   "answered": 0,
   "retained": 6,
   "correct": null,
   "abstained": null,
   "answerability_correct": null,
   "unsupported_correct_guess": null,
   "lost_answer": null,
   "qa_available": false,
   "strata": {
    "tool_dependent": {
     "questions": 4,
     "answered": 0,
     "retained": 4,
     "correct": null,
     "abstained": null,
     "answerability_correct": null,
     "unsupported_correct_guess": null,
     "lost_answer": null
    },
    "permanent_control": {
     "questions": 2,
     "answered": 0,
     "retained": 2,
     "correct": null,
     "abstained": null,
     "answerability_correct": null,
     "unsupported_correct_guess": null,
     "lost_answer": null
    }
   },
   "per_question": [
    {
     "question": "cc1",
     "kind": "tool_dependent",
     "retained": true,
     "gold": "no",
     "answer": null,
     "correct": null,
     "abstained": null,
     "answerability_correct": null,
     "unsupported_correct_guess": null,
     "lost_answer": null
    },
    {
     "question": "cc2",
     "kind": "permanent_control",

… truncated, 2995 of 37122 bytes shown. Complete safe artifact

openjev (Codiv free hosted) - NOT JEVresults/v2_codiv_full_r1.json
{
 "schema_version": 2,
 "project": "compaction-check",
 "subset": "full",
 "repeat": 1,
 "timestamp": "2026-09-24T16:06:10Z",
 "input_hashes": {
  "samples/compaction-check/fixtures/questions.json": "0d348f75b58e9df36b14babdfe9ae57f192a70d2452d757385839706712f00d1",
  "samples/compaction-check/fixtures/sessions.jsonl": "6b6e01855f16358971ca6734810ebec8d3a7b2240d9a8c05dbfd7c3f0ca115e6"
 },
 "source_hashes": {
  "samples/_shared/community.py": "88155b9fcb87dbc2af50a59a02cdcc85a87e9404debaa6e6e3a96744c104d2cd",
  "samples/_shared/jev.py": "6bd98ea32afdccc886148ee92aaa1213a4ed2c9d2c797739450a1a1856d16a06",
  "samples/compaction-check/build_excerpts.py": "28a991f6ddcd8324b2336543bfde6f3d3ba379dd78c45a876d24d93d54d1e981",
  "samples/compaction-check/compact.py": "80df701cac548bf404ad02794635baab5bb7cf92b2b9e7acd92c263b6c9fd5ec"
 },
 "label": "openjev (Codiv free hosted) - NOT JEV",
 "mode": "live",
 "provider": "codiv",
 "rows": [
  {
   "session": "claude-code__opus-5-5",
   "case_id": "claude-code__opus-5-5",
   "repeat": 1,
   "method": "full",
   "context_chars": 16157,
   "tokens_estimate": 4040,
   "tokens_full_estimate": 4040,
   "tokens_saved_estimate": 0,
   "saved_fraction": 0.0,
   "decisions": {
    "keep": 30,
    "truncate": 0,
    "drop": 0
   },
   "result_decisions": {
    "1": "keep",
    "3": "keep",
    "5": "keep",
    "7": "keep",
    "9": "keep",
    "12": "keep",
    "14": "keep",
    "16": "keep",
    "18": "keep",
    "21": "keep",
    "23": "keep",
    "25": "keep",
    "27": "keep",
    "30": "keep",
    "31": "keep",
    "33": "keep",
    "35": "keep",
    "37": "keep",
    "41": "keep",
    "43": "keep",
    "45": "keep",
    "47": "keep",
    "49": "keep",
    "51": "keep",
    "53": "keep",
    "56": "keep",
    "58": "keep",
    "60": "keep",
    "62": "keep",
    "64": "keep"
   },
   "unknown_result_associations": 2,
   "questions": 6,
   "answered": 6,
   "retained": 6,
   "correct": 5,
   "abstained": 0,
   "answerability_correct": 5,
   "unsupported_correct_guess": 0,
   "lost_answer": 0,
   "qa_available": true,
   "strata": {
    "tool_dependent": {
     "questions": 4,
     "answered": 4,
     "retained": 4,
     "correct": 4,
     "abstained": 0,
     "answerability_correct": 4,
     "unsupported_correct_guess": 0,
     "lost_answer": 0
    },
    "permanent_control": {
     "questions": 2,
     "answered": 2,
     "retained": 2,
     "correct": 1,
     "abstained": 0,
     "answerability_correct": 1,
     "unsupported_correct_guess": 0,
     "lost_answer": 0
    }
   },
   "per_question": [
    {
     "question": "cc1",
     "kind": "tool_dependent",
     "retained": true,
     "gold": "no",
     "answer": "no",
     "correct": true,
     "abstained": false,
     "answerability_correct": true,
     "unsupported_correct_guess": false,
     "lost_answer": false
    },
    {
     "question": "cc2",
     "kind": "permanent_control",
     "retained": true,
     "gold": "commons",
     "answer": "commons",
     

… truncated, 2997 of 96962 bytes shown. Complete safe artifact

openjev (Codiv free hosted) - NOT JEVresults/v2_codiv_full_r2.json
{
 "schema_version": 2,
 "project": "compaction-check",
 "subset": "full",
 "repeat": 2,
 "timestamp": "2026-09-24T16:27:52Z",
 "input_hashes": {
  "samples/compaction-check/fixtures/questions.json": "0d348f75b58e9df36b14babdfe9ae57f192a70d2452d757385839706712f00d1",
  "samples/compaction-check/fixtures/sessions.jsonl": "6b6e01855f16358971ca6734810ebec8d3a7b2240d9a8c05dbfd7c3f0ca115e6"
 },
 "source_hashes": {
  "samples/_shared/community.py": "88155b9fcb87dbc2af50a59a02cdcc85a87e9404debaa6e6e3a96744c104d2cd",
  "samples/_shared/jev.py": "5a0a37f5097a0ba9c7ccfc4a75990c7d825c809e48f2e407341526b6b8bbcc28",
  "samples/compaction-check/build_excerpts.py": "28a991f6ddcd8324b2336543bfde6f3d3ba379dd78c45a876d24d93d54d1e981",
  "samples/compaction-check/compact.py": "80df701cac548bf404ad02794635baab5bb7cf92b2b9e7acd92c263b6c9fd5ec"
 },
 "label": "openjev (Codiv free hosted) - NOT JEV",
 "mode": "live",
 "provider": "codiv",
 "rows": [
  {
   "session": "claude-code__opus-5-5",
   "case_id": "claude-code__opus-5-5",
   "repeat": 2,
   "method": "full",
   "context_chars": 16157,
   "tokens_estimate": 4040,
   "tokens_full_estimate": 4040,
   "tokens_saved_estimate": 0,
   "saved_fraction": 0.0,
   "decisions": {
    "keep": 30,
    "truncate": 0,
    "drop": 0
   },
   "result_decisions": {
    "1": "keep",
    "3": "keep",
    "5": "keep",
    "7": "keep",
    "9": "keep",
    "12": "keep",
    "14": "keep",
    "16": "keep",
    "18": "keep",
    "21": "keep",
    "23": "keep",
    "25": "keep",
    "27": "keep",
    "30": "keep",
    "31": "keep",
    "33": "keep",
    "35": "keep",
    "37": "keep",
    "41": "keep",
    "43": "keep",
    "45": "keep",
    "47": "keep",
    "49": "keep",
    "51": "keep",
    "53": "keep",
    "56": "keep",
    "58": "keep",
    "60": "keep",
    "62": "keep",
    "64": "keep"
   },
   "unknown_result_associations": 2,
   "questions": 6,
   "answered": 6,
   "retained": 6,
   "correct": 5,
   "abstained": 0,
   "answerability_correct": 5,
   "unsupported_correct_guess": 0,
   "lost_answer": 0,
   "qa_available": true,
   "strata": {
    "tool_dependent": {
     "questions": 4,
     "answered": 4,
     "retained": 4,
     "correct": 4,
     "abstained": 0,
     "answerability_correct": 4,
     "unsupported_correct_guess": 0,
     "lost_answer": 0
    },
    "permanent_control": {
     "questions": 2,
     "answered": 2,
     "retained": 2,
     "correct": 1,
     "abstained": 0,
     "answerability_correct": 1,
     "unsupported_correct_guess": 0,
     "lost_answer": 0
    }
   },
   "per_question": [
    {
     "question": "cc1",
     "kind": "tool_dependent",
     "retained": true,
     "gold": "no",
     "answer": "no",
     "correct": true,
     "abstained": false,
     "answerability_correct": true,
     "unsupported_correct_guess": false,
     "lost_answer": false
    },
    {
     "question": "cc2",
     "kind": "permanent_control",
     "retained": true,
     "gold": "commons",
     "answer": "commons",
     

… truncated, 2997 of 96904 bytes shown. Complete safe artifact

openjev (Codiv free hosted) - NOT JEVresults/v2_codiv_full_r3.json
{
 "schema_version": 2,
 "project": "compaction-check",
 "subset": "full",
 "repeat": 3,
 "timestamp": "2026-09-24T16:34:45Z",
 "input_hashes": {
  "samples/compaction-check/fixtures/questions.json": "0d348f75b58e9df36b14babdfe9ae57f192a70d2452d757385839706712f00d1",
  "samples/compaction-check/fixtures/sessions.jsonl": "6b6e01855f16358971ca6734810ebec8d3a7b2240d9a8c05dbfd7c3f0ca115e6"
 },
 "source_hashes": {
  "samples/_shared/community.py": "88155b9fcb87dbc2af50a59a02cdcc85a87e9404debaa6e6e3a96744c104d2cd",
  "samples/_shared/jev.py": "5a0a37f5097a0ba9c7ccfc4a75990c7d825c809e48f2e407341526b6b8bbcc28",
  "samples/compaction-check/build_excerpts.py": "28a991f6ddcd8324b2336543bfde6f3d3ba379dd78c45a876d24d93d54d1e981",
  "samples/compaction-check/compact.py": "80df701cac548bf404ad02794635baab5bb7cf92b2b9e7acd92c263b6c9fd5ec"
 },
 "label": "openjev (Codiv free hosted) - NOT JEV",
 "mode": "live",
 "provider": "codiv",
 "rows": [
  {
   "session": "claude-code__opus-5-5",
   "case_id": "claude-code__opus-5-5",
   "repeat": 3,
   "method": "full",
   "context_chars": 16157,
   "tokens_estimate": 4040,
   "tokens_full_estimate": 4040,
   "tokens_saved_estimate": 0,
   "saved_fraction": 0.0,
   "decisions": {
    "keep": 30,
    "truncate": 0,
    "drop": 0
   },
   "result_decisions": {
    "1": "keep",
    "3": "keep",
    "5": "keep",
    "7": "keep",
    "9": "keep",
    "12": "keep",
    "14": "keep",
    "16": "keep",
    "18": "keep",
    "21": "keep",
    "23": "keep",
    "25": "keep",
    "27": "keep",
    "30": "keep",
    "31": "keep",
    "33": "keep",
    "35": "keep",
    "37": "keep",
    "41": "keep",
    "43": "keep",
    "45": "keep",
    "47": "keep",
    "49": "keep",
    "51": "keep",
    "53": "keep",
    "56": "keep",
    "58": "keep",
    "60": "keep",
    "62": "keep",
    "64": "keep"
   },
   "unknown_result_associations": 2,
   "questions": 6,
   "answered": 6,
   "retained": 6,
   "correct": 5,
   "abstained": 0,
   "answerability_correct": 5,
   "unsupported_correct_guess": 0,
   "lost_answer": 0,
   "qa_available": true,
   "strata": {
    "tool_dependent": {
     "questions": 4,
     "answered": 4,
     "retained": 4,
     "correct": 4,
     "abstained": 0,
     "answerability_correct": 4,
     "unsupported_correct_guess": 0,
     "lost_answer": 0
    },
    "permanent_control": {
     "questions": 2,
     "answered": 2,
     "retained": 2,
     "correct": 1,
     "abstained": 0,
     "answerability_correct": 1,
     "unsupported_correct_guess": 0,
     "lost_answer": 0
    }
   },
   "per_question": [
    {
     "question": "cc1",
     "kind": "tool_dependent",
     "retained": true,
     "gold": "no",
     "answer": "no",
     "correct": true,
     "abstained": false,
     "answerability_correct": true,
     "unsupported_correct_guess": false,
     "lost_answer": false
    },
    {
     "question": "cc2",
     "kind": "permanent_control",
     "retained": true,
     "gold": "commons",
     "answer": "commons",
     

… truncated, 2997 of 96919 bytes shown. Complete safe artifact

REAL JEV (OpenCode Zen free, jev-1.13-free)results/v2_zen_zen_r1.json
{
 "schema_version": 2,
 "project": "compaction-check",
 "subset": "zen",
 "repeat": 1,
 "timestamp": "2026-09-24T16:17:39Z",
 "input_hashes": {
  "samples/compaction-check/fixtures/questions.json": "0d348f75b58e9df36b14babdfe9ae57f192a70d2452d757385839706712f00d1",
  "samples/compaction-check/fixtures/sessions.jsonl": "6b6e01855f16358971ca6734810ebec8d3a7b2240d9a8c05dbfd7c3f0ca115e6"
 },
 "source_hashes": {
  "samples/_shared/community.py": "88155b9fcb87dbc2af50a59a02cdcc85a87e9404debaa6e6e3a96744c104d2cd",
  "samples/_shared/jev.py": "6bd98ea32afdccc886148ee92aaa1213a4ed2c9d2c797739450a1a1856d16a06",
  "samples/compaction-check/build_excerpts.py": "28a991f6ddcd8324b2336543bfde6f3d3ba379dd78c45a876d24d93d54d1e981",
  "samples/compaction-check/compact.py": "80df701cac548bf404ad02794635baab5bb7cf92b2b9e7acd92c263b6c9fd5ec"
 },
 "label": "REAL JEV (OpenCode Zen free, jev-1.13-free)",
 "mode": "live",
 "provider": "zen",
 "rows": [
  {
   "session": "claude-code__opus-5-5",
   "case_id": "claude-code__opus-5-5",
   "repeat": 1,
   "method": "full",
   "context_chars": 16157,
   "tokens_estimate": 4040,
   "tokens_full_estimate": 4040,
   "tokens_saved_estimate": 0,
   "saved_fraction": 0.0,
   "decisions": {
    "keep": 30,
    "truncate": 0,
    "drop": 0
   },
   "result_decisions": {
    "1": "keep",
    "3": "keep",
    "5": "keep",
    "7": "keep",
    "9": "keep",
    "12": "keep",
    "14": "keep",
    "16": "keep",
    "18": "keep",
    "21": "keep",
    "23": "keep",
    "25": "keep",
    "27": "keep",
    "30": "keep",
    "31": "keep",
    "33": "keep",
    "35": "keep",
    "37": "keep",
    "41": "keep",
    "43": "keep",
    "45": "keep",
    "47": "keep",
    "49": "keep",
    "51": "keep",
    "53": "keep",
    "56": "keep",
    "58": "keep",
    "60": "keep",
    "62": "keep",
    "64": "keep"
   },
   "unknown_result_associations": 2,
   "questions": 6,
   "answered": 6,
   "retained": 6,
   "correct": 6,
   "abstained": 0,
   "answerability_correct": 6,
   "unsupported_correct_guess": 0,
   "lost_answer": 0,
   "qa_available": true,
   "strata": {
    "tool_dependent": {
     "questions": 4,
     "answered": 4,
     "retained": 4,
     "correct": 4,
     "abstained": 0,
     "answerability_correct": 4,
     "unsupported_correct_guess": 0,
     "lost_answer": 0
    },
    "permanent_control": {
     "questions": 2,
     "answered": 2,
     "retained": 2,
     "correct": 2,
     "abstained": 0,
     "answerability_correct": 2,
     "unsupported_correct_guess": 0,
     "lost_answer": 0
    }
   },
   "per_question": [
    {
     "question": "cc1",
     "kind": "tool_dependent",
     "retained": true,
     "gold": "no",
     "answer": "no",
     "correct": true,
     "abstained": false,
     "answerability_correct": true,
     "unsupported_correct_guess": false,
     "lost_answer": false
    },
    {
     "question": "cc2",
     "kind": "permanent_control",
     "retained": true,
     "gold": "commons",
     "answer": "commons",
     

… truncated, 3000 of 87542 bytes shown. Complete safe artifact

Run it
./run.sh --offline --repeat-start 4
JEV_PROVIDER=codiv JEV_KEY_FILE="$CODIV_KEY_FILE" JEV_MAX_USD=0 ./run.sh --repeat-start 4
JEV_PROVIDER=zen JEV_KEY_FILE="$ZEN_KEY_FILE" JEV_MAX_USD=0 JEV_MIN_INTERVAL_S=20 ./run.sh --subset zen --repeat-start 4
./run.sh --mock --repeat-start 4