Jev Labs

Free Routing

Free routing and lean harness

Measured routing results

OpenCode Zen free chat via the CLI — NOT real Jev. Real Jev is queued, not run.

Frozen 24-task set; each row identifies its committed result source.
PolicyCorrectSource file and keyAcquisition caveat
Primary router24/24experiments/free-routing/results/analysis.json
primary_run1.policies.router
60-second call deadlines; offline policy replay over the original call matrix.
Secondary live router v219/24experiments/free-routing/results/analysis.json
live_router_v2
Fresh calls with 180-second deadlines; a separate acquisition, not a paired replay.
Fixed MiMo21/24experiments/free-routing/results/analysis.json
secondary_run2.fixed_arms.mimo-v2.6-flash-free
Retained run1 attempts use 60 seconds; new attempts use 180 seconds. Unavailable cells are not latency observations.

Mixed-deadline comparison: retained 60-second attempts and new 180-second attempts are not a controlled speed comparison. Fixed MiMo is not the designated strongest rung. Missing cost receipts remain unknown; reported zero cost does not fill missing receipts.

Details · results, methods & full write-up
← All projects

Project 13

Free Routing

Free routing and lean harness

Details

Reproduce offline

Python's standard library is sufficient. Set FREE_ROUTING_SCRATCH to an external scratch directory; runtime logs and spill files must stay outside the checkout. The examples assume this variable is set. Tests default to a sibling directory named after the checkout with a -scratch suffix.

nice -n 10 python3 experiments/free-routing/run.py analyse --backend replay \
  --records experiments/free-routing/results/calls.jsonl \
  --output "$FREE_ROUTING_SCRATCH/replay.json" \
  --scratch "$FREE_ROUTING_SCRATCH/context"

nice -n 10 python3 experiments/free-routing/run.py measure --backend mock \
  --output "$FREE_ROUTING_SCRATCH/mock_calls.jsonl" \
  --scratch "$FREE_ROUTING_SCRATCH/context"

nice -n 10 python3 -m unittest discover -s tests -p 'test_free_routing*.py'
nice -n 10 scripts/check.sh

Output collection refuses to overwrite a recording. Replay rejects mixed routes, missing task/model/prompt matches, malformed accounting and parsed answers that disagree with the retained raw text. Mock calls always say MOCK - NOT JEV. freeze.py deterministically rebuilds tasks.jsonl and its provenance manifest.

Run-1 real collection

OPENCODE_BIN must point to an already installed OpenCode executable. Set the scratch argument to its prepared external home. No package or repository install is performed. Only the explicit free model allowlist is accepted; the process inherits no API keys. It uses the CLI, a new session per call, native tool schemas with all permissions set to ask, and headless automatic rejection of tool requests. No browser is used. harness.json requires the API preference; the CLI is the necessary exception because Zen rejects direct free-tier chat API calls.

nice -n 10 python3 experiments/free-routing/run.py calibrate --backend opencode \
  --opencode "$OPENCODE_BIN" --scratch "$FREE_ROUTING_SCRATCH/oc" \
  --output "$FREE_ROUTING_SCRATCH/calibration.jsonl"

nice -n 10 python3 experiments/free-routing/run.py measure --backend opencode \
  --opencode "$OPENCODE_BIN" --scratch "$FREE_ROUTING_SCRATCH/oc" \
  --config "$FREE_ROUTING_SCRATCH/calibration.ladder.json" \
  --output "$FREE_ROUTING_SCRATCH/calls.jsonl"

The saved ladder.json was calibrated before scored calls. Cheaper means lower median wall latency on three small calibration calls, including CLI startup. “Strongest” denotes the final escalation rung; better quality is unverified. Calls are sequential and bounded by 60 seconds. Two consecutive transport failures cost stops the campaign immediately. Missing receipts remain unknown.

Scored model order rotates per task. Every policy is replayed over this same matrix, so the comparison consumes no additional model calls. Router latency sums the calls it uses; it excludes offline analysis and mock-adviser overhead. Latency percentiles include attempted tasks and their timeouts, excluding wholly skipped circuit cells; latency_n states that denominator. Every frozen case stays in the success-rate denominator. Valid wrong labels do not trigger escalation. The grader alone reads fixture answers. Reported token totals retain the provider's tokens.total; component counters are retained separately because they can disagree.

Stakes, shadow mode and the pending queue

--mode shadow is the default. The mock adviser independently picks the first allowed option; its agreement rate measures plumbing, not Jev quality. off uses the default router without consultation; on acts on adviser choices. Only initial model selection and failure-driven escalate/stop choices consult. Parsing is counted as a local decision. Forced stops are queued and recorded as code-only, with no consultation. Agreement excludes forced and local decisions.

The published pending_jev.jsonl is QUEUED — REAL JEV NOT RUN. To validate it later, use the existing external Jev credential file and strict zero-budget client:

JEV_PROVIDER=zen JEV_MAX_USD=0 JEV_MAX_CALLS=48 \
nice -n 10 python3 experiments/free-routing/run.py drain \
  --queue experiments/free-routing/pending_jev.jsonl \
  --output "$FREE_ROUTING_SCRATCH/jev_results.jsonl" \
  --scratch "$FREE_ROUTING_SCRATCH/context"

The shared client enforces its Zen lock, pacing, timeout and free-only guards. The drain resumes successful IDs, stops on the first error, and will not retry a retained refusal automatically. It reports REAL agreement only for verified live Zen receipts. Keep mock and live drain results in separate files. No drain has been executed against real Jev for this milestone.

The separate v2 queue has 57 stakes rows: 33 potential consultations and 24 forced stops. To validate that trace later, use its own receipt file:

JEV_PROVIDER=zen JEV_MAX_USD=0 JEV_MAX_CALLS=33 \
nice -n 10 python3 experiments/free-routing/run.py drain \
  --queue experiments/free-routing/pending_jev_v2.jsonl \
  --output "$FREE_ROUTING_SCRATCH/jev_results_v2.jsonl" \
  --scratch "$FREE_ROUTING_SCRATCH/context-v2"

This command is a queued follow-up, not part of the completed measurement.

Pass --jev-results "$FREE_ROUTING_SCRATCH/jev_results.jsonl" and --queue to analyse to replay completed real advice in shadow or on mode without more Jev calls. Advice must match the exact decision, including its default choice. If active advice reaches an unvalidated decision, replay fails closed; agreement on the original default path is not a complete active-policy evaluation.

Lean context contract

The classifier and escalation gate consume bounded JSON handoffs containing only routing features, prompt hashes, model and failure status. Oversized tool logs are saved intact under <scratch>/tool-output/; the consumed log includes that path and capped head/tail text. Classification evidence and rubrics stay complete in every model prompt. Caps are configurable in harness.json.

The comparison uses the same prompts, calls and handoff payloads with lean rules off and on. The off baseline inlines full task/history/log payloads at each handoff. The shared estimate_tokens helper estimates JSON characters divided by four. These are harness-context estimates, not measured model input-token savings. A smaller OpenCode system prompt remains untested.

Round 2 protocol and collection

Collection is complete at the 100-new-attempt cap: 60 planned fixed cells, 32 fresh router calls and eight excluded resource interruptions. The repaired fixed scores are MiMo 21/24, Space Bunny 24/24, Ultra 12/24 and Lightning 19/24. The primary v1 router remains 24/24; secondary live v2 is 19/24, with median 10.81 s and p90 193.80 s. All 24 v2 tasks were attempted; one additional escalation was blocked to reserve initial calls for remaining tasks. The commands below document the collection procedure; the published campaign budget is exhausted.

DESIGN.md and round2-plan.json register the extension before new calls. Run 1 remains the primary result; its original analysis is byte-preserved in results/run1/analysis.json, and its original recordings remain unchanged. Round 2 retries the 36 timeout/circuit-skipped cells, adds 24 Ultra fixed-arm cells, then runs a fresh router after the second ladder is frozen and committed.

All new calls use 180-second deadlines, run sequentially, and share a durable 100-attempt budget. Five consecutive errors stop one model; a free-tier, quota, transport and refuses mixed mock/real resumes or unresolved reservations. Keep the same ledger across phases; use unused output paths for each export.

nice -n 10 python3 experiments/free-routing/round2.py fixed --backend opencode \
  --opencode "$OPENCODE_BIN" --scratch "$FREE_ROUTING_SCRATCH/oc" \
  --ledger "$FREE_ROUTING_SCRATCH/round2-ledger.jsonl" \
  --output "$FREE_ROUTING_SCRATCH/round2-fixed.jsonl"

nice -n 10 python3 experiments/free-routing/round2.py freeze --backend opencode \
  --opencode "$OPENCODE_BIN" --scratch "$FREE_ROUTING_SCRATCH/oc" \
  --ledger "$FREE_ROUTING_SCRATCH/round2-ledger.jsonl" \
  --output experiments/free-routing/ladder-v2.json

freeze reads existing receipts and makes no calls. It uses original calibration and round-2 status, parseability and latency; no scoring answer is read. Commit the frozen ladder before grading new answers or running the next command. The router verifies the committed file and evidence hashes before making calls.

nice -n 10 python3 experiments/free-routing/round2.py router --backend opencode \
  --opencode "$OPENCODE_BIN" --scratch "$FREE_ROUTING_SCRATCH/oc" \
  --ledger "$FREE_ROUTING_SCRATCH/round2-ledger.jsonl" \
  --queue "$FREE_ROUTING_SCRATCH/pending-jev-v2.jsonl" \
  --output "$FREE_ROUTING_SCRATCH/router-v2.json"

The router reserves one initial call for each remaining task before escalation. The measured fixed arm for the final rung also serves as always-strongest, with all 24 task receipts reported directly and counted once. This is one measured arm, not an independent duplicate sample. Retained run-1 cells and new repairs remain distinguishable by provenance and deadline. Real Jev remains queued.

Memory guard and interrupted collection

Before every new call the runner reads /proc/meminfo. It launches only when 1 - MemAvailable / MemTotal is below 75%; otherwise it logs a wait and polls every 30 seconds inside the runner. Waiting is recorded separately from model latency. Memory above 85% during a call terminates the owned process group and ends that invocation with resource_interruption.

Resume the same phase with the same ledger and a new output filename. Completed cells are reused; resource-interrupted cells receive a new attempt sequence. interrupted_by_supervisor and aborted_memory_pressure consume budget slots but are excluded from accuracy, latency, reliability and model-error streaks. An unresolved reservation requires an explicit interruption receipt before resuming; it must never silently become a timeout or a successful call. A

Offline round-2 recount

After the frozen ladder is committed, the reporter reads the append-only ledger and retained run-1 recordings. It makes no model calls. The original analysis remains at results/run1/analysis.json; the combined result uses schema version 2 and separates primary_run1, secondary_run2 and live_router_v2.

nice -n 10 python3 experiments/free-routing/round2_report.py \
  --ledger experiments/free-routing/round2-ledger.jsonl \
  --router experiments/free-routing/results/run2/router-v2.json \
  --output "$FREE_ROUTING_SCRATCH/recount-round2.json"

The published ledger and results/run2/calls.jsonl are two representations of the same attempts. Count each call identity once. The ledger also retains reservations and protocol identity; the flattened file supplies top-level route labels for the result manifest. Source hashes and collection accounting are in results/run2/metadata.json.

Recounting the exported ledger and live summary reproduces the published results/analysis.json byte-for-byte. Acquisition totals include 769,493 known tokens and 14 missing usage/cost receipts; all received costs are $0. The report separates the mixed-deadline fixed comparison from each router version and explains the limits of this small authored set.

Raw result files
OpenCode Zen free chat via CLI - NOT JEVresults/analysis.json
{
 "acquisition": {
  "new_run2": {
   "attempts": 100,
   "known_reported_cost_usd": 0,
   "known_reported_tokens": 769493,
   "phase_attempts": {
    "matrix-repair": 44,
    "router-v2": 32,
    "ultra-fixed": 24
   },
   "receipt_ids": [
    "62dde0656795a3d51762be83",
    "b19034ae6f05586513e23ac9",
    "00ea932cfe6e48dfd0bc5282",
    "3acbf84e64efefcd3582e3e1",
    "976ea38738d0cf0270cb6ca6",
    "6ce49ca8bb5b748c1fa914b5",
    "e308c001b4c642d613fce1b5",
    "e2bce319e567d03e2b51b97f",
    "433285a62073d4b5f046ddb9",
    "32224561e827bcf57694b392",
    "fd7ddf45c1b3d4592468ac82",
    "98160d452f7ef488847a4cb0",
    "40b628e2029307acb49a5acc",
    "f40be22e64ece4739985a640",
    "dae9b908df5a79ec6603ac70",
    "6052a4ea6614aa501839b6ab",
    "f4bb038db7562fd2c745803f",
    "2cee28e92e56afafe5f06c8b",
    "c52b1607c577758551494bca",
    "5a8d4d28a8a84616222cdbd9",
    "9516fb481df0094fd6235bb2",
    "a07d762cadea3386572639e9",
    "c1a80aef9c1a2e20ce1112e3",
    "c20bab225f37f3acf60d8d83",
    "c52174cdfbfc2f25c565d6a1",
    "5394288201e4be3d474d635a",
    "3cd2ba3614929b18dcb67a9f",
    "2f7855cc9782e3b6478d92fa",
    "fc637a6161f4aac5fb4722d8",
    "88dce017501d6142acd829da",
    "ff1d8610eaefb6588e89e1d6",
    "a70eac1e9cf099800cb2629b",
    "b623315427fc5b86c7f07cbd",
    "e8ed32bcb8c2253cbd9b94d7",
    "92f859ea5837bdba5983d627",
    "586209217d5dc17c7ba60356",
    "7a4c8a86d866dc8411585b83",
    "0482021288c367e8e8d7e88b",
    "77ee1d5619e236a1aa2474ab",
    "5c1da1dba8c8599567776fe4",
    "00a2e73a785a3a1377ee006b",
    "7276fb092c7ca8358e372f57",
    "0e60b8e595f3a5a5683c7041",
    "de812c06ddc62a02adf1b758",
    "3bd827f6f0f76793cd3721e8",
    "a9047d0420853d5d328ffe78",
    "f65256c40d09998e9ec7ba62",
    "ba0b32486b3b8439c62ff6c8",
    "db811dc34686d8215b89da02",
    "ef12e0e19a8da3f06da1e012",
    "64ceb9dbc56b4818fd1dcf49",
    "a769aa33065877e7edb5c942",
    "3b9af047fc4a73f2bdfb198c",
    "e11214922abc3abea3f62f6c",
    "8db7f772b44dfa02b57a6ea3",
    "017d924dc8006b3a120b039b",
    "66683c36e06e5c6a163e297b",
    "ec70d6ac0ab7d76bbae57cb9",
    "75a3b603e85a35739f9ba0e8",
    "67a4f3b9b0ec1c9a05ed8ea0",
    "1d7654f162ed1b326b3c28bf",
    "7ad84c2965da64009790f10f",
    "11c2a2ee4cdda4b74a7efd77",
    "9f62b85d3e461f938567cfc5",
    "5c994b57ba137afa2f14adbf",
    "6105f210bfb9f9b6157e1756",
    "36683f25a3a0bf1d9fa6c214",
    "8a9c99719c20408c0120bd8e",
    "1028cd91726edfc3ff859a7c",
    "ceb4090c274e426755972154",
    "70bd4dcd8837d4a28ab952c2",
    "4824a857a24c9b07df7b44aa",
    "22d53946faf371edc14f1d4b",
    "a2d15069c36bf45e09ad88d2",
    "f888103fe6a27677238c8b7d",
    "a9dc557fd3aa563a1e2471f5",
    "46ede36a3e60427f1a89cb86",
    "89d7aff485c8849e7a6d1db1",
    "87e1c025e9403802f7b9d68f",
    "1a7654bc6864f1833afb2324",
    "9a0655df191ce6569ac83b03",
    "5e51cbbe9baf019844490309",
    "4f2509aeb66cdb9804880d6e",
    "8ee3ef55260df6bd614a4a92",
    "69cb8b0668c2ec939d806e61",
    "996a7d75115a7828f6ac3a2b",
    

… truncated, 2998 of 373245 bytes shown. Complete safe artifact

OpenCode Zen free chat via CLI - NOT JEVresults/metadata.json
{
 "access_check_attempts": 2,
 "calibration_attempts": 9,
 "call_timeout_s": 60,
 "cli_version": "1.18.32",
 "frozen_tasks_sha256": "2f867c4c6cd27059c76568cadf0bcdcff3cbbd5fa32c98a36098972cea9c4c77",
 "harness_sha256": "5ecf376a355b7003c3fdd58c81d3237ae639b33e8880e45b5fe2c0ee0bea7d13",
 "known_reported_cost_usd": 0,
 "label": "OpenCode Zen free chat via CLI - NOT JEV",
 "ladder_sha256": "70fff8c22ac1431634c44c8a82e74a70fd6e2fea336f5a32071cb7de06b0a7d2",
 "latency_limit": "Shared host; includes CLI startup. Calibration order was not stable during scored availability changes.",
 "matrix_cells": 72,
 "normalization": "Provider totals and original usage receipts restored from retained CLI events; raw text, timestamps, statuses and latency unchanged. Full raw events stay in <scratch>/oc/raw; exported usage receipts support offline recount.",
 "pending_decisions": 56,
 "profile": "isolated CLI, --pure, --auto=false, ask-all permissions (headless requests rejected)",
 "real_attempts_total": 55,
 "real_jev": "QUEUED - NOT RUN",
 "recordings_sha256": {
  "access_checks.jsonl": "8f4208cc75db4df7e780fb5178eaca1cfd6b7af68a283dabb1376d60f25ac54b",
  "calibration.jsonl": "c6e27f8b6006a0c9e58972493430c3b89856f9ba85bad1fd9d73bb75fdaac86c",
  "calls.jsonl": "e016cebc8c48c8c3bc9f7252aae499c8ab8b42d9502cd049cc009d5e500e04eb"
 },
 "reported_tokens_known": 383481,
 "schema_version": 1,
 "scored_attempts": 44,
 "scored_status_counts": {
  "ok": 36,
  "skipped_unavailable": 28,
  "timeout": 8
 },
 "sequential": true,
 "status": "complete",
 "timestamp_utc": "2026-09-25T05:21:02Z",
 "unknown_cost_attempts": 9,
 "unknown_token_attempts": 9
}

1646 of 1646 bytes shown. Complete safe artifact