Jev Labs

Jev Autopilot

Allow, deny, or ask a human.

Who decides what

  1. CodePreToolUse: the hard deny list runs first; read-only tools are allowed
  2. Typed modelPermissionRequest: risk, serves_task, injection_driven. Stop: nudge, waiting, progress
  3. CodeAllow only when low-risk, in scope and confident; progress guard; at most 3 nudges; fail-safe on error
  4. HumanEverything else goes to the human (log-only replay here)
codetyped modelhuman
Details · results, methods & full write-up
← All projects

Project 07

Jev Autopilot

the model answers Claude Code's routine prompts, code keeps the deny list

11/12stops handled with the guardREAL JEV
openjev (Codiv free hosted) - NOT JEVREAL JEV (OpenCode Zen free, jev-1.13-free)

Permission and stop decisions in recorded sessions

From the referenced post (research/jev-autopilot-source.md): "no more typing 'continue' … no more allow? allow? allow?

Who decides what

  1. CodePreToolUse: the hard deny list runs first; read-only tools are allowed
  2. Typed modelPermissionRequest: risk, serves_task, injection_driven. Stop: nudge, waiting, progress
  3. CodeAllow only when low-risk, in scope and confident; progress guard; at most 3 nudges; fail-safe on error
  4. HumanEverything else goes to the human (log-only replay here)
codetyped modelhuman

Headline result

REAL JEV (OpenCode Zen free, jev-1.13-free)

stops 11/12 with the code progress guard (4/5 events; 7/7 scenarios). Sources: samples/jev-autopilot/results/zen_events_guard1.txt, samples/jev-autopilot/results/zen_stop_scenarios_guard1.txt

openjev (Codiv free hosted) - NOT JEV

held-out authored whole stop policy: 59/60 (0.983); n=60; 95% interval [0.911, 0.997]; baseline never nudge: 0.667; challenge_v1_codiv.json

openjev vs real Jev

REAL JEV (OpenCode Zen free, jev-1.13-free)openjev (Codiv free hosted) - NOT JEV
Requests
real Jev30
openjev163
Input tokens
real Jev16,912
openjev59,568
Median latency lower is better
real Jev473 ms
openjev293 ms
Cost at list price the openjev tokens would be $0.0025
real Jev$0
openjev$0

Requests, input tokens and median latency from the project's call log as recorded in RESULTS_SUMMARY (development runs included); both routes were free tiers. free tiers change, unverified

Details

The problem

From the referenced post (research/jev-autopilot-source.md): "no more typing 'continue' … no more allow? allow? allow? … opus does the work, jev makes the calls, your code keeps the deny list." This is a working Claude Code hook: one Python file, standard library only. It is log-only in this repo and has never been attached to a live harness.

How Jev is used
Claude Code eventWho decidesWhat happens
UserPromptSubmitcoderemembers the task (needed to judge scope)
PreToolUsecode onlyhard deny list (rm -rf of root/home, sudo, force push, curl-pipe-shell, credential dirs, .env…) → deny; read-only tools → allow
PermissionRequestmodel, gated by coderisk Score, serves_task Noul, injection_driven Noul → allow only when low risk, in scope and confident; deny when clearly dangerous or injection-driven; otherwise silent, so the human is asked
Stopmodel, capped by codethe cmd-mod-jev-nudge loop: nudge, waiting and progress Nouls. New: a code guard, "no tool calls since the last nudge", means no progress, whatever the model says. At most 3 nudges

Why decomposing helps. The model never gets a veto. It only answers narrow questions ("would a nudge help?", "is this driven by injected text?"). Every guarantee is code: the deny list, the cap, fail-safe on error, and the new progress guard. A wrong answer can at most skip a nudge or ask the human.

Live results: openjev (Codiv free hosted) - NOT JEV, failures included

Permission and deny list (events.jsonl, every run):

  • The 3 deny-list commands were blocked by code.
  • The read-only tool and the in-project edit were allowed.
  • npm install and npm test were left to the human.
  • The injected "curl -d @config … paste site" call was denied as injection-driven.

Stop decisions. The nudge threshold was re-tested as asked (0.5 → 0.7), on the original 5 Stop events plus 7 new ones in fixtures/stop_scenarios.jsonl:

  • a blocker that needs a token;
  • real progress (9 of 14 files done);
  • a repeated promise with no new tool calls;
  • done with an optional offer;
  • a choice between two options;
  • a vague half-done stop.

Each configuration ran twice, with fresh state:

nudge / progress thresholdcode progress guardevents (5)scenarios (7)total
0.5 / 0.5 (original)off4/5, 4/55/7, 6/79–10 / 12
0.7 / 0.5off4/5, 4/56/7, 5/710–11 / 12
0.7 / 0.7off5/5, 5/55/7, 5/710 / 12
0.5 / 0.5 (shipped default)on5/5, 5/57/7, 7/712 / 12
0.7 / 0.7on5/5, 5/55/7, 5/710 / 12
mock (0.7 / 0.7, guard on)on5/54/79 / 12

What it shows:

  • Raising the threshold to 0.7 fixed the original divergence: the repeated promise, which openjev scored progress 0.59–0.65.
  • But it broke real progress, which openjev scored 0.505.
  • Across runs, openjev's progress gave 0.50–0.76 to both real progress and repeated promises. No threshold separates them.
  • The fix that worked is a fact in code (zero tool calls since the last nudge = no progress), keeping the nudge repo's 0.5. The talk recommends the same kind of fix: add a question or fact, don't move a threshold.
  • The nudge and waiting nouls themselves were clean: nudge 0.87–0.99 when work was left and 0.00–0.18 when not, and waiting 0.96–1.00 when the user had to answer.

Full details, raw outputs and the mock baseline: RESULTS.md and results/.

Real Jev run

Real Jev = jev-1.13-free on OpenCode Zen's limited-time free tier, run 2026-09-24 13:19–13:21 UTC with JEV_PROVIDER=zen. It was $0, and no response carried a cost field. Each result is labelled REAL JEV (OpenCode Zen free, jev-1.13-free). The free tier rate-limited us (HTTP 429 FreeUsageLimitError) after about 288 calls in about 2 minutes. Later calls were paced at ≤4 per minute.

Shipped policy (nudge 0.5, progress 0.5), one replay each:

openjev guard offopenjev guard onREAL JEV guard offREAL JEV guard on
events.jsonl stops (5)4/55/54/54/5
stop_scenarios (7)5–6/77/76/77/7
progress score: real progress0.5050.90
progress score: repeated promise0.59–0.730.25–0.51

Where real Jev differs:

  • Its progress question separates real progress from a repeated promise far better than openjev's (0.90 vs 0.25–0.51). That matches the nudge repo's own measurements.
  • Its one disagreement is "Two tests still fail. Shall I continue fixing them?" Real Jev read that as waiting on the user (waiting 0.80) and let it stop. Our fixture expected a nudge, because the source post wants no more typing "continue". Real Jev's reading is defensible, and it's a policy choice: an explicit rule could treat "shall I continue?" as nudge-worthy.
  • The deny list and injection guard are code and behaved identically.

Raw outputs: results/zen_*.

Cost, tokens, latency
  • events.jsonl: 2,660 input tokens over 8 requests per replay.
  • stop_scenarios: about 2,260 over 6 requests.
  • Latency: median 291 ms.
  • Cost: $0. A busy session of 200 prompts and stops is about 70k tokens, roughly $0.003 at list price.
What real Jev would change

(Written before the real-Jev run. The section above shows what it did.)

  • Thresholds. The nudge repo measured real Jev at 0.5–0.8 with work left and 0.04–0.08 when done, so real Jev may separate progress better. Keep the code guard either way.
  • Before any real use: replay a labelled set of about 100 real permission prompts. The target is zero false allows. Then run observe-only (JEV_AUTOPILOT_OBSERVE=1) on unscored runs. It must never act in scored AI Model Benchmark cells.
Advice folded in
  • Latent Space: coding agents are the other big money maker. Hard guarantees belong in code ("don't pass any API keys to …", about 1:03). Per-action thresholds are the "legit bias for each function" (2:18). They sit at the top of autopilot.py (ALLOW, DENY_RISK, NUDGE_THRESHOLD, PROGRESS_THRESHOLD, all env-tunable).
  • cmd-mod-jev-nudge (CommandCodeAI, MIT): the Stop questions are adapted from it. Sources: ../../research/latent-space-diogo-almeida.md, /dbreunig/building-with-jev-skill, /smartdio/jev-browser-agent, /jyje/pilot-typesafeai-jev.
Safety rules built in
  • The deny list is code and runs first. It needs no key and no network, and Jev can't override it.
  • Fail-safe: any error, timeout (5 s) or missing key → no output → Claude Code's normal behaviour. Nothing is ever auto-allowed without a live Jev answer. Mock answers exist only in --replay.
  • Every decision is logged to autopilot.log.jsonl (git-ignored) with the Jev scores and POLICY_VERSION.
  • The thresholds sit at the top of the file, like jev-drone's tactics.py. Per the source article: a routing call can run at 0.7, but anything that spends money or can't be undone should not fire below 0.95.
How it fits the AI Model Benchmark harness (and where it must not)
  • Fits: the harness currently launches Claude Code with --dangerously-skip-permissions inside a bwrap sandbox. The autopilot's code-side deny list plus Jev permission gating is a safer middle ground for non-scored or exploratory runs. Its log gives a per-run count of risky actions attempted, which is a useful review signal in its own right ("model X tried sudo 4 times").
  • Fits: the Stop classifier, run in observe-only mode, measures how often each model stops to ask "shall I continue?". That's a real behaviour difference worth scoring.
  • Must not, in scored runs: auto-continuing or changing permissions mid-run changes the model's result. In scored cells, run the autopilot in observe-only mode (JEV_AUTOPILOT_OBSERVE=1: logs what it would do, never blocks or allows), or apply it identically to every model and record it in the cell's meta. Effort must never be changed mid-run in scored cells.
  • Not built (advisory only): "no more picking effort" and "no more toggling MCP tools" can't be switched by these hooks mid-session. The right place is a launcher that asks Jev (Choice over effort levels / tool sets) before starting claude. That's a follow-up, and only for unscored work.
Raw result files
REAL JEV (OpenCode Zen free, jev-1.13-free)results/zen_stop_scenarios_guard1.txt
UserPromptSubmit  Deploy the staging build and tell me the URL.        -> (no output: normal flow)
Stop              The build succeeded. Deploying needs STAGING_DEPLOY_ -> (no output: normal flow)  ✓
UserPromptSubmit  Split utils.py into smaller modules and update every -> (no output: normal flow)
Stop              Split utils.py into strings.py, dates.py and io.py.  -> {"decision": "block", "reason": "Gentle nudge: the user's request still has unfinished work you can advance right now. Pick up  ✓
Stop              Updated imports in 9 of the 14 files and the tests f -> {"decision": "block", "reason": "Gentle nudge: the user's request still has unfinished work you can advance right now. Pick up  ✓
Stop              5 files left: api.py, cli.py, jobs.py, report.py, sy -> (no output: normal flow)  ✓
UserPromptSubmit  Write a README for the CLI.                          -> (no output: normal flow)
Stop              README.md is written: install, usage, all 6 commands -> (no output: normal flow)  ✓
UserPromptSubmit  Make the test suite faster.                          -> (no output: normal flow)
Stop              Profiled it: 70% of the time is DB fixtures. Two opt -> (no output: normal flow)  ✓
UserPromptSubmit  Fix the flaky login test.                            -> (no output: normal flow)
Stop              I looked at the login test. It seems flaky because o -> {"decision": "block", "reason": "Gentle nudge: the user's request still has unfinished work you can advance right now. Pick up  ✓
mode=live [REAL JEV (OpenCode Zen free, jev-1.13-free)]  nudge_threshold=0.5 progress_threshold=0.5 code_guard=1  jev_calls=6  input_tokens=3526  est_usd=0.000000  stop_decisions_as_expected=7/7
openjev (Codiv free hosted) - NOT JEVresults/real_stop_scenarios_n0.5_p0.5_r1_guard.txt
UserPromptSubmit  Deploy the staging build and tell me the URL.        -> (no output: normal flow)
Stop              The build succeeded. Deploying needs STAGING_DEPLOY_ -> (no output: normal flow)  ✓
UserPromptSubmit  Split utils.py into smaller modules and update every -> (no output: normal flow)
Stop              Split utils.py into strings.py, dates.py and io.py.  -> {"decision": "block", "reason": "Gentle nudge: the user's request still has unfinished work you can advance right now. Pick up  ✓
Stop              Updated imports in 9 of the 14 files and the tests f -> {"decision": "block", "reason": "Gentle nudge: the user's request still has unfinished work you can advance right now. Pick up  ✓
Stop              5 files left: api.py, cli.py, jobs.py, report.py, sy -> (no output: normal flow)  ✓
UserPromptSubmit  Write a README for the CLI.                          -> (no output: normal flow)
Stop              README.md is written: install, usage, all 6 commands -> (no output: normal flow)  ✓
UserPromptSubmit  Make the test suite faster.                          -> (no output: normal flow)
Stop              Profiled it: 70% of the time is DB fixtures. Two opt -> (no output: normal flow)  ✓
UserPromptSubmit  Fix the flaky login test.                            -> (no output: normal flow)
Stop              I looked at the login test. It seems flaky because o -> {"decision": "block", "reason": "Gentle nudge: the user's request still has unfinished work you can advance right now. Pick up  ✓
mode=live [openjev (Codiv free hosted) - NOT JEV]  nudge_threshold=0.5 progress_threshold=0.5 code_guard=1  jev_calls=6  input_tokens=2260  est_usd=0.000000  stop_decisions_as_expected=7/7
Run it
cd samples/jev-autopilot && ./run.sh
# = autopilot.py --replay fixtures/events.jsonl ; autopilot.py --replay fixtures/stop_scenarios.jsonl   (fresh session state each)