Keep it running.
Choose the obstacles. Jev picks the maneuver.
Low Confidence. Code Takes Over.
Historical simulator · confidence veto enabled
openjev (Codiv free hosted) - NOT JEV
Comparison & limits
These are deterministic obstacle abstractions with perfect timing, not browser gameplay results or a frontier-LLM baseline. All arms use a code rule table below the confidence gate. A fallback is a code decision, not an unassisted model success.
Current browser play uses a newly recorded category set at fixed speed. Its score is separate from these historical runs.
Details · results, methods & full write-up
Project 05
Dino Runner
one decision per obstacle, code owns timing
Choose a maneuver; code controls timing
The smallest real-time control loop: a deterministic abstraction of the Chrome dino game (cacti of four sizes, birds at five heights). A wrong maneuver ends the run.
Who decides what
- CodeTurn geometry into a named flight_path category (the ablation sends raw pixel heights)
- Typed modelA Choice maneuver (jump / duck / keep_running) plus a speculative jump_profile
- CodeBelow 0.5 confidence the rule table takes over (the veto); cached answers are reused; timing stays in code
- ResultOne maneuver per obstacle
Headline result
category: mean obstacles cleared: 60.000; n=3; 95% interval [60.000, 60.000]; baseline always jump; code rule also shown: 4.333; zen_category.txt
held-out collision-safe direct action: 75/100 (0.750); n=100; 95% interval [0.657, 0.825]; baseline always jump: 0.290; challenge_v1_codiv.json
openjev vs real Jev
Requests, input tokens and median latency from the project's call log as recorded in RESULTS_SUMMARY (development runs included); both routes were free tiers. free tiers change, unverified
Details
The problem
The smallest real-time control loop: a deterministic abstraction of the Chrome dino game (cacti of four sizes, birds at five heights). A wrong maneuver ends the run. It tests the plumbing any game or agent demo needs:
- one call per event, not per frame;
- a fingerprint cache;
- a confidence fallback;
- a call budget.
It also measures TypeSafe's warning that the model is weak with raw numbers.
How Jev is used
One request per new obstacle:
- a Choice
maneuver(jump / duck / keep_running); - a speculative Choice
jump_profile, used only on a jump.
Code turns geometry into a named flight_path category first. --raw-state is the ablation: it sends pixel heights instead. Below 0.5 confidence, code applies the rule table instead (the veto). Identical situations reuse the cached answer.
Why decomposing helps. Timing and geometry stay in code. The model answers "given a bird in the head-height lane, which move?". That is a category judgment. The arithmetic stays in code.
Live results: openjev (Codiv free hosted) - NOT JEV, failures included
| policy | seed 1 | seed 2 | seed 3 |
|---|---|---|---|
| always_jump | died at obstacle 2 | died at 2 | died at 9 |
| rules (code, perfect by construction) | cleared 60 | cleared 60 | cleared 60 |
openjev, flight_path category | cleared 60 (1 fallback) | cleared 60 (1 fallback) | cleared 60 (1 fallback) |
openjev, --raw-state (pixels) | died at obstacle 2 (ducked a bird at 70 px) | died at 0 (ducked a bird at 10 px) | died at 11 (ducked a bird at 10 px) |
| mock, raw state | cleared 60 | cleared 60 | cleared 60 |
What it shows:
- The category is doing the work. With code-computed categories openjev matched the rules table on all three seeds, using 19 calls for 180 obstacles thanks to the cache.
- Given raw pixel heights, it died within 11 obstacles every time, always by ducking a low bird it should have jumped.
- This is TypeSafe's jaggedness warning, reproduced.
- The mock's raw-state pass means nothing: the mock computes the category internally.
Full details, raw outputs and the mock baseline: RESULTS.md and results/.
Real Jev run
Real Jev = jev-1.13-free on OpenCode Zen's limited-time free tier, run 2026-09-24 13:19–13:21 UTC with JEV_PROVIDER=zen. It was $0, and no response carried a cost field. Each result is labelled REAL JEV (OpenCode Zen free, jev-1.13-free). The free tier rate-limited us (HTTP 429 FreeUsageLimitError) after about 288 calls in about 2 minutes. Later calls were paced at ≤4 per minute.
| policy (60 obstacles × 3 seeds) | openjev (NOT JEV) | REAL JEV |
|---|---|---|
| category state | 60/60 ×3, 1 fallback per seed | 60/60 ×3, 1 fallback per seed |
| raw pixel state | died at 2, 0, 11 | 60/60 ×3, but 5 fallbacks per seed |
| calls / tokens | 19 + 8 / 5,309 + 2,466 | 19 + 25 / 10,420 + 14,375 |
Calibration explains the raw-state difference:
- On raw pixels, openjev was confidently wrong. It ducked low birds at high confidence and died.
- Real Jev was unsure on the same situations. It fell below the 0.5 gate 5 times per seed, so code's rule table took over and the run survived.
- The code veto only helps a model that knows when it doesn't know. Calibration gives it that. The category design is still the right one: zero raw-state fallbacks would be better than five.
Raw outputs: results/zen_*.
Cost, tokens, latency
- Category run: 19 calls, 5,309 input tokens (about 280 per call), median 293 ms.
- Raw-state run: 8 calls before dying, 2,466 tokens, median 321 ms.
- Cost: $0.
- Latency budget. One decision per obstacle at about 0.3 s means code must ask when an obstacle appears, not when it arrives. That is the talk's point that intelligence per second is "not Jev's niche" and the model should stay out of the frame loop.
What real Jev would change
(Written before the real-Jev run. The section above shows what it did.)
- Raw numbers. Real Jev may do better on raw numbers than openjev, but TypeSafe's own docs say not to rely on it. The category design should stay.
- Speed. The real difference would be latency. The talk describes Jev as fast per request (a few hundred ms), so the per-obstacle design is the same either way.
Advice folded in
- Latent Space: real-time use is described as "not Jev's niche", so don't call it in the game loop (53:00, 2:10). The speculative
jump_profileis the "ask everything you might need at once" pattern. - building-with-jev (dbreunig): "errors on numeric nearness → move the arithmetic to code" is exactly the raw-state result. Sources: ../../research/latent-space-diogo-almeida.md, /dbreunig/building-with-jev-skill, /smartdio/jev-browser-agent, /jyje/pilot-typesafeai-jev.
Patterns it borrows
- jev-t-rex-runner: per-obstacle Choice with a semantic
flight_path; code handles timing. - jev-drone: decide when to ask, reuse fingerprinted judgments, keep the veto in code (low confidence falls back to rules).
- Speculative fan-out: the jump profile is asked every time and used only on a jump.
Raw result files
results/real_raw_state.txtmode=live [openjev (Codiv free hosted) - NOT JEV] state=raw pixels
seed 1 {"policy": "always_jump", "cleared": 2, "died_on": ["bird", "single", 70], "action": ["jump", "full"], "fallbacks": 0, "calls_saved_by_cache": 0}
seed 1 {"policy": "rules", "cleared": 60, "died_on": null, "fallbacks": 0, "calls_saved_by_cache": 0}
seed 1 {"policy": "jev", "cleared": 2, "died_on": ["bird", "single", 70], "action": ["duck", null], "fallbacks": 0, "calls_saved_by_cache": 0}
seed 2 {"policy": "always_jump", "cleared": 2, "died_on": ["bird", "single", 45], "action": ["jump", "full"], "fallbacks": 0, "calls_saved_by_cache": 0}
seed 2 {"policy": "rules", "cleared": 60, "died_on": null, "fallbacks": 0, "calls_saved_by_cache": 0}
seed 2 {"policy": "jev", "cleared": 0, "died_on": ["bird", "single", 10], "action": ["duck", null], "fallbacks": 0, "calls_saved_by_cache": 0}
seed 3 {"policy": "always_jump", "cleared": 9, "died_on": ["bird", "single", 85], "action": ["jump", "full"], "fallbacks": 0, "calls_saved_by_cache": 0}
seed 3 {"policy": "rules", "cleared": 60, "died_on": null, "fallbacks": 0, "calls_saved_by_cache": 0}
seed 3 {"policy": "jev", "cleared": 11, "died_on": ["bird", "single", 10], "action": ["duck", null], "fallbacks": 1, "calls_saved_by_cache": 7}
jev calls=8 input_tokens=2466 est_usd=0.000000results/zen_raw_state.txtmode=live [REAL JEV (OpenCode Zen free, jev-1.13-free)] state=raw pixels
seed 1 {"policy": "always_jump", "cleared": 2, "died_on": ["bird", "single", 70], "action": ["jump", "full"], "fallbacks": 0, "calls_saved_by_cache": 0}
seed 1 {"policy": "rules", "cleared": 60, "died_on": null, "fallbacks": 0, "calls_saved_by_cache": 0}
seed 1 {"policy": "jev", "cleared": 60, "died_on": null, "fallbacks": 5, "calls_saved_by_cache": 51}
seed 2 {"policy": "always_jump", "cleared": 2, "died_on": ["bird", "single", 45], "action": ["jump", "full"], "fallbacks": 0, "calls_saved_by_cache": 0}
seed 2 {"policy": "rules", "cleared": 60, "died_on": null, "fallbacks": 0, "calls_saved_by_cache": 0}
seed 2 {"policy": "jev", "cleared": 60, "died_on": null, "fallbacks": 5, "calls_saved_by_cache": 51}
seed 3 {"policy": "always_jump", "cleared": 9, "died_on": ["bird", "single", 85], "action": ["jump", "full"], "fallbacks": 0, "calls_saved_by_cache": 0}
seed 3 {"policy": "rules", "cleared": 60, "died_on": null, "fallbacks": 0, "calls_saved_by_cache": 0}
seed 3 {"policy": "jev", "cleared": 60, "died_on": null, "fallbacks": 5, "calls_saved_by_cache": 51}
jev calls=25 input_tokens=14375 est_usd=0.000000Run it
cd samples/dino-runner && ./run.sh
# = runner.py --obstacles 60 --seeds 1 2 3 ; runner.py --obstacles 60 --seeds 1 2 3 --raw-state