Jev Labs

Connect Four.

Pick a column. Inspect each reply and its source.

Details · results, methods & full write-up
← All projects

Project 04

Connect4

better facts beat more calls, measured against a solver

2/8facts: game win rate: (0.250); n=8; 95% interval…REAL JEV
openjev (Codiv free hosted) - NOT JEVREAL JEV (OpenCode Zen free, jev-1.13-free)

Choose a legal move from code-computed facts

A controlled test of the claim behind typesafe-chess: a single-shot judgment model plays well when code supplies the facts (wins now, blocks, lets the opponent win on top, threats), and poorly without them.

Who decides what

  1. CodeList the legal columns and attach code-computed facts to each option (removed in the no-facts ablation)
  2. Typed modelA Choice over the legal columns only, plus a Score for the position value
  3. CodePlay the move against negamax depth 2 (or random)
  4. CodeGrade it with a depth-4 negamax: solver-best move? blunder?
codetyped modelhuman

Headline result

REAL JEV (OpenCode Zen free, jev-1.13-free)

facts: game win rate: 2/8 (0.250); n=8; 95% interval [0.071, 0.591]; baseline random hero against same opponent: not saved; zen_facts_8g.json

openjev (Codiv free hosted) - NOT JEV

held-out exact-search optimal move: 199/200 (0.995); n=200; 95% interval [0.972, 0.999]; baseline prefer center legal column: 0.755; challenge_v1_codiv.json

openjev vs real Jev

REAL JEV (OpenCode Zen free, jev-1.13-free)openjev (Codiv free hosted) - NOT JEV
Requests
real Jev46
openjev131
Input tokens
real Jev39,565
openjev82,480
Median latency lower is better
real Jev460 ms
openjev337 ms
Cost at list price the openjev tokens would be $0.0035
real Jev$0
openjev$0

Requests, input tokens and median latency from the project's call log as recorded in RESULTS_SUMMARY (development runs included); both routes were free tiers. free tiers change, unverified

Details

The problem

A controlled test of the claim behind typesafe-chess: a single-shot judgment model plays well when code supplies the facts (wins now, blocks, lets the opponent win on top, threats), and poorly without them. The historical game report uses a depth-limited move grader. The separate challenge evaluates near-terminal positions with exact search.

How Jev is used

Every move is one request:

  • a Choice over the legal columns only, so an illegal move is impossible;
  • a Score for the position value.

With facts on, each option's description carries code-computed facts. --no-facts removes them (the ablation). A depth-4 negamax is the yardstick: did the model pick a solver-best move, and did it blunder (pick a losing move when a safe one existed)? The opponent is negamax depth 2, or random.

Why decomposing helps. Search stays in code, which it is good at. The model only picks among labelled options, which is the "choice is a switch on an enum" shape from the talk.

Live results: openjev (Codiv free hosted) - NOT JEV, failures included
configuration (8 games each)W / D / Lsolver-best movesblunderscalls (moves)distinct games
openjev + facts, vs negamax-2, seed 74 / 0 / 40.580.05831 (69)5
openjev + facts, vs negamax-2, seed 114 / 0 / 40.600.03729 (80)—
openjev, no facts, vs negamax-2, seed 70 / 0 / 80.220.2229 (36)2
openjev, no facts, vs negamax-2, seed 110 / 0 / 80.220.2229 (36)identical to seed 7
openjev + facts, vs random8 / 0 / 00.770.00039 (44)—
mock + facts, vs negamax-21 / 0 / 70.470.200—5
mock, no facts, vs negamax-22 / 0 / 60.490.111—7
(earlier smoke test, 2 games, facts)1 / 0 / 10.430.2114

What it shows:

  • Facts are the whole game. Without them openjev lost all 8, blundered on 22% of moves and picked the best move 22% of the time. With them it went 4–4 against a depth-2 searcher, with 4–6% blunders.
  • The facts are a few words of code per option.

Honest limits:

  • Without facts openjev answers the same position the same way and the opponent's ties are seeded, so its 8 games are really 2 distinct games (seed 11 replays them exactly). Treat "0/8" as "lost both lines it knows", not 8 independent losses.
  • Repeated positions are served from the cache, which is why calls are fewer than moves.
  • The mock "no facts" did better than openjev "no facts" because the mock secretly computes the same facts. Treat it as a plumbing check only.

Full details, raw outputs and the mock baseline: RESULTS.md and results/.

Real Jev run

Real Jev = jev-1.13-free on OpenCode Zen's limited-time free tier, run 2026-09-24 13:19–13:21 UTC with JEV_PROVIDER=zen. It was $0, and no response carried a cost field. Each result is labelled REAL JEV (OpenCode Zen free, jev-1.13-free). The free tier rate-limited us (HTTP 429 FreeUsageLimitError) after about 288 calls in about 2 minutes. Later calls were paced at ≤4 per minute.

8 games vs negamax-2openjev (NOT JEV)REAL JEV
with facts: W / D / L4 / 0 / 42 / 0 / 6
with facts: solver-best / blunders0.58 / 0.0580.52 / 0.19
without facts0 / 0 / 8, blunders 0.22not run on real Jev: exceeds free-tier budget
tokens per call / calls~640 / 31~930 / 33

Honest result: with the same code facts, real Jev played worse than openjev here. It blundered on 19% of moves against openjev's 6%, over 8 partly repeated games. It's a small sample, but the direction is clear enough to say out loud. It fits the talk's warning that multi-hop lookahead isn't native: the facts cover one move deep, and the losses come from two-move traps.

Raw outputs: results/zen_*.

Cost, tokens, latency
  • Facts on: about 640 input tokens per call (19,773 for 31 calls).
  • Facts off: about 410 per call (3,695 for 9).
  • Latency: median 335–355 ms per move.
  • Cost: about 118 requests and 73k tokens in all, $0.
What real Jev would change

(Written before the real-Jev run. The section above shows what it did.)

  • What real Jev might add. Better quiet-position play (the non-fact moves, where openjev is at about chance) would show up as a higher solver-best rate with facts on. The no-facts ablation should stay poor, because the talk says multi-hop lookahead isn't native.
  • Cascade (from the talk). When move confidence is low, call the solver instead. We log confidence per move (results/*.json) so the cascade rate can be measured.
Advice folded in
  • Latent Space: games are a named family. "choice is a switch on an enum", multi-hop degrades, and System 2 search isn't native (1:23, 1:55). The ablation shows exactly that. Sources: ../../research/latent-space-diogo-almeida.md, /dbreunig/building-with-jev-skill, /smartdio/jev-browser-agent, /jyje/pilot-typesafeai-jev.
Patterns it borrows
  • typesafe-chess: legal-move Choice, facts in option descriptions, Score as the value head (ready for MCTS later), and a fixed external yardstick.
  • jev-drone: "the state has to contain the answer". The ablation shows what happens when it doesn't.
  • Jaggedness page: the board goes in as text rows with a legend; nothing asks the model to count.
Raw result files
openjev (Codiv free hosted) - NOT JEVresults/real_facts_8g.txt
[openjev (Codiv free hosted) - NOT JEV]
{
 "mode": "live",
 "facts": true,
 "opponent": "negamax:2",
 "games": 8,
 "win": 4,
 "draw": 0,
 "loss": 4,
 "jev_moves": 69,
 "solver_best_rate": 0.58,
 "blunder_rate": 0.058,
 "calls": 31,
 "input_tokens": 19773,
 "est_usd": 0.0
}
REAL JEV (OpenCode Zen free, jev-1.13-free)results/zen_facts_8g.txt
[REAL JEV (OpenCode Zen free, jev-1.13-free)]
{
 "mode": "live",
 "facts": true,
 "opponent": "negamax:2",
 "games": 8,
 "win": 2,
 "draw": 0,
 "loss": 6,
 "jev_moves": 63,
 "solver_best_rate": 0.524,
 "blunder_rate": 0.19,
 "calls": 33,
 "input_tokens": 30707,
 "est_usd": 0.0
}
openjev (Codiv free hosted) - NOT JEVresults/real_nofacts_8g.txt
[openjev (Codiv free hosted) - NOT JEV]
{
 "mode": "live",
 "facts": false,
 "opponent": "negamax:2",
 "games": 8,
 "win": 0,
 "draw": 0,
 "loss": 8,
 "jev_moves": 36,
 "solver_best_rate": 0.222,
 "blunder_rate": 0.222,
 "calls": 9,
 "input_tokens": 3695,
 "est_usd": 0.0
}
Run it
cd samples/connect4 && ./run.sh
# = connect4.py --games 8 --opponent negamax:2 --judge-depth 4            (facts)
#   connect4.py --games 8 --opponent negamax:2 --judge-depth 4 --no-facts (ablation)