Jev Labs

Planner Actor Poker

Test who should decide the next bet.

Who decides what

  1. CodeName the spot: made hand, draws, equity bucket, position, street. No raw cards reach the model
  2. Typed modelPlanner every 25 hands: reads code-counted opponent stats, picks one of 5 strategy notes
  3. Typed modelActor on every decision: keep_playing and aggress (Nouls)
  4. CodeCompose the action: fold / check-call / bet-raise
codetyped modelhuman
Details · results, methods & full write-up
← All projects

Project 08

Planner Actor Poker

a planner writes the strategy, the model makes every betting decision

100none planner vs equity: bb/ hands: 67.800; n=200…openjev · NOT JEV
openjev (Codiv free hosted) - NOT JEVREAL JEV (OpenCode Zen free, jev-1.13-free)claims by browser_use unverified

Compare a betting actor with and without a planner

browser_use posted "Luna (planner) -> Jev (actor) playing poker and winning" (source post, 2026-09-24). This project is our version of that split.

Who decides what

  1. CodeName the spot: made hand, draws, equity bucket, position, street. No raw cards reach the model
  2. Typed modelPlanner every 25 hands: reads code-counted opponent stats, picks one of 5 strategy notes
  3. Typed modelActor on every decision: keep_playing and aggress (Nouls)
  4. CodeCompose the action: fold / check-call / bet-raise
codetyped modelhuman

Headline result

REAL JEV (OpenCode Zen free, jev-1.13-free)

not run on real Jev: exceeds free-tier budget; baseline not available; saved results

openjev (Codiv free hosted) - NOT JEV

none planner vs equity: bb/100 hands: 67.800; n=200; 95% interval [38.500, 97.100]; baseline code equity hero: 0.000; real_suite_200deals.json

openjev vs real Jev

REAL JEV (OpenCode Zen free, jev-1.13-free)openjev (Codiv free hosted) - NOT JEV
Requests
real Jevpending
openjev2,276
Input tokens
real Jevpending
openjev763,324
Median latency lower is better
real Jevpending
openjev350 ms
Cost at list price the openjev tokens would be $0.0321
real Jevpending
openjev$0

pending: operator live run

Requests, input tokens and median latency from the project's call log as recorded in RESULTS_SUMMARY (development runs included); both routes were free tiers. free tiers change, unverified

Details

The problem

browser_use posted "Luna (planner) -> Jev (actor) playing poker and winning" (source post, 2026-09-24). This project is our version of that split:

  • A planner reads the opponent's tendencies every 25 hands and writes a short strategy note.
  • An actor makes every betting decision from small typed questions.

We measure whether the actor wins, and whether the planner helps.

How Jev is used
  • Game: heads-up fixed-limit Texas Hold'em, chosen over no-limit so every decision is fold / check-call / bet-raise, with lower variance. Blinds are 1/2 chips (1 bb = 2 chips). Bets are 1 bb on preflop and flop and 2 bb on turn and river, capped at 4 bets per street.
  • Code facts: the made hand, named ("top pair", "overpair", "AK suited"); draws; equity against a random hand from Monte Carlo, bucketed into words ("strong", "weak"); whether calling is profitable on raw equity; position; street; the opponent's last action. No raw cards reach the model. Only named categories do, so identical situations share one cached answer. That is the dino-runner lesson.
  • Actor: one request per decision, two Nouls. keep_playing (continue rather than fold?) is asked only when there is something to call. aggress (bet or raise now?) is asked only when a raise is legal. Code composes the answer: fold if keep_playing < 0.5, raise if aggress ≥ 0.5, else check or call.
  • Planner, every 25 hands. It reads code-counted opponent stats (fold-to-bet rate, raise share, showdown rate) and picks one of 5 strategy notes, which goes into the actor's state. There are three variants:
    • (a) heuristic: deterministic rules;
    • (b) openjev: a Choice over the same 5 notes. openjev can't write free text, so the planner picks a note from a menu. A text-writing planner like browser_use's Luna would be a separate LLM, which we don't use: no paid models;
    • (c) none: the actor-only ablation.
  • Opponents: random; tight_passive (plays about 20% of hands, rarely raises); equity (raises at ≥65% equity vs a random hand, calls when equity beats the pot odds).
  • Reference hero: the same equity bot, as a code-only baseline with no model.
  • Duplicate format. Every deal is played twice with seats and cards swapped, so card luck mostly cancels. The equity-bot mirror match scores exactly 0.0 ± 0.0, which checks the engine.

Why decomposing helps. Odds, equity and hand reading are code. The model makes two yes/no judgments per spot, and the planner changes only one field of the state. Each piece can be ablated alone, which is how we found the planner was hurting.

Live results: openjev (Codiv free hosted) - NOT JEV, failures included

bb/100 for our hero, with 95% confidence intervals. 200 duplicate deals = 400 hands per cell, 12 cells, 4,800 hands in all:

opponentcode-only equity bot (no model)actor only (no planner)actor + heuristic planneractor + openjev planner
random+103.2 ± 35.8+68.8 ± 31.5+55.6 ± 26.2+55.6 ± 26.2
tight-passive-4.4 ± 28.7+38.9 ± 23.6+15.2 ± 32.1+12.9 ± 34.5
equity+0.0 ± 0.0+67.8 ± 29.3-4.4 ± 16.4-4.4 ± 16.4

What it shows:

  • The actor alone wins against all three bots, and every interval is clear of zero. Against tight-passive and equity it beats the code-only equity bot playing the same seats and cards (+38.9 vs −4.4, and +67.8 vs 0.0).
  • Against random it wins less than the equity bot does (+68.8 vs +103.2). The actor folds more marginal hands than a pure pot-odds rule.
  • The planner made things worse. Against the equity bot, the planner's read was "call_down_maniac": the bot raises a lot, but only with strong hands. That note turned +67.8 into −4.4. Against tight-passive, "exploit_folds" (bluff more) turned +38.9 into +15.2.
  • The planner reads frequencies, not hand strength. It sees a high raise share and concludes "maniac", so its advice is wrong for a value-raiser.
  • The openjev planner picked the same note as the heuristic in at least 46 of 48 design steps, so it made the same mistake. It's cheap (16–43 calls per 400 hands), but it isn't smarter here.
  • Honest takeaway: a planner on top added no edge here. It needs a better read (showdown hand strength as well as frequencies) before it helps. The actor-only ablation is the one to beat.
  • The notes each planner used are listed per cell in results/real_suite_200deals.json.
Real Jev run

REAL JEV (OpenCode Zen free, jev-1.13-free): not run on real Jev: exceeds free-tier budget. No retry is scheduled. The saved results contain the openjev pass; its measured sample and interval are generated above. Free-route evaluation is reserved for the frozen held-out projects.

Cost, tokens, latency
  • openjev suite: 2,072 requests, 698,491 input tokens, $0. That's about 330 tokens per decision.
  • Cache: 4,800 hands needed only 2,072 requests, because the state has no raw cards and identical situations reuse the cached answer.
  • Latency: median 335–385 ms per decision.
  • At TypeSafe's list price ($0.042 per 1M input tokens) the whole suite would cost about $0.05.
What real Jev would change
  • Better-calibrated keep_playing / aggress should mostly show up as fewer marginal folds. See the REAL JEV section for what the small real-Jev sample showed.
  • The planner problem is a design problem. A different model would get the same frequency-only read. The next version should feed the planner showdown hand strength and use the "add a question" fix: a Noul like "are this opponent's raises mostly value?" rather than a raise-frequency rule.
Source and credit

Inspired by browser_use's "Luna (planner) -> Jev (actor) playing poker and winning" post (2026-09-24): claims unverified. The numbers on this page are ours, from our own simulation.

Advice folded in
  • Latent Space: games are a named family. Put state in as JSON, keep facts in code, and ask small independent questions (both Nouls travel in one request). Don't call the model inside a tight loop.
  • building-with-jev (dbreunig): "the final decision is wrong while each answer is right → change the policy in code". That describes the planner, which is policy.
  • browser_use (the post above): the planner/actor split itself.
Raw result files
openjev (Codiv free hosted) - NOT JEVresults/real_suite_200deals.txt
[code-only hero (no model)] hero=equitybot planner=n/a       vs random        hands=400  bb/100=+103.2 ±35.8 (95% CI)  calls=0 tokens=0 median_ms=None wall_s=4.9 notes={}
[code-only hero (no model)] hero=equitybot planner=n/a       vs tight_passive hands=400  bb/100=-4.4 ±28.7 (95% CI)  calls=0 tokens=0 median_ms=None wall_s=4.3 notes={}
[code-only hero (no model)] hero=equitybot planner=n/a       vs equity        hands=400  bb/100=+0.0 ±0.0 (95% CI)  calls=0 tokens=0 median_ms=None wall_s=6.5 notes={}
[openjev (Codiv free hosted) - NOT JEV] hero=actor     planner=none      vs random        hands=400  bb/100=+68.8 ±31.5 (95% CI)  calls=595 tokens=196329 median_ms=336.2 wall_s=281.3 notes={}
[openjev (Codiv free hosted) - NOT JEV] hero=actor     planner=none      vs tight_passive hands=400  bb/100=+38.9 ±23.6 (95% CI)  calls=91 tokens=25823 median_ms=354.9 wall_s=47.6 notes={}
[openjev (Codiv free hosted) - NOT JEV] hero=actor     planner=none      vs equity        hands=400  bb/100=+67.8 ±29.3 (95% CI)  calls=149 tokens=46123 median_ms=344.7 wall_s=76.5 notes={}
[openjev (Codiv free hosted) - NOT JEV] hero=actor     planner=heuristic vs random        hands=400  bb/100=+55.6 ±26.2 (95% CI)  calls=609 tokens=218368 median_ms=366.8 wall_s=309.2 notes={'call_down_maniac': 16}
[openjev (Codiv free hosted) - NOT JEV] hero=actor     planner=heuristic vs tight_passive hands=400  bb/100=+15.2 ±32.1 (95% CI)  calls=395 tokens=134721 median_ms=335.1 wall_s=195.2 notes={'exploit_folds': 14, 'balanced': 2}
[openjev (Codiv free hosted) - NOT JEV] hero=actor     planner=heuristic vs equity        hands=400  bb/100=-4.4 ±16.4 (95% CI)  calls=158 tokens=51258 median_ms=384.5 wall_s=84.2 notes={'call_down_maniac': 16}
[openjev (Codiv free hosted) - NOT JEV] hero=actor     planner=openjev   vs random        hands=400  bb/100=+55.6 ±26.2 (95% CI)  calls=16 tokens=5482 median_ms=None wall_s=24.0 notes={'call_down_maniac': 16}
[openjev (Codiv free hosted) - NOT JEV] hero=actor     planner=openjev   vs tight_passive hands=400  bb/100=+12.9 ±34.5 (95% CI)  calls=43 tokens=14913 median_ms=438.5 wall_s=44.2 notes={'exploit_folds': 16}
[openjev (Codiv free hosted) - NOT JEV] hero=actor     planner=openjev   vs equity        hands=400  bb/100=-4.4 ±16.4 (95% CI)  calls=16 tokens=5474 median_ms=None wall_s=10.5 notes={'call_down_maniac': 16}
total calls=2072 input_tokens=698491 est_usd=0.000000
Run it
cd samples/planner-actor-poker && ./run.sh      # = poker.py --suite --deals 200 --threads 3