Jev Labs

Sift

Find the rows that meet your condition.

Who decides what

  1. CodeCheap exact filters (--where) run first
  2. Typed modelMany rows in one request: shared state {condition, rows}, one Noul per row
  3. CodeOne sweepable threshold; per-row cache makes re-thresholding free
  4. ResultMatching rows, with precision, recall and F1 on labelled data
codetyped modelhuman
Details · results, methods & full write-up
← All projects

Project 03

Sift

filter or rank rows with a plain-language condition

20lost_card, batch , fixed 0.5 F1: 0.818; n=240; 9…REAL JEV
openjev (Codiv free hosted) - NOT JEVREAL JEV (OpenCode Zen free, jev-1.13-free)

Select rows by a plain-language condition

Lots of useful data is unlabelled text: tickets, prompts, transcripts, reviews. sift keeps the rows that match a plain-language condition ("the customer wants to close their account"), or ranks them.

Who decides what

  1. CodeCheap exact filters (--where) run first
  2. Typed modelMany rows in one request: shared state {condition, rows}, one Noul per row
  3. CodeOne sweepable threshold; per-row cache makes re-thresholding free
  4. ResultMatching rows, with precision, recall and F1 on labelled data
codetyped modelhuman

Headline result

REAL JEV (OpenCode Zen free, jev-1.13-free)

lost_card, batch 20, fixed 0.5 F1: 0.818; n=240; 95% interval [0.722, 0.897]; baseline select every row: 0.286; zen_banking77_lost_card_batch20.json

openjev (Codiv free hosted) - NOT JEV

lost_card, batch 5, fixed 0.5 F1: 0.864; n=240; 95% interval [0.769, 0.935]; baseline select every row: 0.286; real_banking77_lost_card_batch5.json

openjev vs real Jev

REAL JEV (OpenCode Zen free, jev-1.13-free)openjev (Codiv free hosted) - NOT JEV
Requests
real Jev117
openjev403
Input tokens
real Jev105,743
openjev145,599
Median latency lower is better
real Jev445 ms
openjev301 ms
Cost at list price the openjev tokens would be $0.0061
real Jev$0
openjev$0

Requests, input tokens and median latency from the project's call log as recorded in RESULTS_SUMMARY (development runs included); both routes were free tiers. free tiers change, unverified

Details

The problem

Lots of useful data is unlabelled text: tickets, prompts, transcripts, reviews. sift keeps the rows that match a plain-language condition ("the customer wants to close their account"), or ranks them. It's a reusable tool for everything else in this repo.

How Jev is used

Several rows are packed into ONE request as a shared state {condition, rows}, with one Noul per row pointing at rows[i] (pg-jev's recipe). Cheap exact filters (--where) run in code first. The per-row cache makes re-thresholding free. --label reports precision and recall, and --sweep reports precision, recall and F1 across 11 thresholds.

Why decomposing helps. Each row is its own yes/no question, so a threshold is one number you can sweep on labelled data. There's no prompt to rewrite and no prose to parse.

Live results: openjev (Codiv free hosted) - NOT JEV, failures included

A bigger labelled set. Two sets were built from BANKING77 (PolyAI, CC-BY-4.0, test split, via the Hugging Face datasets server). The labels are human intent labels, not ours:

  • close_account: 220 rows, 40 positive (terminate_account), with hard negatives cancel_transfer, request_refund and edit_personal_details, plus 60 random rows.
  • lost_card: 240 rows, 40 positive (lost_or_stolen_card), with hard negatives compromised_card, lost_or_stolen_phone, card_not_working and card_swallowed, plus 40 random rows.

Only the text field is sent. The intent and label never are.

setrows per requestF1 at 0.5 (P / R)best threshold → F1input tokens (per row)
close_account20 (pg-jev default)0.758 (0.96 / 0.63)0.2 → 0.83119,627 (89)
close_account101.000 (1.00 / 1.00)any 0.2–0.9 → 1.00015,211 (69)
close_account51.000 (1.00 / 1.00)any 0.2–0.9 → 1.00016,938 (77)
lost_card200.613 (0.86 / 0.48)0.2 → 0.78621,552 (90)
lost_card100.787 (0.69 / 0.93)0.95 → 0.88116,680 (70)
lost_card50.864 (0.79 / 0.95)0.9 → 0.89518,589 (77)
lost_card1 (one row per request)0.857 (0.82 / 0.90)0.7 → 0.88934,044 (142)
tickets (24, our own labels)200.667 (0.75 / 0.60)0.3 → 0.8002,958
mock (word overlap), close_account—0.800 at 0.50.6 → 0.889—
mock, lost_card—0.345 at 0.50.6 → 0.771—

What the sweep found:

  • On openjev, the batch size mattered far more than the threshold.
  • At 20 rows per request, recall collapsed. Blunt positives like "I am highly unsatisfied with this company and want to delete my account!" scored 0.01, which looks like answers drifting off their row index.
  • At 5–10 rows, close_account was perfect at any threshold, and lost_card went from F1 0.61 to 0.86 at 0.5.
  • Smaller batches did not cost more tokens per row (70–77 vs about 90). One row per request cost about twice as much and was no better.
  • So on openjev the default batch is 5. On real Jev it stays 20, pg-jev's measurement, which the real-Jev run below confirms (F1 0.99 / 0.82 at 20 rows).
  • Remaining errors on lost_card are the genuinely confusable hard negatives: "someone used my card" (compromised_card) and "the ATM took my card" (card_swallowed).
  • Mock caveat: the mock beat openjev-at-batch-20 on close_account. That shows batch 20 was broken. Word overlap is still weak: mock lost_card is 0.35.

Full details, raw outputs and the mock baseline: RESULTS.md and results/.

Real Jev run

Real Jev = jev-1.13-free on OpenCode Zen's limited-time free tier, run 2026-09-24 13:19–13:21 UTC with JEV_PROVIDER=zen. It was $0, and no response carried a cost field. Each result is labelled REAL JEV (OpenCode Zen free, jev-1.13-free). The free tier rate-limited us (HTTP 429 FreeUsageLimitError) after about 288 calls in about 2 minutes. Later calls were paced at ≤4 per minute.

setrows per requestopenjev F1 at 0.5 (best)REAL JEV F1 at 0.5 (best)
close_account (220)200.758 (0.831 at 0.2)0.988 (1.000 at 0.6)
close_account51.0001.000
lost_card (240)200.613 (0.786 at 0.2)0.818 (0.907 at 0.8)
lost_card50.864 (0.895 at 0.9)0.854 (0.914 at 0.7)
tickets (24)200.667, P 3/4 R 3/51.000, P 5/5 R 5/5

pg-jev's batch-20 claim holds on real Jev and fails on openjev. Real Jev at 20 rows per request is close to its own 5-row numbers. So the default of 5 is an openjev workaround. Real Jev does not need it. The hard negatives (compromised_card, card_swallowed) are still where both models make most of their mistakes. Real Jev tokens: about 91 per row at batch 20, 132 at batch 5; median latency 420–480 ms per request.

Raw outputs: results/zen_*.

Cost, tokens, latency
  • Tokens: 70–90 input tokens per row in batches, 142 per row one at a time.
  • Latency: median 250–530 ms per request. Larger batches were slower per request but faster per row.
  • Everything in this upgrade: 401 requests, about 143k tokens, $0. At list price, 10,000 rows is about 750k tokens, roughly $0.03.
What real Jev would change

(Written before the real-Jev run. The section above shows what it did.)

  • Verify the batch knee. pg-jev measured real Jev at 100% for 1–20 rows per request. If that holds, real Jev keeps batch 20 and costs fewer requests. Measure it with --batch 5/10/20 --sweep on these same two sets before trusting it.
  • Threshold depends on the set. Best thresholds differed by set (1.0 at anything on close_account, 0.9 on lost_card), so thresholds must come from labelled examples of the actual job, not a default.
Advice folded in
  • Latent Space: "dark data" is one of the two biggest money makers he names. "Choose thresholds from real examples" (about 1:07) is what --sweep does.
  • building-with-jev (dbreunig): "accuracy falls as inputs grow → filter in code, send only the needed fields" matches both --fields text and the batch-size finding. Sources: ../../research/latent-space-diogo-almeida.md, /dbreunig/building-with-jev-skill, /smartdio/jev-browser-agent, /jyje/pilot-typesafeai-jev.
Patterns it borrows
  • pg-jev: 20 rows share one state with one Noul per row (rows[i]). pg-jev measured 100% accuracy at ≤20 rows per batch, 92–98% at 40 and 77–94% at 80, and ~175 tokens/row at 20 vs ~435 for a single row. Batch size is fixed at 20 until we re-measure it ourselves.
  • Answers are cached per row content (live mode only, so mock answers can't poison the cache). --max-rows and the call cap act as spend guards.
  • jev-curate: cheap host-side filters run before the model call.
Raw result files
REAL JEV (OpenCode Zen free, jev-1.13-free)results/zen_lost_card_batch20.txt
mode=live [REAL JEV (OpenCode Zen free, jev-1.13-free)]  rows=240  requests=12  median_latency_ms=423.5  input_tokens=21853  est_usd=0.000000  matched 48 at p>=0.5
vs 'is_lost_or_stolen_card': precision 36/48  recall 36/40
threshold  picked  tp  fp  fn  precision  recall  f1
     0.05     192  40  152   0      0.208     1.0  0.345
      0.1     144  40  104   0      0.278     1.0  0.435
      0.2     101  39  62   1      0.386   0.975  0.553
      0.3      84  39  45   1      0.464   0.975  0.629
      0.4      64  38  26   2      0.594    0.95  0.731
      0.5      48  36  12   4       0.75     0.9  0.818
      0.6      43  36   7   4      0.837     0.9  0.867
      0.7      40  35   5   5      0.875   0.875  0.875
      0.8      35  34   1   6      0.971    0.85  0.907   <- best F1
      0.9      31  30   1  10      0.968    0.75  0.845
     0.95      24  23   1  17      0.958   0.575  0.719
REAL JEV (OpenCode Zen free, jev-1.13-free)results/zen_tickets_leave.txt
mode=live [REAL JEV (OpenCode Zen free, jev-1.13-free)]  rows=24  requests=2  median_latency_ms=632.4  input_tokens=2979  est_usd=0.000000  matched 5 at p>=0.5
vs 'threatens_to_leave': precision 5/5  recall 5/5
threshold  picked  tp  fp  fn  precision  recall  f1
     0.05       6   5   1   0      0.833     1.0  0.909
      0.1       6   5   1   0      0.833     1.0  0.909
      0.2       5   5   0   0        1.0     1.0  1.000
      0.3       5   5   0   0        1.0     1.0  1.000
      0.4       5   5   0   0        1.0     1.0  1.000
      0.5       5   5   0   0        1.0     1.0  1.000   <- best F1
      0.6       4   4   0   1        1.0     0.8  0.889
      0.7       4   4   0   1        1.0     0.8  0.889
      0.8       4   4   0   1        1.0     0.8  0.889
      0.9       3   3   0   2        1.0     0.6  0.750
     0.95       3   3   0   2        1.0     0.6  0.750
Run it
cd samples/sift && ./run.sh
# = sift.py fixtures/banking77_close_account.jsonl "the customer wants to close or delete their account" --fields text --label wants_to_close_account --sweep --batch 5
#   sift.py fixtures/banking77_lost_card.jsonl "the customer says their card has been lost or stolen" --fields text --label is_lost_or_stolen_card --sweep --batch 5
#   sift.py fixtures/tickets.jsonl "the customer threatens to cancel or leave" --label threatens_to_leave --sweep