Route the lead.
Make the right handoff before writing a reply.
Decide First. Write Later.
Format illustration. No frontier-LLM benchmark.
openjev (Codiv free hosted) - NOT JEV
Limits & sources
Authored leads, not production traffic. Matching the broad fixture expectation is not a complete correctness measure. Both providers passed; a plain LLM was not measured.
The initial request batches its questions. A reply gets its own request. Code handles scheduling, stops and confidence gates. Message writing remains a stub; nothing is sent.
Historical results and free-route estimates: REAL JEV log, openjev log. Typed request: Source artifact. Provider cost fields were absent; estimates are not receipts.
Details · results, methods & full write-up
Project 06
Lead Router
the decision layer of an automated sales workflow
Choose a sales follow-up action
From the sales-automation guide the human sent (research/sales-agent-guide.md). For each inbound lead it decides.
Who decides what
- CodeHours, counts, the price list and stop-on-reply
- Typed modelOne request, five questions: is_spam, is_urgent, intent, asks_price_not_on_list, reply_kind
- CodePick the route; below 0.75 confidence a human takes over
- ResultOnly the winning route would spend chat tokens (the writer is a stub)
Headline result
expected routing action: 7/7 (1.000); n=7; 95% interval [0.646, 1.000]; baseline keyword mock (plumbing only): 1.000; zen_leads.json, mock_leads.txt
held-out authored all routing fields correct: 81/100 (0.810); n=100; 95% interval [0.722, 0.875]; baseline fieldwise majority: 0.150; challenge_v1_codiv.json
openjev vs real Jev
Requests, input tokens and median latency from the project's call log as recorded in RESULTS_SUMMARY (development runs included); both routes were free tiers. free tiers change, unverified
Details
The problem
From the sales-automation guide the human sent (research/sales-agent-guide.md). For each inbound lead it decides:
- spam or lead seller (ignore);
- urgent (alert the human now);
- what they want;
- whether they're asking a price that isn't on the human's list (never guess one).
For each reply it decides yes / question / not_now / stop / other.
Code owns the schedule: a reply within 2 minutes, 4 follow-ups over 10 days in daytime only, stop the moment they reply, and hand off to the human with a note. It passes the follow-up number explicitly, which fixes the guide's one bug ("this is my last text" sent on follow-up 2).
How Jev is used
One request per item carries all the questions (fan-out):
| id | type | asks |
|---|---|---|
is_spam | Noul | spam, bot, or someone selling to the business |
is_urgent | Noul | needs handling today |
intent | Choice (5) | book / price / question / complaint / other |
asks_price_not_on_list | Noul | a price not in price_list (then don't guess) |
reply_kind | Choice (5) | yes / question / not_now / stop / other |
Below 0.75 confidence, the human takes over.
Why decomposing helps. The split works like this:
- Each judgment is independent, with its own threshold.
- Hours, counts, the price list and stop-on-reply are code, never model decisions.
- An LLM would only write the words, and only for the route that won. That's the jyje/pilot-typesafeai-jev pattern: one request asks everything, code picks the route, and only the winning route spends chat tokens. The writer is still a stub here.
Live results: openjev (Codiv free hosted) - NOT JEV, failures included
| openjev (Codiv), real | mock | |
|---|---|---|
routing decisions agreeing with expected (7 leads) | 7/7 (twice: 12:08Z and 12:51Z) | 7/7 |
| requests / input tokens | 10 / 2,897 | 10 / est. |
| median latency | 259 ms | — |
Differences from the mock, both defensible:
- "Late-night leak" got
first_reply_book(real) instead offirst_reply_question. - "Listed price" was also flagged urgent.
Honest limit: 7 hand-written leads is a smoke test. There is no real lead data in this repo, and nothing here messages anyone.
Full details, raw outputs and the mock baseline: RESULTS.md and results/.
Real Jev run
Real Jev = jev-1.13-free on OpenCode Zen's limited-time free tier, run 2026-09-24 13:19–13:21 UTC with JEV_PROVIDER=zen. It was $0, and no response carried a cost field. Each result is labelled REAL JEV (OpenCode Zen free, jev-1.13-free). The free tier rate-limited us (HTTP 429 FreeUsageLimitError) after about 288 calls in about 2 minutes. Later calls were paced at ≤4 per minute.
| openjev (NOT JEV) | REAL JEV | |
|---|---|---|
| routing decisions (7 leads) | 7/7 | 7/7 |
| requests | 10 | 10 |
One difference: real Jev did not flag the "Listed price" lead as urgent, and openjev did. The fixture's expected label accepts both, and real Jev's reading ("just a price question") is the more natural one.
Raw outputs: results/zen_*.
Cost, tokens, latency
- Tokens: about 290 input tokens per request, 1–2 requests per lead.
- Latency: median 259 ms.
- Cost: $0. At list price, 10,000 leads is about $0.001.
What real Jev would change
(Written before the real-Jev run. The section above shows what it did.)
- Calibration. Better calibration would make the 0.75 human-handoff gate more trustworthy. That is the main thing to measure, on a real, labelled lead export.
- From the talk:
- split "spam or seller" into separate nouls with their own thresholds;
- add an escalation noul (an angry or sarcastic high-value lead goes to the human), gated on confidence.
Advice folded in
- Latent Space: many independent guardrail questions, each with its own threshold (1:06–1:07), plus escalation by confidence.
- jev-browser-agent (smartdio): escalates below 0.6. Here the stricter 0.75 gate sends the lead to the human, which is the right bias when the cost of a wrong reply is a lost customer.
- pilot-typesafeai-jev (jyje): one request, many questions, code routes, and only the winner spends chat tokens. Sources: ../../research/latent-space-diogo-almeida.md, /dbreunig/building-with-jev-skill, /smartdio/jev-browser-agent, /jyje/pilot-typesafeai-jev.
Patterns it borrows
- Guide: trigger → first reply with one next step → follow-ups that stop on reply → human handoff with a note; test on mock leads first.
- TypeSafe intent routing + confidence gating: below 0.75 confidence, the human takes over.
- Code keeps the rules: hours, counts, the price list and stop-on-reply are never model decisions.
Raw result files
results/real_leads.txtmode=live [openjev (Codiv free hosted) - NOT JEV] leads=7 jev_calls=10 est_usd=0.000000 Late-night leak first_reply_book -> 4 follow-ups scheduled | URGENT: [urgent_note] for Late-night leak ✓ Unlisted price no_price_guess_offer_confirm -> 4 follow-ups scheduled ✓ Lead seller ignore (spam or seller) ✓ Listed price first_reply_price -> 4 follow-ups scheduled | URGENT: [urgent_note] for Listed price ✓ Booker, says yes first_reply_book -> answer once, then human_handoff with note ✓ Says stop first_reply_book -> stop all messages ✓ Not now first_reply_price -> human_handoff with note (lead said later) ✓
results/zen_leads.txtmode=live [REAL JEV (OpenCode Zen free, jev-1.13-free)] leads=7 jev_calls=10 est_usd=0.000000 Late-night leak first_reply_book -> 4 follow-ups scheduled | URGENT: [urgent_note] for Late-night leak ✓ Unlisted price no_price_guess_offer_confirm -> 4 follow-ups scheduled ✓ Lead seller ignore (spam or seller) ✓ Listed price first_reply_price -> 4 follow-ups scheduled ✓ Booker, says yes first_reply_book -> answer once, then human_handoff with note ✓ Says stop first_reply_book -> stop all messages ✓ Not now first_reply_price -> human_handoff with note (lead said later) ✓
Run it
cd samples/lead-router && ./run.sh # = router.py fixtures/leads.jsonl