Jev Labs

Post Virality

Test what predicts engagement.

Who decides what

  1. Code5 code checks: number, media, link, list format, length bucket
  2. Typed modelOne request per post (text only): 7 Nouls plus a hook-type Choice
  3. CodeScorer: mean of hook-group mean, 5 nearest posts and per-check lift, learned on training folds
  4. Code5-fold held-out test: Spearman vs log(likes), against three baselines
codetyped modelhuman
Details · results, methods & full write-up
← All projects

Project 09

Post Virality

typed checks and hook type from the model, a deterministic scorer, an honest held-out test

0.262held-out Spearman correlation: ; n=84; 95% inter…openjev · NOT JEV
openjev (Codiv free hosted) - NOT JEVREAL JEV (OpenCode Zen free, jev-1.13-free)claims unverified

Rank posts with typed checks and a code scorer

@0xMovez's post (source post, 2026-09-24) described a "viral post prediction analyser".

Who decides what

  1. Code5 code checks: number, media, link, list format, length bucket
  2. Typed modelOne request per post (text only): 7 Nouls plus a hook-type Choice
  3. CodeScorer: mean of hook-group mean, 5 nearest posts and per-check lift, learned on training folds
  4. Code5-fold held-out test: Spearman vs log(likes), against three baselines
codetyped modelhuman

Headline result

REAL JEV (OpenCode Zen free, jev-1.13-free)

not run on real Jev: exceeds free-tier budget; baseline not available; saved results

openjev (Codiv free hosted) - NOT JEV

held-out Spearman correlation: 0.262; n=84; 95% interval [0.069, 0.449]; baseline followers only: 0.017; real_openjev.json

openjev vs real Jev

REAL JEV (OpenCode Zen free, jev-1.13-free)openjev (Codiv free hosted) - NOT JEV
Requests
real Jevpending
openjev84
Input tokens
real Jevpending
openjev64,475
Median latency lower is better
real Jevpending
openjev362 ms
Cost at list price the openjev tokens would be $0.0027
real Jevpending
openjev$0

pending: operator live run

Requests, input tokens and median latency from the project's call log as recorded in RESULTS_SUMMARY (development runs included); both routes were free tiers. free tiers change, unverified

Details

The problem

@0xMovez's post (source post, 2026-09-24) described a "viral post prediction analyser":

  • Jev parses 800 viral posts;
  • it groups them by hook type (build demo, receipt, launch, contrarian…);
  • it runs 12 typed checks on every post (hook, numbers, media, CTA);
  • a new post is scored against its hook type and its 5 most similar viral posts.

This is a small, honest version with a held-out test of whether the score ranks higher-engagement posts higher.

How Jev is used
  • Data: 84 public posts. 78 are the posts behind X links the human has sent (saved link-index, the Sep 24 batch, and links in our research notes) and 6 are posts quoted inside them.
    • Fetched read-only through `, one request about every second, with no login (fetch.py`).
    • Median 199 likes, max 94,637.
    • This is not a random sample of X. It's what one person found worth sharing, mostly AI and dev posts.
  • 12 checks:
    • 5 in code: has a number, has media, has a link, list format, length bucket. Code can see these for sure.
    • 7 model Nouls, one request per post: strong hook, concrete result, shows how, novelty, call to action, names a known product, hype language.
    • 1 model Choice: hook type (build demo, receipt, launch, contrarian, tutorial or list, news or commentary, other).
  • The model sees only the post text: no counts, no author, no follower numbers.
  • Scorer (code): the mean of three predictions in log-likes space, learned from the training folds only: (i) the hook group's mean, (ii) the mean of the 5 nearest posts in check space, and (iii) a per-check lift sum. It is reported as a 0–100 percentile against the training posts.
  • Explanation: a short template (hook type, checks passed, 3 nearest posts, score), not a second model.
  • Evaluation: 5-fold cross-validation, so every post is scored by a scorer that never saw it. We report the Spearman rank correlation with log(likes) and a bootstrap 95% CI, against three baselines: followers only, code checks only, and a shuffled control.

Why decomposing helps. Each check is a separate yes/no you can inspect, and the scorer is plain arithmetic. When the result is weak, you can see which part is weak. Nothing is a black-box "virality score".

Live results: openjev (Codiv free hosted) - NOT JEV, stated plainly

Spearman rank correlation with log(likes) on held-out folds, 84 posts, with 95% bootstrap CI:

scoreropenjevmock
model checks + hook groups + 5-NN (the project)+0.262 [+0.07, +0.45]+0.159 [-0.05, +0.35]
code checks only (no model)-0.022 [-0.23, +0.19]same
followers only+0.017 [-0.22, +0.24]same
shuffled control+0.185 [-0.04, +0.41]same
the project's score vs likes per follower+0.239 [+0.01, +0.43]+0.097 [-0.13, +0.32]

What it shows:

  • The result is weak. The model-check scorer has a positive rank correlation of about 0.26, and its CI only just clears zero.
  • The shuffled control reached +0.19 by chance, which shows how noisy 84 posts are.
  • The code-only checks and follower count carried no signal on this set.
  • So the model's judgments (hook, concrete result, novelty…) are the only part with any signal. But this is not a virality predictor anyone should trust yet.
  • What would make it meaningful: a baseline of several hundred posts from one niche, drawn systematically (not hand-shared), with engagement normalised by audience and age.
  • openjev hook types: {'launch': 17, 'build_demo': 19, 'news_or_commentary': 21, 'other': 5, 'tutorial_or_list': 13, 'receipt': 9}.
Real Jev run

REAL JEV (OpenCode Zen free, jev-1.13-free): not run on real Jev: exceeds free-tier budget. No retry is scheduled. The saved results contain the openjev pass; its measured sample and interval are generated above. Free-route evaluation is reserved for the frozen held-out projects.

Cost, tokens, latency
  • openjev: 84 requests (one per post), 64,475 input tokens (about 767 per post), median 362 ms, $0.
  • At list price a 10,000-post baseline is about 7.7M tokens, roughly $0.32.
What real Jev would change
  • Calibration. Better-calibrated checks give cleaner check vectors and so better neighbours. Data size and selection limit this project more than the model does.
  • The scorer is deliberately dumb. With a real baseline, it should be refit per niche.
Source and credit

Inspired by a published viral-post predictor post (2026-09-24): claims unverified. The numbers on this page are ours, from our own held-out test.

Advice folded in
  • @0xMovez (the post above): the 12-typed-checks and hook-group design.
  • Latent Space: "dark data" is the biggest money maker he names, and "evaluate on your own data, not public benchmarks" is why this project has a held-out test.
  • building-with-jev (dbreunig): one property per question, and counts in code (numbers, media, links, length are code checks).
Raw result files
openjev (Codiv free hosted) - NOT JEVresults/real_openjev.json

highest 3 and lowest 2 of 84 posts by the project's score

authorlikeshookscore (log)
chetanankola670build_demo2.95
Bhavani_00007668receipt2.94
Stefan_3D_AI515news_or_commentary2.92
DataChaz9other2.01
RoundtableSpace54tutorial_or_list2.0
Run it
cd samples/post-virality && ./run.sh      # = virality.py fixtures/posts.jsonl   (fetch.py refreshes the posts)