Jev Labs

Post Ranking Association

A retrospective ranking test on 84 selected posts. Not a validated predictor.

Details · results, methods & full write-up
← All projects

Project 16

Post Ranking Association

a retrospective test, not a validated predictor

Details

Measured result

The committed-input recount includes 84/84 posts and 57 author groups. Route: OpenCode Zen free chat via CLI - NOT JEV. Real Jev is queued, not run.

ArmHeld-out Spearman95% author-cluster interval
Registered typed-feature scorer0.4251[0.2495, 0.5838]
Followers only0.0174[-0.2353, 0.3019]
Registered shuffled scores-0.0809[-0.2913, 0.1241]
Additional shuffled training labels0.0308[-0.1779, 0.2282]

Source: analysis.json. The added label control is not preregistered. These are conditional estimates on a small, selected corpus; see the report for feature errors and limits.

Run the free collection

Run these commands from the repository root. Point RUN_DIR outside the repository and use an existing OpenCode CLI installation. The route requires no API key. Protocol registration must already be committed before acquisition.

RUN_DIR=/path/to/external/viral-predictor
OPENCODE_BIN=/path/to/existing/opencode

nice -n 10 python3 experiments/viral-predictor/cli.py collect \
  --opencode "$OPENCODE_BIN" --scratch "$RUN_DIR/oc" \
  --ledger "$RUN_DIR/calls.jsonl" --output "$RUN_DIR/features.json"

Calls run sequentially, first through space-bunny-free, then mimo-v2.6-flash-free only after a transport or schema failure. Each call has a 180-second deadline. The shared-host memory guard waits for usage below 75% before launch and aborts active calls above 85%; its wait is recorded separately from latency. Refusals, invalid cost, and nonzero reported cost stop the campaign.

The append-only ledger retains the 150-call ceiling across restarts, including interrupted attempts. Reuse the same ledger to resume; completed cells are replayed. Six predetermined public posts exercise model-authored explanations. The completed corpus run is retained; no new acquisition is needed for a recount. The current corpus pin differs only in unused source-credit metadata, verified by corpus-binding.json. Original receipt and registration hashes remain unchanged; future acquisition under the re-pinned protocol needs a new ledger. Raw recordings and runtime state remain under the external paths. Output files must be new paths; existing output is never overwritten.

Score a private draft list

Supply an external JSON list containing objects with exactly id and text. The output includes every input ID: successful rows contain score, category, up to three supported reasons, explanation provenance, and nearest reference IDs; feature failures have a null score and an explicit failure status.

nice -n 10 python3 experiments/viral-predictor/cli.py score \
  --posts "$RUN_DIR/drafts.json" --output "$RUN_DIR/ranked.json" \
  --features experiments/viral-predictor/results/features.json \
  --opencode "$OPENCODE_BIN" --scratch "$RUN_DIR/oc" \
  --ledger "$RUN_DIR/calls.jsonl"

Input texts, private IDs, raw calls, and ranked output stay outside the repository. Text-only drafts supply no attachment metadata, so has_visual is false with that uncertainty recorded. The model's explanation receives only validated checks and category. Its vocabulary and reason codes are bounded; rejected explanations fall back to labelled deterministic wording.

Interpret the score

For each query, the five nearest training posts minimize Boolean Hamming distance plus two for a different category. Source order breaks distance ties; engagement counts never choose the neighbors. Published post and neighbor references are SHA-256 hashes of text. The original tie keys are restored only in memory, preserving the registered algorithm and seeded control ordering.

Let y = log1p(likes). The raw prediction combines the neighbor mean with a category mean shrunk toward the overall training mean:

category_estimate = (sum(category training y) + 5 * mean(training y))
                    / (category training count + 5)
raw_prediction = 0.6 * mean(neighbor y) + 0.4 * category_estimate
score = 100 * (training targets below raw_prediction
               + 0.5 * training targets equal to raw_prediction) / training count

All fitting and percentile calibration use training rows only. Floating-point ties use relative and absolute tolerances of 1e-12. With fewer than five training posts, all available references are used. An unseen category falls back to the training global mean. A score expresses position against that training corpus; it is neither a success probability nor a forecast of likes.

Reproduce evaluation offline

The public sanitized feature file contains opaque hashes, checks, category, outcomes, author-group hashes, and preassigned folds needed for evaluation. This command recomputes correlations and intervals without a model call or the external raw ledger:

nice -n 10 python3 experiments/viral-predictor/publication.py recount

To verify original raw receipts and reproduce acquisition accounting, use the CLI with the matching external ledger. Sanitized receipt fields are retained in the committed feature artifact:

nice -n 10 python3 experiments/viral-predictor/cli.py evaluate \
  --features "$RUN_DIR/features.json" --ledger "$RUN_DIR/calls.jsonl" \
  --output "$RUN_DIR/evaluation.json"

Every author's posts occupy the same fold, assigned as int(sha256(author UTF-8), 16) % 5. Predictions for each fold use only the other folds. Spearman correlation uses average ranks for ties. The followers baseline uses the same eligible rows; the shuffled control permutes held-out scores once with seed 20260925 and is descriptive, not a p-value. A separately labelled, unregistered shuffled-label control permutes training likes within each fold using seed 20260925 + fold, refits once, and scores the unchanged held-out likes. It is a diagnostic, not a significance test.

The 95% interval uses 2,000 paired author-cluster bootstrap resamples with seed 20260925, preserving score/target pairs and all posts from each selected author. Models are not refitted. The interval is conditional on these fitted models and corpus; it does not capture repeated feature-call variability or selection bias. Constant-rank resamples have no defined correlation and are excluded with their counts reported. Missing or invalid feature rows are reported in coverage, never converted into negative check answers. A fold without available training rows is explicitly unavailable.

This is a small, selected corpus with dependent posts and limited authors. The earlier sample's saved correlation used different folds and features, so it is context rather than a controlled comparison.

Pending Jev checks

The same checks are queued in pending_jev.jsonl as opaque text hashes with current source and contract pins. No live Jev call is part of this measurement. The updated queue uses durable ledger schema 2; older ledgers fail closed and must not be reused or discarded to bypass budgets. To generate a fresh queue:

nice -n 10 python3 experiments/viral-predictor/cli.py queue \
  --output "$RUN_DIR/pending_jev.jsonl"

After quota is available, the following command can drain the checked-in queue using an already configured free-client credential file. Keep the results external and retain the same results file across invocations:

JEV_PROVIDER=zen JEV_MAX_USD=0 JEV_MAX_CALLS=84 \
nice -n 10 python3 experiments/viral-predictor/jev_queue.py drain \
  --queue experiments/viral-predictor/pending_jev.jsonl \
  --source samples/post-virality/fixtures/posts.jsonl \
  --results "$RUN_DIR/viral-jev-drain.jsonl" --max-calls 84 --batch-calls 4

The drain uses the existing strict free-client guards, memory admission, and a 30-second HTTP timeout. Its durable call ceiling survives restarts; the batch limit caps one invocation. A refusal, error, or unresolved reservation stops automatic retries. Results from these future calls must be labelled separately from the free-chat measurement.

Offline checks
nice -n 10 python3 -m unittest discover -s tests -p 'test_viral*.py'
nice -n 10 scripts/check.sh

The synthetic scoring tests exercise target independence, author isolation, neighbor ordering, known tied-rank statistics, coverage, and cluster-bootstrap behavior. Model and queue tests use local fixtures and mocked clients.

Raw result files
OpenCode Zen free chat via CLI - NOT JEVresults/analysis.json
{
 "acquisition": {
  "attempts": 97,
  "excluded_interruptions": 5,
  "known_reported_cost_usd": 0,
  "median_call_latency_s": 19.583069000000002,
  "memory_wait_s": 630.2136670000006,
  "models": {
   "mimo-v2.6-flash-free": 3,
   "space-bunny-free": 94
  },
  "status_counts": {
   "aborted_memory_pressure": 5,
   "ok": 86,
   "schema_error": 6
  },
  "tokens": {
   "cache_read": 201737,
   "cache_write": 0,
   "input": 583259,
   "output": 8952,
   "reasoning": 146410
  },
  "unknown_cost_receipts": 5
 },
 "additional_shuffled_label_control": {
  "bootstrap": {
   "conditional": true,
   "confidence": 0.95,
   "limitation": "fixed permuted fits; no repeated permutations, refitting, or corpus selection correction",
   "method": "percentile bootstrap of author clusters over fixed held-out score/target pairs",
   "paired": true,
   "performed_resamples": 2000,
   "refit": false,
   "samples": 2000,
   "seed": 20260925,
   "undefined_policy": "exclude constant-rank resamples; report counts; interval conditional on valid resamples",
   "unit": "author_group"
  },
  "coverage": {
   "complete_feature_rows": 84,
   "evaluated_author_groups": 57,
   "evaluated_fraction": 1.0,
   "evaluated_rows": 84,
   "excluded": [],
   "total_rows": 84
  },
  "folds": [
   {
    "fold": 0,
    "permutation_seed": 20260925,
    "test_groups": 9,
    "test_rows": 12,
    "train_groups": 48,
    "train_rows": 72
   },
   {
    "fold": 1,
    "permutation_seed": 20260926,
    "test_groups": 8,
    "test_rows": 11,
    "train_groups": 49,
    "train_rows": 73
   },
   {
    "fold": 2,
    "permutation_seed": 20260927,
    "test_groups": 13,
    "test_rows": 19,
    "train_groups": 44,
    "train_rows": 65
   },
   {
    "fold": 3,
    "permutation_seed": 20260928,
    "test_groups": 15,
    "test_rows": 23,
    "train_groups": 42,
    "train_rows": 61
   },
   {
    "fold": 4,
    "permutation_seed": 20260929,
    "test_groups": 12,
    "test_rows": 19,
    "train_groups": 45,
    "train_rows": 65
   }
  ],
  "is_permutation_test": false,
  "label": "DESCRIPTIVE SHUFFLED-LABEL CONTROL - NOT PREREGISTERED",
  "method": "permute training likes within each fold; fit once; score original held-out likes",
  "metrics": {
   "shuffled_label_control": {
    "ci95": [
     -0.17785803739516706,
     0.2281988949542269
    ],
    "spearman": 0.03079560646366433,
    "undefined_bootstrap_resamples": 0,
    "valid_bootstrap_resamples": 2000
   }
  },
  "permutation_unit": "individual training post; Random(seed + fold).sample in ID order",
  "permutations_per_fold": 1,
  "predictions": [
   {
    "category": "Launch",
    "fold": 2,
    "followers": 13194,
    "followers_log1p": 9.487593248937424,
    "group": "befc92b3a019d04cffb46eaec3c6dca8befcdd8cdaf0410fad67dd7d9c04bb65",
    "id": "81e1978fbd05bc40ab9e54593e4a533d91698219a01242403e2d0877d683b1f2",
    "likes": 646,
    "nearest_distances": [
     3,
     4,
     4,
     4,
     5
    ],
    "nearest_ids": [
     

… truncated, 2986 of 153301 bytes shown. Complete safe artifact

OpenCode Zen free chat via CLI - NOT JEVresults/features.json
{
 "acquisition": {
  "attempts": 97,
  "excluded_interruptions": 5,
  "known_reported_cost_usd": 0,
  "median_call_latency_s": 19.583069000000002,
  "memory_wait_s": 630.2136670000006,
  "models": {
   "mimo-v2.6-flash-free": 3,
   "space-bunny-free": 94
  },
  "status_counts": {
   "aborted_memory_pressure": 5,
   "ok": 86,
   "schema_error": 6
  },
  "tokens": {
   "cache_read": 201737,
   "cache_write": 0,
   "input": 583259,
   "output": 8952,
   "reasoning": 146410
  },
  "unknown_cost_receipts": 5
 },
 "acquisition_protocol_sha256": "568c7efd27b3d6bc63bbb47a54b1c62a5e9a31eabe7eb3596dfcb12988b989a8",
 "call_receipts": [
  {
   "call_id": "748f1db5388d505d0dd39cf6",
   "excluded_from_scoring": false,
   "label": "OpenCode Zen free chat via CLI - NOT JEV",
   "latency_s": 18.23062,
   "memory_wait_s": 0.000127,
   "model": "space-bunny-free",
   "phase": "features",
   "post_sha256": "81e1978fbd05bc40ab9e54593e4a533d91698219a01242403e2d0877d683b1f2",
   "prompt_sha256": "6afc0240141cd5d8f3f22d4ee23339e72b3710782d6f4e5d86a1efa697771263",
   "raw_text_sha256": "5bdcf1fe10f85be4dc9a18e76b93095c2852cb98e5851b687287045e680ebd41",
   "reported_cost_usd": 0.0,
   "sequence": 1,
   "status": "ok",
   "timestamp": "2026-09-25T08:54:08Z",
   "tokens": {
    "cache_read": 1929,
    "cache_write": 0,
    "input": 6477,
    "output": 98,
    "reasoning": 872
   }
  },
  {
   "call_id": "8c0c87ecc66d1b92c85654a5",
   "excluded_from_scoring": false,
   "label": "OpenCode Zen free chat via CLI - NOT JEV",
   "latency_s": 61.964169,
   "memory_wait_s": 7.9e-05,
   "model": "space-bunny-free",
   "phase": "features",
   "post_sha256": "a2ee5aa7b62e056fe6cf21ff4ca517870ca87c60bb381bdf8877a9d5cddb91b1",
   "prompt_sha256": "6532ac06fe5b0105cc47f4f31e8c8b1d16776a94bd5c0af83b011b149a6d9450",
   "raw_text_sha256": "ec2153c58019ebaf514febb35040e1ba240fb76840999c7bfb7e38e6773a2f5b",
   "reported_cost_usd": 0.0,
   "sequence": 2,
   "status": "ok",
   "timestamp": "2026-09-25T08:54:26Z",
   "tokens": {
    "cache_read": 2272,
    "cache_write": 0,
    "input": 6152,
    "output": 99,
    "reasoning": 5661
   }
  },
  {
   "call_id": "05757d8847149e84afe3c981",
   "excluded_from_scoring": false,
   "label": "OpenCode Zen free chat via CLI - NOT JEV",
   "latency_s": 19.402567,
   "memory_wait_s": 0.000131,
   "model": "space-bunny-free",
   "phase": "features",
   "post_sha256": "a5dda5c0ca90e39170a7ddf282b96e34571b8fa19f97818c6920b0ccbf953fc2",
   "prompt_sha256": "f877742967256205a14f78a6b97fb0d626f4a8fabb3b80d05074b20d54a58262",
   "raw_text_sha256": "b0466bf9e3aeaad207dbd8df597003d03832cf6ac4122e12df64560ef97d0333",
   "reported_cost_usd": 0.0,
   "sequence": 3,
   "status": "ok",
   "timestamp": "2026-09-25T08:55:28Z",
   "tokens": {
    "cache_read": 2273,
    "cache_write": 0,
    "input": 6217,
    "output": 98,
    "reasoning": 1150
   }
  },
  {
   "call_id": "03c3bf6990092f759af137a1",
   "excluded_from_scoring": false,
   "label": 

… truncated, 2974 of 153559 bytes shown. Complete safe artifact