Post ranking association: held-out retrospective evaluation
Evidence boundary. A selected, single-source, retrospective ranking association on 84 posts from 57 authors, recomputed offline from saved feature answers. It is not validated future-post prediction. The raw model-call ledger is excluded from the repository (only sanitized receipts ship), so the data collection cannot be independently checked.
This merged-tree run is an offline recomputation from saved model scores (the previously recorded typed feature answers) using unchanged deterministic scoring code. It makes no new model calls or feature re-scoring. Held-out Spearman is 0.4251, with a 95% author-cluster bootstrap interval of [0.2495, 0.5838], on 84 posts from 57 authors. All 84 saved feature rows are included. These unchanged values describe a ranking association with observed likes in this small sample; they do not establish real-world virality prediction.
| Arm | Held-out Spearman | 95% author-cluster interval |
|---|---|---|
| Registered typed-feature scorer | 0.4251 | [0.2495, 0.5838] |
| Followers only | 0.0174 | [-0.2353, 0.3019] |
| Registered shuffled scores | -0.0809 | [-0.2913, 0.1241] |
| Additional shuffled training labels | 0.0308 | [-0.1779, 0.2282] |
Source: results/analysis.json. The first three rows are under metrics; the added label control is under additional_shuffled_label_control.metrics. Features and sanitized call receipts are in features.json. The fresh recomputation log records the merged-tree run. Every metric and prediction exactly matches the prior saved analysis. It reuses the original feature acquisition and outcomes, providing a reproducibility check of those saved data. The committed scoring inputs originated at b09484e.
Design and controls
Fourteen Boolean checks and one of seven categories feed an unchanged five-neighbor/category-mean scorer. Five folds hold out entire author groups; fold sizes are 12, 11, 19, 23 and 19. Fitting, neighbor selection and percentile calibration use only training rows. Prompts included post text and media flags; engagement, follower-count and author metadata were withheld. Public text can itself mention people or quantities. No prediction uses its own post, author or fold as a neighbor. Ties retain the registered source order even though published references are opaque text hashes.
Spearman uses average ranks for ties against likes; log1p is monotone and does not change those ranks. Each 95% percentile interval uses 2,000 paired resamples of the 57 author clusters with seed 20260925, keeping every selected author's posts together. All 2,000 resamples were valid for every reported metric. The bootstrap holds fitted models and their out-of-fold predictions fixed.
The registered control shuffles held-out scores once with seed 20260925. The additional shuffled-label control was not preregistered: during resume, each fold's training likes were permuted once with seed 20260925 + fold, the same scorer was fitted, and original held-out likes supplied the evaluation target. Neither control is a permutation significance test. No control seed, feature, weight or model was selected by the reported results.
Acquisition and costs
Route: OpenCode Zen free chat via CLI - NOT JEV. The completed acquisition used 97 sequential attempts: 94 Space Bunny and 3 MiMo, with 180-second deadlines. It produced 84 valid feature rows (82 Space Bunny, 2 MiMo), 2 accepted free-model explanations and 4 labelled deterministic explanation fallbacks. Six schema errors remain recorded. Five memory-pressure interruptions consumed budget slots but are excluded from scoring and model-failure accounting.
Launches required memory usage below 75%; active calls aborted above 85%. Recorded admission waiting totals 630.214 seconds and is outside call latency; median non-interrupted call latency is 19.583 seconds. All 92 received cost receipts report $0. Five interrupted calls have no cost receipt, so their cost is unknown. The corpus recount and compaction repeat made zero new model calls. All 84 Jev checks remain queued; real Jev was not run for M5.
Provenance and privacy
Released main changed only unused source-credit metadata in 78 corpus rows. The binding manifest verifies that ordered text, author groups, outcomes and media flags are unchanged while recording original and current digests. Original protocol registration and raw receipt hashes remain intact; RE-PINNED.md records each re-pin. The input commit precedes this recount. Pinned scoring code and current corpus are checked by the read-only reproduction command.
Private registry text and IDs are absent from the repository. Feature, neighbor and queued references use opaque text hashes. Raw call logs stay external. The private-draft CLI was exercised with synthetic offline inputs only; no private-draft prediction accuracy or real private scoring run is claimed.
Limits
This is a small, selected, single-source corpus, not a representative future feed. Authors contribute dependent posts, and the folds are uneven. The interval is conditional on this corpus and these fits: it omits refitting uncertainty, feature-call variability, publication selection and distribution shift. Public post recognition or prior exposure by the free feature model is unmeasured.
Free-model features are imperfect. Direct checks find first-line-digit mismatches in 14/84 posts and length-check mismatches in 2/84; visual-metadata checks agree on all 84. The measured answers were retained, not corrected after seeing outcomes. An evidence reference does not verify its claim, and media flags do not inspect images. The relative 0–100 score is not a calibrated probability or a forecast of likes. The older sample's 0.262 result uses different features and splits, so it is not a controlled comparison.
Reproduce
From the repository root, without model calls or private inputs:
PYTHONDONTWRITEBYTECODE=1 nice -n 10 python3 experiments/viral-predictor/publication.py recount
See the experiment README for private input boundaries, acquisition guards and the separately guarded future Jev drain.