alpha-signal-lab — pre-registered research
Eight pre-registered checks, run against a committed synthetic fixture in CI with no live network calls. Hover or tap a badge for what it verifies.
Identical features, target, model, and hyperparameters. Only the cross-validation changes: PurgedWalkForward (honest) vs naive random 5-fold shuffled CV with no purge or embargo (leaky twin). The gap is pure leakage.
Delta between the two mean rank ICs above, computed from the live numbers once both sides are available.
5-day rebalance, equal-weight top-3 long / bottom-3 short, model vs 12-1 momentum rank vs equal-weight buy-and-hold.
Daily cross-sectional Spearman rank IC, smoothed. The dashed line and label mark the pre-registered 0.02 bar.
One LightGBM model per walk-forward fold; bars show that fold's mean OOS rank IC. Gray bars mean no data for that fold, not a measured zero.
SHAP on the final (most recent) fold's model, computed on that fold's OOS test rows only. Stability heatmap shows fold-to-fold rank churn (1 = most important).
Everything below was computed after the primary pre-registered verdict above was sealed. These are diagnostic/robustness checks, not a re-run or reframing of the primary result.
Does the model add anything orthogonal to plain momentum? Rank IC of raw predictions vs predictions cross-sectionally residualized against 12-1 momentum.
Rank IC split by realized-volatility regime (median split on trailing 21-day realized volatility), plus per-ticker time-series IC across the universe.
Mean |SHAP| per feature, one line per feature, across all walk-forward folds, how importance ranks drift over time rather than a single snapshot.
Combinatorial Purged Cross-Validation: the same frozen model refit across many purged/embargoed train/test path combinations, giving a distribution of OOS rank IC instead of the single walk-forward realization above.
Computed after the primary verdict, from the already-sealed OOS predictions: a diagnostic look at why the model trails the 12-1 momentum baseline, not a re-run or reframing of the primary result.