Alt-data signals at institutional scale
Small panels left the question open: were the odd results noise, or something real? So we scaled the same test to 500 names, where even a one-percent-a-year effect becomes measurable, and let the statistics settle it.
Why scale up
The sign tests in Note 02 ran on hand-picked 10-name panels. They were suggestive but underpowered, and two results looked backwards relative to the textbook. Ten names can't distinguish a small real effect from noise, but five hundred can. At institutional breadth, with tens of thousands of position decisions per test, the statistics can detect effects near 1% a year. A null result at that power means the effect is genuinely small, not that the test was weak.
Each stock is judged against roughly a year of its own signal history: long in its top quartile, short in its bottom quartile, position sizes capped and equal-weighted per side. Three design questions run as paired comparisons: with and without a trend filter, with and without volatility targeting, and a monthly-refreshed universe versus one frozen at January 2019. Three signals times six configurations: 18 cells.
What breadth settled
The small-panel anomalies dissolve. The strange insider result from Note 02 does not appear at 500 names: insider books land within ±1% a year of zero everywhere. The borrow-fee direction flips back to the textbook reading (expensive-borrow names underperform, weakly). Both oddities were properties of small hand-picked panels, not of the signals.
News stays positive, and stays too small. Positive in all nine of its cells, as in every study in this line. Its best cell reaches +4.0% a year at p = 0.065, short of significance even before discounting for having run an 18-cell grid.
The structural choices barely matter. Volatility targeting rescales return and risk together and changes no conclusion. Freezing the universe changes little; its main effect is telling us that monthly refresh churns the news books without adding information.
Annualised long-short returns before borrow costs, p-values from Newey–West standard errors on each book's return series. No cell of 18 reaches p < 0.05; the two closest are 2 of 36 statistics, consistent with chance.
Making sure the null isn't hiding anything
A fixed-direction book has one blind spot: if a signal's true direction differed stock by stock, the opposing pieces would cancel and the book would read as empty even though information existed. We tested for exactly that, computing each stock's own signal-to-return correlation across the full period and asking whether those correlations are more dispersed than sampling noise. They aren't, for any of the three signals. The share of individually significant names sits at or below the 5% you'd expect by chance. There is no hidden per-name structure for a cleverer book to harvest.
Where this leaves us
This closes the line, at depth and now at breadth. Three free alternative-data signals, expressed as own-history extremes on liquid US names, carry directional information worth at most a few percent a year before costs, and nothing in 18 cells clears the standard significance bar. Any further work should change the signal construction itself, not re-tune this design.
For an allocator, the takeaway is the process: a hypothesis pursued through three escalating studies, statistics powerful enough to make the null informative, robustness checks that rule out the loopholes, and a written verdict. That machinery is the asset.
Period 2019-01-01 to 2026-06-30. Universe: top 500 US names by trailing dollar volume above $5 (top 150 for news, whose event volume is ~5 million headlines per run), monthly refresh or frozen at January 2019. Data: Quiver insider transactions, Tiingo news sentiment, Interactive Brokers borrow fees, point-in-time. Positions capped at 2% per name, re-checked every 10 trading days; volatility targeting at 10% with a 3× leverage cap where enabled. Significance from weekly returns with Newey–West standard errors; CAPM alpha vs SPY; sign-heterogeneity via a slope-homogeneity test on per-name correlations. The books measure signal content, so borrow costs are not modeled. Backtested results are hypothetical; this note is research and is not investment advice.