Look-Ahead Contamination in LLM Risk-Factor Evaluation
DOI:
https://doi.org/10.13021/jssr2026.5637Abstract
Large language models (LLMs) are increasingly used to analyze corporate disclosures and predict stock market outcomes. A working paper by Gao et al. investigates whether an LLM scoring SEC 10-K risk-factor disclosures is genuinely reading the filing or recalling firm-specific information from training. However, this placebo’s validity depends on whether company identity is truly concealed. The paper scores 209 SEC Form 10-K risk factor disclosures twice with Llama-4-Scout (109B): once on the original filing and once on a machine-anonymized version with company name, tickers, dates, and states redacted. It compares each version’s ability to predict 12-month buy-and-hold abnormal returns (BHAR) before and after the model’s approximate mid-2024 training cutoff, since only pre-cutoff filings have outcomes realized before the cutoff and could be in the model's training data. The paper reports no statistically significant evidence of look-ahead contamination: raw and anonymized scores are highly correlated (ρ = 0.65), and neither predicts BHAR in either cutoff window (triple interaction = +0.0297, t = 0.37). This result assumes the anonymization actually conceals firm identity–an assumption the paper never tests. That test is addressed here. Manually comparing 15 machine-anonymized filings against their raw versions, 7 remained identifiable despite automated-redactions, through residual clues including flagship products, executives, subsidiaries, and unique geographic or business descriptions. Because leaky anonymization biases the placebo toward finding no difference between raw and anonymized scores, this finding does not confirm the paper’s null result, but raises the possibility that the null partly reflects leaky anonymization rather than real absence of look-ahead bias. These observations motivate a hand-built gold standard, currently underway, covering 100 filings, paired with a blind human identification test to measure how much leakage the automated process leaves, before the placebo is re-run on cleaner text.


