What 52,777 replay rows said about stock selection versus QQQ

The expanded model produced positive average raw returns, but every tested horizon still had negative average excess return versus QQQ.

V4.67.1120 signal datesNext-session-open entry
Bottom line: this replay did not establish that the stock-screening model beat a passive QQQ benchmark. Publishing the negative comparison is more useful than promoting the positive raw-return column alone.

Research design

Replay window2026-02-03 to 2026-07-27, 120 U.S. trading dates
Universe requested516 stocks; 514 had usable daily bars
Total signal rows52,777 across seven controlled experiments
Unique signaled stocks483
Entry assumptionNext session open after daily discovery
Holding horizons1, 3, 5, 10 and 20 sessions
Benchmark basisSame entry-to-horizon-close return for QQQ

Strict filter versus expanded filter

The strict control generated 3,881 rows. The expanded V4.67.1 experiment generated 19,788. Expansion increased coverage and average raw returns at longer horizons, but not benchmark-relative performance.

Holding horizonStrict raw returnStrict excess vs QQQExpanded raw returnExpanded excess vs QQQ
1 session-0.02%-0.08%0.00%-0.05%
3 sessions+0.06%-0.23%+0.16%-0.12%
5 sessions+0.19%-0.20%+0.25%-0.20%
10 sessions+0.29%-0.58%+0.48%-0.50%
20 sessions+0.14%-0.91%+0.65%-0.97%

What changed our view

A positive average stock return is not enough in a rising market. At 20 sessions, the expanded experiment averaged approximately +0.65%, yet lagged QQQ by about 0.97 percentage points on average. More opportunities were found, but the additional signal density was not evidence of alpha.

What remains unresolved

Practical conclusion

The evidence supports using broad ETFs as the default benchmark and treating individual-stock signals as an unproven research overlay. It does not support presenting the model as a reliable replacement for long-term index investing.

Research record

Source: V4.67.1 dated study record and ablation summary generated 2026-08-25. The public table rounds source decimals to two percentage points for readability.

Research record: 120 dates, 514 tickers with bars, 52,777 replay rows. Results have not been independently audited.