Audit deepening studies: pre-registered specification
Written 2026-10-03 by lane 2, during the model audit (docs/briefs/model_audit.md, register in docs/research/AUDIT.md), before any result below was computed. The commit that adds this file is the freeze. Changes after results are logged at the bottom with the date and the reason, and the original result is kept. Null results are published as plainly as positive ones.
Three studies:
- D1: does starting valuation forecast Indian equity returns, with inference that survives overlapping windows and small samples, and out of sample?
- D2: uncertainty intervals on every “what happened next” table on the site.
- D3: one ledger of every confirmatory test on the site, corrected for multiplicity across the whole site.
D1. Valuation and subsequent returns, done properly
C1 (site_calculations_spec.md) regressed overlapping 5- and 10-year returns on starting CAPE with Newey-West errors. The audit found three weaknesses (register rows E1–E3): Newey-West with 60 or 120 lags on 200 to 260 monthly observations is badly undersized; a persistent predictor biases the slope in small samples (Stambaugh 1999); and nothing was tested out of sample. D1 replaces C1’s inference; C1’s results stay in its log.
Data
- Returns: Nifty 50 total-return index (from June 1999) as the main target; Nifty 500 TRI as the second (its own valuation series start in 1999). Month-end levels (last session of each completed month).
- Real returns: deflated by all-India CPI (the chained series in
macro_cpi_monthly.cpi_combined_index, base 2024 with NSO back series and IDH splices before 2013). Realised returns are used, so no CPI publication lag applies to the target. - Predictors, each known at month-end t:
log_cape: IIMA’s Sensex CAPE10 (from April 2000; their series does not adjust for NSE’s 2021 earnings switch, which affects only start months after March 2021). Our own consolidated-basis CAPE starts only in 2009, so it is a robustness check, not the main series.log_pe: log trailing P/E on the consolidated basis (pe_consolidated_basis, NSE’s P/E with pre-April-2021 earnings scaled by the switch-day ratio).log_pb: log price to book (no break at the switch).dy: trailing 12-month dividend yield implied by total-return over price-return indices (dividend_yield_tri), in percent.
- Valuation series are matched to the index whose returns are forecast (Nifty 50 for Nifty 50, Nifty 500 for Nifty 500); IIMA’s Sensex CAPE is used for both.
- Lag: predictors enter with a one-month lag (value at the end of month t−1 forecasts the return from the end of month t), as in C1. NSE’s P/E is computed from reported results and published daily, so this is conservative.
- Horizons: h = 1, 3 and 5 years (12, 36, 60 months). Ten years is dropped: the sample holds fewer than three independent 10-year periods.
- Target: annualised log real return over the next h years, in percent a year. Nominal returns are reported beside every real result, unchanged otherwise.
In-sample inference
For each predictor × horizon × index:
- Slope and R² of the overlapping monthly regression.
- Hodrick (1992) 1B standard errors. They are better sized than Newey-West in samples like this (Ang and Bekaert 2007).
- Non-overlapping regressions: for each of the h×12 possible start offsets, the regression on every (12h)-th month. Reported: the median slope and the range. Offsets share data, so their share of negative slopes is not reported as evidence.
- Bias-corrected bootstrap under the null of no predictability (the decisive test).
- Fit an AR(1) to the monthly predictor.
- Draw (return, predictor innovation) pairs jointly by stationary block bootstrap (mean block 12 months) from demeaned monthly real returns and AR(1) residuals. This keeps their contemporaneous correlation, which is the source of the Stambaugh bias.
- Rebuild the predictor from the AR(1) with the bias-corrected coefficient (Kendall’s correction), and rebuild the forward h-year returns from the drawn monthly returns, so returns are unpredictable by construction.
- Run the same overlapping regression on each of 10,000 samples (seed 20261003).
- p-value: the share of null slopes at least as far from zero, in the predicted direction, as the observed slope (one-sided; sign below).
- Bias-corrected slope: observed slope minus the mean null slope.
- Predicted signs: negative for
log_cape,log_peandlog_pb; positive fordy.
Out of sample (Goyal and Welch 2008; Campbell and Thompson 2008)
- At each month-end t from the first date the rule allows, fit the regression only on start months s whose forward window closed by t (s + 12h ≤ t). At least 60 such start months are needed, so the first forecast at 5 years is in 2009 for Nifty 50.
- Forecast the h-year return starting at t.
- Benchmark forecast: the mean of the same closed-window returns (the historical-mean forecast).
- OOS R² = 1 − Σ(actual − model)² / Σ(actual − historical mean)², over every forecast whose outcome is known.
- Clark and West (2007) test of equal accuracy (one-sided), with Newey-West errors at 12h − 1 lags.
- Campbell–Thompson variant: the slope is set to zero when it has the wrong sign, and the forecast is floored at zero nominal. Reported alongside.
- The number of independent h-year periods in the evaluation window is printed beside every OOS statistic.
Hypotheses (family of 12 = 4 predictors × 3 horizons, on Nifty 50; Bonferroni α = 0.05/12 ≈ 0.0042)
- D1-H(p, h): predictor p forecasts the h-year real return with the predicted sign. It passes only if all three hold:
- the bootstrap p-value is below 0.0042;
- the bias-corrected slope keeps the predicted sign;
- the OOS R² is positive.
- Nifty 500 and nominal returns are reported as the same table, not as more hypotheses.
Predictions (written before computing)
- Slopes for
log_cape,log_peandlog_pbwill be negative at 3 and 5 years. - None of the 12 passes at the Bonferroni bar. The bootstrap p-values at 5 years will be far larger than C1’s Newey-West t-statistics implied (C1: t = −5.4 at 5 years).
- OOS R² at 1 year will be within ±5% for every predictor.
log_capewill have the largest in-sample R² at 5 years.
What the page may say
- If nothing passes: “Starting valuation and later returns have moved together in India’s short history, but the relationship is too weak, and the history too short, to tell apart from chance once overlapping windows are accounted for.” The effect size is shown with its bootstrap interval.
- “At today’s valuation” line, shown only as a description. The fitted 5-year real return at today’s CAPE, with an 80% interval from the bootstrap (parameter uncertainty plus residual resampling), labelled as the historical relationship, not a forecast. It is shown only if
log_capeat 5 years has a positive OOS R²; otherwise the page says the historical-mean forecast did better.
D2. Uncertainty on every “what happened next” table
- Where:
- the barometers’ band tables (
barometers_stats.what_next: fear and greed, bull and bear, trend); - C1’s tercile table;
- C2’s after-signal means;
- any later table of forward returns conditional on a reading.
- the barometers’ band tables (
- Interval: a 90% interval for each band’s median forward return and hit rate, by moving-block bootstrap over sessions (block length = the horizon h, 2,000 draws, seed 20261003). Bands are recomputed in each draw from the resampled readings.
- Difference from the unconditional row: a 90% interval for (band median − all-sessions median) from the same draws. The table says “inside noise” when that interval contains zero.
- Independent periods stay beside every row. A band with fewer than 5 independent periods shows its numbers greyed and the words “too few to read”.
- No hypothesis is tested here. D2 changes presentation only, and no band is singled out after the intervals are seen.
- Prediction: most band medians at 63 and 252 sessions will be “inside noise”.
D3. One multiple-testing ledger for the site
- What it is:
docs/research/AUDIT.mdgains a table of every confirmatory hypothesis tested on the site to date, with its own family’s α and p-value (or the pass bar it was graded on). Sources:trend_momentum_spec.md,trend_barometer_spec.md,market_mood_spec.md,momentum_barometer_spec.md;site_calculations_spec.mdC1–C3;portfolio_lab_v2_spec.mdP1–P6,sip_studies_spec.mdS1–S5;ipo_spec.md,baf_equity_spec.md;- D1 above.
- Rule:
- Holm–Bonferroni across every confirmatory p-value on the site at familywise 0.05, and Benjamini–Hochberg at a false-discovery rate of 0.10.
- Hypotheses graded only by a descriptive bar (no p-value) are listed but not corrected.
- What a page may say: a result that passes its own family but fails site-wide Holm keeps its pass. Its page adds one sentence: “Across the N tests on this site, this result would not survive a site-wide correction.”
- Prediction: at most one hypothesis on the site survives site-wide Holm (MO-H1, the stock momentum premium).
Results log
Implementation notes, 2026-10-03, written before the first run
- D1 target units: annualised log returns in percent (100/h times the sum of monthly log returns), so the Hodrick sums and the bootstrap work in additive units.
- D1 Campbell–Thompson floor for real returns: “zero nominal” becomes minus the trailing 12-month CPI inflation known at the forecast date (CPI lagged one month).
- D1 sample: each predictor uses months from its own first value. Gaps of one month in a predictor are carried forward; a longer gap ends the sample at the gap.
- D1 non-overlapping slopes: need at least 3 points per offset. At 5 years most offsets have 4 or 5, so they are shown and not interpreted.
- D1 seeds: 20261003 plus a fixed offset per predictor, basis and index, so each regression draws its own null samples.
- D2 for C2: breadth-thrust signals are discrete events at least 60 sessions apart, so the interval for their mean resamples the events (with replacement, 2,000 draws), not sessions. Every band table uses the moving-block bootstrap over sessions as written.
- Audit fixes before D1 runs: the earnings yield is on the consolidated P/E basis (register V3). This changes no D1 input except the P/E, which was always meant to be consolidated.
2026-10-03: first run (data to 1 Oct 2026; CPI to Aug 2026; IIMA CAPE to Aug 2026)
Code: compute/valuation_forecast.py (D1), compute/barometers_stats.band_intervals and compute/evidence._tercile_ci (D2). Tables: .cache/derived/audit_d1_valuation_forecast.parquet; bundle evidence/valuation_forecast.
D1, Nifty 50, real (the 12 hypotheses; Bonferroni bar 0.0042).
| Predictor | Horizon | Independent periods | Slope | Hodrick t | Null mean slope | Bias-corrected slope | Bootstrap p | OOS R² | Clark–West p | Passes |
|---|---|---|---|---|---|---|---|---|---|---|
| log CAPE | 1y | 26.2 | −27.0 | −1.6 | −18.7 | −8.4 | 0.25 | −7.1% | 0.22 | no |
| log CAPE | 3y | 8.8 | −24.0 | −1.6 | −14.7 | −9.3 | 0.19 | +28.0% | 0.10 | no |
| log CAPE | 5y | 5.2 | −24.6 | −2.5 | −12.4 | −12.2 | 0.055 | +39.2% | 0.12 | no |
| log P/E | 1y | 27.0 | −33.0 | −2.1 | −18.9 | −14.1 | 0.22 | −26.0% | 0.04 | no |
| log P/E | 3y | 9.0 | −18.2 | −1.2 | −10.4 | −7.8 | 0.29 | −142.5% | 0.08 | no |
| log P/E | 5y | 5.4 | −13.1 | −0.9 | −7.7 | −5.4 | 0.32 | −127.2% | 0.09 | no |
| log P/B | 1y | 27.0 | −49.8 | −2.0 | −22.7 | −27.1 | 0.094 | 0.0% | 0.13 | no |
| log P/B | 3y | 9.0 | −29.9 | −1.5 | −11.5 | −18.4 | 0.064 | +3.6% | 0.12 | no |
| log P/B | 5y | 5.4 | −26.2 | −2.0 | −8.9 | −17.3 | 0.025 | −7.2% | 0.16 | no |
| Dividend yield | 1y | 26.1 | +18.6 | +1.5 | +2.9 | +15.7 | 0.061 | +1.7% | 0.15 | no |
| Dividend yield | 3y | 8.7 | +14.1 | +1.3 | +0.5 | +13.6 | 0.065 | +48.6% | 0.045 | no |
| Dividend yield | 5y | 5.2 | +4.6 | +0.4 | −0.2 | +4.8 | 0.27 | −21.6% | 0.82 | no |
- What the columns mean:
- Slope: points of annualised real return per unit of the predictor. For log CAPE, −24.6 at 5 years means a 10% higher CAPE went with about 2.4 points a year less.
- Out-of-sample evaluation: 72 to 120 monthly forecasts, which is 3.2 to 20 independent periods.
- Prediction: none of the 12 passes. Held. The closest are log P/B at 5 years (p = 0.025, but negative OOS R²) and log CAPE at 5 years (p = 0.055, OOS R² +39%, Clark–West p 0.12, on 3.3 independent periods).
- Prediction: bootstrap p-values far larger than C1’s Newey-West t implied. Held. For CAPE at 5 years: C1’s t of −5.4 against a bootstrap p of 0.055.
- About half of the in-sample CAPE slope is small-sample bias. Under no predictability at all, the simulated slope averages −12.4, against the observed −24.6. This is the Stambaugh bias of a persistent predictor, and C1 did not account for it.
- Prediction: OOS R² within ±5% at 1 year for every predictor. Failed: −7.1% (CAPE) and −26.0% (P/E) are worse than that. P/B and dividend yield are inside.
- Prediction: log CAPE has the largest in-sample R² at 5 years. Held: 0.72, against P/B 0.50, P/E 0.17 and dividend yield 0.06.
- Trailing P/E is the worst predictor out of sample at every horizon (OOS R² −26% to −143%): one-year earnings swing more than prices.
- Nifty 500 and nominal returns (reported, not tested) look the same: no bootstrap p below 0.025, CAPE OOS R² at 5 years +34% (real) and +20% (nominal).
- “At today’s valuation” line (allowed, because CAPE’s 5-year OOS R² is positive):
- At the August 2026 Sensex CAPE of 29.0, the historical relationship implies 1.8% a year real over five years, with an 80% interval of −4.0% to +7.1%.
- It is a description of 2000–2021 history with five independent periods, not a forecast.
D2, intervals on “what happened next”.
- Prediction: most band medians at 63 and 252 sessions are inside noise. Held: 21 of 30 band-horizon cells across the trend barometer, fear and greed, and bull and bear.
- Outside noise:
- Fear and greed, extreme greed: followed by higher 63- and 252-session returns than all sessions (+2.5 to +6.3 points at 63 sessions). This is the opposite of the folk reading, and consistent with FG-H4b’s failure.
- Fear and greed, fear (63 sessions) and neutral (252): slightly lower.
- Trend barometer, 40–60 band (63 and 252) and 20–40 (21): lower.
- Bull and bear: three cells, two of them with 1 to 3 independent periods, flagged “too few to read”.
- C1 terciles: every tercile has under 5 independent periods, so all are flagged “too few”. The CAPE tercile intervals overlap (Sensex 5-year: cheapest 8.9–29.2%, dearest −4.7% to +6.6%).
- C2: the mean 6-month return after a breadth thrust is 14.9% with a 90% interval of 4.1–25.0% (11 events), against 8.8% for all sessions.
D3, the site-wide ledger. 35 confirmatory tests (table in docs/research/AUDIT.md).
- Prediction: at most one survives site-wide Holm (MO-H1). Failed. Three survive:
- MO-H1 (p ≈ 1e-6);
- C3’s two “returns predict next week’s flows” tests (DII and FPI, p < 1e-6).
- Benjamini–Hochberg at 10%: adds FG-H4a (extreme fear → higher 3-month returns, p = 0.0061).
- Nothing that predicts market returns survives Holm. Every positive claim about forecasting the index rests on FG-H4a alone, which survives only the softer false-discovery reading.