Calculations across the site: pre-registered specification
Written 2026-10-02 by lane 2, before any of these was computed. Each block states the question, the data, the method, the overlap handling and a prediction to be graded. Changes go under “Amendments” with the date and the reason.
Common rules:
- No look-ahead: every signal uses only data published before the date it is used (one-month lag for monthly macro and valuation series unless stated).
- Overlap: forward returns over h months computed every month overlap by h−1 months. We report the effective number of independent periods (sample months ÷ h), Newey-West (Hansen-Hodrick style) standard errors with h lags, and the same statistic on non-overlapping subsamples (every h-th month, for each of the h possible offsets, with the range across offsets).
- Nulls are results. A calculation that finds nothing is published with the same prominence as one that finds something.
C1. Valuation and subsequent returns (CAPE, earnings yield, yield gap)
- Question: did starting valuation tell you anything about the next 5 and 10 years of Indian equity returns?
- Data: IIMA Sensex CAPE10 (month-end, from April 2000) and Nifty 500 CAPE10 (from April 2006) from
docs/reference/iima_india_cape_*.csv; Nifty 500 trailing earnings yield and the yield gap (earnings yield minus 10-year G-sec) fromvaluation_nifty_500.parquet(1999 on); forward returns from the Nifty 500 TRI, nominal and real (all-India CPI, spliced). - Method: regress the annualised forward 5-year and 10-year real return on log CAPE (and separately on the earnings yield and the yield gap) using the value known at the start month (one-month lag). Report the slope, R², Newey-West t-statistic with 60 or 120 lags, the non-overlapping-offset range of slopes, and a table of median forward returns by CAPE tercile with the number of independent periods in each.
- Prediction C1: the slope on log CAPE is negative at both horizons, but at 10 years the sample holds fewer than three independent periods, so we will not call it significant whatever the t-statistic says.
C2. Breadth thrusts and their record
- Question: do NSE breadth thrusts (Zweig 1986) precede above-normal returns?
- Data: daily advances and declines of NSE stocks from bhavcopy (
breadth_stocks_daily.parquet), Nifty 500 TRI. - Method: the 10-session EMA of advances / (advances + declines) rises from below 0.40 to above 0.615 within 10 sessions; a new signal needs 60 sessions since the last. For each signal, forward 1, 3, 6 and 12-month Nifty 500 returns, against the unconditional distribution of forward returns over the same sample. Significance by a block bootstrap of signal dates (random dates with the same count, mean block 63 sessions).
- Prediction C2: fewer than 15 signals since 1995; the average forward 6-month return after a signal is above the unconditional average, but the bootstrap p-value is above 0.05.
C3. Flows and returns
- Question: do foreign (FPI) and domestic institutional (DII) flows lead market returns, or follow them?
- Data: daily net flows (
flows_daily.parquet, NSDL FPI and exchange-reported DII), weekly sums scaled by Nifty 500 market capitalisation where available; Nifty 500 TRI weekly returns. - Method: cross-correlations of weekly flows and returns at leads and lags of up to 8 weeks, and bivariate Granger tests (lag length by BIC, up to 4) in each direction, on the full sample and two halves. Caveats stated: Granger causality is predictability, not cause; flows and returns are measured with timing mismatches.
- Prediction C3: returns predict next week’s FPI flows (positive, significant) more strongly than flows predict next week’s returns; DII flows respond negatively to past returns (domestic institutions buy dips).
C4. The IPO cycle and market returns
- Question: does a hot IPO market precede weak returns (Baker and Wurgler 2000, the equity share in new issues)?
- Data: monthly IPO count and money raised (
ipo_fundraising_monthly.parquet), Nifty 500 TRI and its market capitalisation proxy. - Method: IPO money raised over the trailing 12 months as a share of market capitalisation, against forward 12 and 36-month Nifty 500 returns; overlap handling as above.
- Prediction C4: a negative relation at 36 months, not significant after accounting for overlap.
C5. Leverage (MTF) and positioning against drawdowns
- Question: does a rising margin-trading book or crowded futures positioning precede drawdowns?
- Data: the MTF book (2004 on) and participant open interest from
leverage(lane 1 tables). - Method: 3-month growth of the MTF book (and its level relative to market capitalisation) against the probability of a 10% Nifty 500 drawdown in the next 3 months (logistic regression) and forward 3-month returns; overlap handling as above.
- Prediction C5: high MTF growth is associated with a higher drawdown probability, with a confidence interval that includes no effect.
C6. SIP outcome distribution
Covered by sip_studies_spec.md.
C7. A survivorship-free factor library from bhavcopy
- Question: what did momentum, low volatility, short-term reversal and size earn in Indian stocks since 1995, net of realistic costs, when delisted stocks are kept?
- Data: split- and bonus-adjusted NSE bhavcopy prices with delistings kept (
equity_prices_adjusted.parquet). Value and quality need fundamentals the warehouse does not hold yet; size needs shares outstanding (request to lane 1). - Method: monthly rebalanced long-only top-quintile and long-short quintile portfolios within the top 500 stocks by traded value (formation at month-end, trade at the next close): momentum 12-1 (Jegadeesh and Titman 1993), 252-day volatility (lowest quintile), 1-month reversal; equal and liquidity-capped weights; costs 0.3% each way for the top 200 and 0.6% beyond, plus impact proportional to trade size over traded value. Delisted stocks: last traded price, then a −30% delisting return for compulsory delistings where the reason is known, 0% otherwise, stated.
- Prediction C7: momentum’s long-short premium is positive before costs and roughly halves after costs; low volatility has a higher Sharpe ratio than the market with a lower return; short-term reversal does not survive costs.
C8. Fear & Greed 1.1.0 with its null results
- Question: port the Aftermarket Report’s Fear & Greed 1.1.0 (NSE-only inputs) onto Data bank inputs, and publish its pre-registered null results next to the gauge.
- Method: as specified in the AMR documentation (
FEAR_GREED_RISK_RESULTS.md), read before porting; the port is checked by reproducing AMR’s published values over their overlap to within rounding. - Prediction C8: the gauge does not predict 1, 3 or 6-month forward returns (as AMR found).
Order of work
C1 first (data ready, small), then C2 and C3 (data ready), C4, C5 (after lane 1’s leverage tables settle), C8 (needs the AMR spec), and C7 (largest; needs a data request for shares outstanding and fundamentals).
Amendments
2026-10-02, before C2 and C3 were computed (data constraints found while wiring them):
- C2 sample: the bhavcopy breadth universe holds fewer than 200 stocks a day before 2004 (9 in 1995, about 100 in 2000), so an advance/decline ratio there is not a market breadth measure. The C2 sample starts on the first day the 60-session median universe reaches 200 stocks. The prediction’s “since 1995” becomes “since that date”.
- C2 significance: the block bootstrap of signal dates is done as random circular shifts of the whole set of signal dates (10,000 shifts), which keeps the signals’ spacing and clustering; the p-value is the share of shifted sets whose mean forward return is at least the observed one.
- C3 data: the warehouse holds daily FPI equity flows only from January 2021 (DII from April 2007), and no Nifty 500 market-capitalisation series. Weekly flows are therefore scaled by their own trailing 52-week standard deviation (known before the week), and the FPI tests cover 2021 onward only. A request for NSDL’s FPI history back to 1999 goes to lane 1.
Results log
2026-10-02: C1, C2 and C3 (first runs, rules as specified and amended above)
Code: pipeline/tipsheet/compute/evidence.py; tables evidence_c1_*, evidence_c2_*, evidence_c3_* in .cache/derived.
C1. Valuation and subsequent real returns (Nifty 500 TRI, CPI-deflated).
| Signal | Horizon | Start months | Independent periods | Slope | R² | Newey-West t | Non-overlapping slopes (min / median / max; share negative) |
|---|---|---|---|---|---|---|---|
| Sensex CAPE10 (log) | 5 years | 2000-05 to 2021-08 | 5.3 | −31.7 | 0.69 | −5.4 | −46.5 / −28.5 / +18.7; 90% (4 points each) |
| Sensex CAPE10 (log) | 10 years | 2000-05 to 2016-08 | 2.6 | −11.3 | 0.76 | −25.1 | too few points (fewer than 3 per offset) |
| Nifty 500 CAPE10 (log) | 5 years | 2006-05 to 2021-08 | 4.1 | −14.2 | 0.35 | −3.4 | −51.3 / −6.6 / +112.5; 65% |
| Nifty 500 CAPE10 (log) | 10 years | 2006-05 to 2016-08 | 2.0 | −10.0 | 0.77 | −14.5 | too few points |
| Nifty 500 earnings yield | 5 years | 1999-02 to 2021-08 | 5.5 | +2.6 | 0.13 | +1.3 | −3.9 / +0.8 / +9.6; 40% |
| Nifty 500 yield gap | 5 years | 1999-02 to 2021-08 | 5.5 | +1.5 | 0.08 | +1.8 | −22.8 / +0.9 / +20.0; 33% |
| Nifty 500 yield gap | 10 years | 1999-02 to 2016-08 | 2.8 | −0.1 | 0.00 | −0.3 | too few points |
By Sensex CAPE tercile, the median real 5-year return was 12.8% a year from the cheapest third (CAPE 12 to 20), 7.9% from the middle and 2.3% from the dearest (CAPE above 24.4); each tercile holds about 1.4 independent 5-year periods. Tercile boundaries use the whole sample, so the table describes history and is not a trading rule.
- Prediction C1 holds. The CAPE slope is negative at both horizons. At 10 years the sample has 2.6 independent periods, so the t-statistic of −25 says nothing: overlapping windows make one or two episodes look like 196 observations. We do not call it significant.
- The 5-year CAPE result is the more informative one: nine in ten non-overlapping subsamples give a negative slope. Even there, a single valuation cycle (2007-08 peaks followed by weak returns, 2003 and 2009 troughs followed by strong ones) drives most of it.
- The earnings yield and the yield gap carry little information at either horizon: R² of 0.13 and 0.08 at 5 years, slopes that change sign across subsamples, and nothing at 10 years. Trailing earnings are noisier than ten-year average earnings, and the yield gap mixes in the bond yield, which moved for its own reasons.
C2. Breadth thrusts (sample from 2003-09-11, when the breadth universe reached 200 stocks).
Twelve signals: 2004-04, 2005-02, 2006-06, 2007-09, 2008-07, 2008-12, 2009-07, 2017-01, 2020-04, 2024-04, 2025-04 and 2026-04.
| Horizon | Signals | Mean after signal | Unconditional mean | Share positive after signal (unconditional) | p (circular-shift test) |
|---|---|---|---|---|---|
| 1 month | 12 | 2.7% | 1.4% | 75% (64%) | 0.25 |
| 3 months | 12 | 4.8% | 4.4% | 67% (65%) | 0.41 |
| 6 months | 11 | 15.5% | 8.8% | 82% (69%) | 0.09 |
| 12 months | 11 | 30.9% | 18.3% | 91% (78%) | 0.02 |
- Prediction C2 holds: fewer than 15 signals, a higher average 6-month return after a signal, and a p-value above 0.05 at 6 months.
- The 12-month result (p = 0.02) was not part of the prediction and leans on two signals that came straight after crashes (December 2008 and April 2020). With 11 events and four horizons tested, it is a lead worth tracking, not a finding. The July 2008 signal was followed by a 42% fall over three months.
C3. Flows and returns (weekly; DII from October 2007, FPI from July 2021).
| Flow | Sample | Flows predict next returns (F, p) | Returns predict next flows (F, p) | Sum of lagged-return coefficients |
|---|---|---|---|---|
| DII | full | 1.1, 0.33 | 37.2, <0.001 | negative |
| DII | first half / second half | p = 0.36 / 0.29 | p < 0.001 in both | negative in both |
| FPI equity | full | 0.8, 0.36 | 44.7, <0.001 | positive |
| FPI equity | first half / second half | p = 0.29 / 0.62 | p < 0.001 in both | positive in both |
- Prediction C3 holds. Returns predict flows; flows do not predict returns. Foreign investors chase: the correlation of a week’s FPI flow with the previous week’s return is 0.47 and with the same week’s 0.33. Domestic institutions lean against the market: DII flows correlate −0.36 with the previous week’s return and −0.28 with the same week’s.
- Neither flow has a correlation above 0.13 in absolute value with any later week’s return (one to eight weeks ahead).
- Caveats: Granger tests measure predictability, not cause; the FPI sample is only five years; and the weekly sums mix trade dates and reporting dates.
2026-10-03: audit corrections (register docs/research/AUDIT.md, rows E1–E4, V3, V7)
- C1, earnings-yield signal (V3, V7). It now uses the consolidated-basis P/E, read from the maintained valuation table. Rescaling pre-2021 earnings by the switch factor scales the signal, so the slope falls and the t-statistic does not change:
- 5 years: slope +2.64 → +2.17, t 1.27 (unchanged);
- 10 years: +1.10 → +0.92.
- The yield gap: 5-year slope +1.48 → +1.44 (R² 0.08 → 0.10).
- CAPE rows are unchanged (they use IIMA’s series).
- C1, inference (E1): superseded.
- What C1’s inference got wrong: Newey-West errors with 60 or 120 lags on 184 to 271 months are badly undersized. The persistent predictor’s small-sample (Stambaugh) bias was not measured. Adjacent non-overlapping offsets share 59 of 60 months, so “nine in ten subsamples negative” is not nine in ten pieces of evidence.
- What replaces it: D1 (
audit_deepening_spec.md) redoes the question properly. Its 5-year CAPE result has a bias-corrected bootstrap p of 0.055 (C1’s Newey-West t was −5.4), with about half the slope being small-sample bias, and an out-of-sample R² of +39% on 3.3 independent periods. - Which statements stand and which do not: C1’s prediction (a negative slope, not called significant at 10 years) still holds. Its sentence that “the 5-year CAPE result is the more informative one: nine in ten non-overlapping subsamples give a negative slope” overstates the evidence and is withdrawn.
- C1, terciles and scatter (E2, E3). The tercile table now carries 90% moving-block intervals (D2). Every tercile is flagged “too few” (under 5 independent periods). The scatter rows carry a
nonoverlapflag marking one non-overlapping set. - C2, entry timing (E4). Breadth is known only after the signal day’s close, so forward returns now start at the next close.
- Changes: 1-month mean 2.7% → 2.0% (p 0.25 → 0.38); 3-month 4.8% → 4.7% (p 0.41 → 0.43); 6-month 15.5% → 14.9% (p 0.09 → 0.11); 12-month 30.9% → 30.7% (p 0.02 → 0.03).
- The 6-month mean now carries an interval: 14.9% (90% interval 4.1–25.0%, 11 events) against 8.8% for all sessions.
- Prediction C2 holds. The 12-month result stays exploratory.
- C3: unchanged.