Audit register: every calculation that produces a claim
Opened: 2026-10-03 by lane 2, for the model audit (docs/briefs/model_audit.md).
Phase 0 (this file, committed before any fix): one row per claim-producing calculation, with the method, the checks run, a verdict and a severity.
After Phase 0: fixes are logged in each spec’s results log (date, reason, old and new numbers) and in docs/status/quant.md. This register gains an “Outcome” column, and its rows are never deleted.
Key
Verdicts:
- sound: survives the checks.
- fix: a defect with a clear correction.
- rethink: the method or framing needs redesign.
- withdraw: the claim should come off the site.
Severities:
- wrong number: a published figure is incorrect.
- misleading: the figure is right, but the framing, label or inference overstates it.
- weak: a defensible result that deserves a caveat or a better method.
- fine: nothing to change.
Checklist used for every row:
- Look-ahead: publication lags, revised data, full-sample percentiles, signal and trade on the same close.
- Survivorship and back-fill: membership, pre-launch history, delistings, proxies.
- Overlapping windows: both the inference and the presentation.
- Multiple testing.
- Effect sizes with uncertainty.
- Costs and taxes: applied consistently and labelled.
- Units and bases: price vs total return, nominal vs real, annualised vs cumulative, the P/E basis break, CPI splices.
- Data quality: bad prints, stale carry-forward, partial periods, duplicates.
Independent re-implementations of headline numbers live in pipeline/tests/audit/, and the rows say which ones are covered.
Who checked what:
- Rows V, E, T and BT were checked in the main audit session.
- Rows L were checked by an independent reviewer agent, and R, I, B and S by another (both read-only, with scratch scripts).
- Rows D (the data layer, owned by lane 1 and others) by a third reviewer agent.
- Numbers come from this worktree’s
.cache(data to 1 October 2026) unless stated otherwise.
Valuation (compute/valuation.py, publish/valuation.py, docs/methods/valuation.md)
| ID | What it claims | Method | Checks run | Verdict | Severity | Evidence |
|---|---|---|---|---|---|---|
| V1 | CAPE on a consistent consolidated basis (22 indices) | Implied EPS = close ÷ P/E. Pre-2021-03-31 EPS scaled by the switch-day ratio. CPI-deflated 120-month mean, from the month before | Splice robustness; CPI splice; deflator cancels; CPI lag | sound | weak | Splice: EPS moved more than 1% on a single day between 15 Mar and 15 Apr 2021, on 03-31 only, for Nifty 50 (×1.205), 500 (×1.200), Bank, IT, FMCG and Next 50. So the switch is a clean single break. Deflator: the latest-CPI level cancels in the ratio. CPI lag: the historical series deflates price by the current month’s CPI, published about the 12th of the next month. The look-ahead is at most one month’s inflation (about 0.5%) on the level; the latest point uses the last published CPI. Basis: the “constant consolidation gap” assumption is untestable and is stated |
| V2 | P/E and P/B percentile of today within own history | Share of history ≤ today, on the consolidated basis | Full-sample vs expanding | sound | fine | Today’s position in history is legitimately full-sample. History starts 1999 (one regime), and the start date is shown |
| V3 | Earnings yield, nominal yield gap, CAPE yield minus real G-sec | 100 ÷ P/E; minus the 10-year yield | P/E basis | fix | wrong number | valuation.py:174 uses raw pe_ratio, so the earnings yield steps up 17–20% on 2021-03-31 (Nifty 500 EY 2.31 → 2.79). The published earnings_yield_pct series and every consumer inherit the break: C1 (E1), the lab’s yield-gap glide and the dynamic SIP (L6) |
| V4 | Dividend yield from total-return over price-return | 250-session TRI growth ÷ price growth − 1 | Method consistency across 2021; TRI repair | sound | fine | One method throughout. Depends on the TRI repair (see D rows) |
| V5 | Tipsheet CAPE vs IIMA | Month-end comparison | — | sound | fine | Unadjusted series: correlation 0.995 over 212 months |
| V7 | (found in Phase 1, 2026-10-03) Inputs of C1, the lab’s yield-gap glide and dynamic SIP, BAF and the site dossiers | Read DERIVED/valuation_nifty_50/500.parquet | Which code writes the file | fix | wrong number (stale) | Nothing wrote these files any more (the valuation step writes valuation_by_index/). The copies had stopped at 2026-09-28 and would have gone staler every day, in production too. This is also why V3’s fix did not reach the lab on the first rerun |
| V6 | Real 10-year G-sec yield (CPI and survey) | Yield minus same-month CPI y/y | Lag | sound | weak | Same-month CPI is not yet published. Withheld anyway (bond yields not published) |
Evidence chapters (compute/evidence.py, site_calculations_spec.md)
| ID | What it claims | Method | Checks run | Verdict | Severity | Evidence |
|---|---|---|---|---|---|---|
| E1 | C1: CAPE slope against 5- and 10-year real returns “negative”; “nine in ten non-overlapping subsamples negative”; NW t = −5.4 (5y) | Overlapping monthly OLS, Newey-West with h lags; per-offset non-overlap slopes | Overlap inference; predictor persistence; signal basis | rethink | misleading | Newey-West: 60 or 120 lags on 196–256 observations is badly undersized (Ang and Bekaert 2007), and the 5-year t of −5.4 is not credible either. Stambaugh bias: a persistent predictor (log CAPE) biases the slope in small samples; not addressed. Offsets: adjacent offsets share 59/60 of their data, so “90% of offsets negative” is not 90% of independent evidence. R²: overlapping R² (0.69/0.76) is inflated. Signal basis: IIMA CAPE is unadjusted for 2021 (affects the last 5 start months); the earnings-yield signal inherits V3. Replaced by D1 (audit_deepening_spec.md) |
| E2 | C1 terciles: 12.8% / 7.9% / 2.3% median real 5-year returns | Full-sample tercile bounds | Look-ahead; uncertainty | fix | weak | Full-sample bounds are stated (“not a trading rule”). About 1.4 independent periods per tercile, with no interval. D2 adds intervals |
| E3 | C1 scatter | All overlapping monthly points | Presentation | fix | misleading | 256 points look like 256 observations; there are 5.3 independent periods. Mark the non-overlapping points in the bundle (nonoverlap flag) |
| E4 | C2 breadth thrusts: 12 signals; 6-month mean 15.5% vs 8.8%, p = 0.09; 12-month p = 0.02 (flagged exploratory) | Zweig rule; circular-shift null | Entry timing; overlap; multiplicity | fix | weak | Timing: the forward return starts at the signal day’s close, but breadth is known only after that close. Next-close entry: 1m mean 2.68 → 2.00, 3m 4.81 → 4.68, 6m 15.53 → 14.90, 12m 30.91 → 30.74. Null: the circular shift keeps the signal spacing, so overlap is handled. Multiplicity: 4 horizons; 12m already flagged exploratory |
| E5 | C3: returns predict flows; flows do not predict returns | Weekly Granger (BIC lag), F-test; cross-correlations | Variance; timing; sample | sound | weak | Variance: the homoskedastic F; the null is robust (p ≥ 0.29 in every sample). Timing: NSDL reporting dates trail trade dates by about a day. Sample: FPI covers only 5 years. All stated |
Trend and momentum models (compute/trend.py, momentum.py, aftertax.py, trend_momentum_spec.md)
| ID | What it claims | Method | Checks run | Verdict | Severity | Evidence |
|---|---|---|---|---|---|---|
| T1 | 35 trend results (rule × asset): CAGR, Sharpe, drawdown, DSR | backtest(): exposure decided at close t earns close t → t+1 | Signal-to-trade timing vs spec (“position changes at the close of t+1”) | fix | wrong number | Trades at the signal’s own close. With the spec’s timing (exposure.shift(1) more): Midcap 150 sma200d 18.00 → 16.48%, Smallcap 250 18.13 → 15.83%, Nifty 500 14.24 → 12.38%, Nifty 50 11.51 → 9.87%. Daily rules lose most; G-sec unaffected. Published: models/trend/*, the sample post trend-rules.mdx (“18.1% … 15.4%”, “18.3% against 13.8%”), and the spec’s results log |
| T2 | Current state of the monthly rules (sma10m, tsmom12, blend) | _month_ends = last date of each month present | Partial month | fix | wrong number | The latest close (e.g. 1 Oct, one day into the month) counts as a month-end. Nifty Next 50 sma10m reads “since 2026-10-01” although the September month-end was above its 10-month average (site Requests, 2026-10-02). /models/ carries a note about it |
| T3 | Every CAGR in trend, momentum (daily), the trend barometer and lab v1 | years = sessions ÷ 252; running cost = annual ÷ 252 per session | Sessions per year in the data | fix | wrong number (small) | The panel has 248.8 sessions a year (Nifty 50 6,780 sessions over 27.25 years). Years are understated 1.3%, so every CAGR is overstated by about 1.3% of itself (12.3% → about 12.45%) and running costs are under-charged by 1.3%. The lab v2 engine already uses calendar days (L1) |
| T4 | Deflated Sharpe ratios (trend, 35 trials) | Bailey and López de Prado 2014 | Formula; trial count | sound | fine | Formula checked term by term: per-period SR, raw kurtosis in (γ4 − 1)/4, SR0 from the variance of trial SRs with the Euler–Mascheroni weights. Trial count: 35, honest. Site-wide multiplicity in D3 |
| T5 | After-tax CAGRs of trend rules (“trailed buy-and-hold on every large-cap index by 1.9–6.1 points”) | aftertax.py: FIFO lots, own rate table, tax at each sale, no set-off | Rules against lab/tax.py (the lab’s audited engine) | fix | wrong number | No set-off: losses never offset gains. Whipsaw losses are exactly what trend rules realise, so this biases the comparison against the rules. No indexation on pre-2023 debt and gold long-term gains: overstates tax on the cash leg and gold. Gold section 50AA: treated as ending 2024-07-23 instead of 2025-04-01. Debt sold after 2024-07-23: taxed at 20% after 3 years instead of 12.5% after 24 months. Pre-2014 debt: long-term threshold is 12 months, not 36. No cess. Two tax engines on one site disagree; use the lab’s |
| T6 | Sector momentum mom12_1/mom6_1 vs equal weight; DSR 0.41/0.73 | Top 3 of NSE sectoral TRIs by 12-1 / 6-1 return | Back-fill; eligibility; costs; DSR | fix | misleading | Back-fill: 16 of 23 sector series start on a shared 2005-04-01 base date; most were launched years later, and which sectors exist is itself hindsight (stated generally in the spec, not quantified). Eligibility: peeks at whether next month’s return exists (minor). Costs: the equal-weight benchmark pays no rebalancing cost, though the spec says same costs. DSR: computed with trial-SR variance 0, so it is a probabilistic Sharpe ratio labelled as deflated |
| T7 | Sector momentum running cost 0.15% a year | Spec cost table | Realism | sound | weak | Sector ETFs charge more (often 0.2–0.5%); pre-registered, so noted |
Barometers (compute/barometers_*.py, three specs)
| ID | What it claims | Method | Checks run | Verdict | Severity | Evidence |
|---|---|---|---|---|---|---|
| BT1 | Trend barometer TR-H1 to H4 (ensemble Sharpe 0.56 vs 0.40, DSR 0.016, etc.) | 24-signal ensemble; _bt lags one extra session before calling backtest | Timing; interaction with T1 | sound | fine | Numbers correct today. Coupled to T1: when backtest is fixed, _bt must drop its extra lag, or every barometer result shifts a further day. Annualisation inherits T3 |
| BT2 | Risk tests (FG-H1–H3, TR-H4, BB-H3) | OLS with HAC (lags = forward window); halves; out of sample from 2016; non-overlapping sign | Look-ahead in OOS training; lag choice | sound | fine | Training rows whose forward window had not closed are excluded (BDay(lags) gap). Lags of 21/63 with n ≈ 4,500 are well-sized |
| BT3 | “What happened next” band tables (mood, bull-bear, trend) | Band medians, IQR, hit rate, greedy independent count | Uncertainty | fix | weak | Independent counts are shown, but there are no intervals on the medians. D2 |
| R1 | Nifty200 Momentum 30 replica “passes the 0.90 gate”, so the timing-luck study is attributed to the index | Correlation of monthly returns | What the gate can tell apart | rethink | misleading | Nifty 500 alone correlates 0.892 with the index. Wrong-schedule replicas score 0.889–0.906 against the replica’s 0.912. Excess returns over Nifty 200 correlate only 0.57. CAGR 13.7% (replica) vs 18.8% (index) over the same 257 months; tracking error 11.2%. The gate can’t tell the replica from the market, so the study must not be attributed to the index |
| R2 | Delisted stocks exit at their last close (momentum, IPO) | Adjusted price panel | ISIN continuity | fix (lane 1 / prices) | wrong number | 35 ISINs are split across two canonical symbols because renames are not chained (BURGERKING → RBA, DSML → DIL, DUDIGITAL → DUGLOBAL, AVANTIFEED, …). The IPO “only mainboard issue that stopped trading” (BURGERKING) kept trading for 1,145 sessions as RBA |
| R3 | Stock momentum robust to delisting | Stress test | Delisting returns of −30% | sound | fine | 44 of 71,654 stock-months stop trading. All at −30% moves long-short from 1.827% to 1.809% a month |
| R4 | MO-H1: top quintile 25.7% vs universe 13.7% a year | long_only | Month alignment | fix | wrong number (small) | Two partial months enter as 0% after costs for Q5 (260 vs 258 months). Matched months: 25.97% vs 13.73%. Both are price returns. Verdict unchanged (NW t 4.74, DSR 0.978) |
| R5 | Partial-month exclusion | last < last + BMonthEnd(0) | Holidays | fix | weak | Drops a complete month whose last weekday is an NSE holiday. Same rule in ipo.py:321 and barometers_indexmom.py:34 |
| R6 | MO-H3 Welch p = 0.035 | Normal approximation | t vs normal | sound | fine | A t-distribution gives about 0.04; it still fails the 0.0167 bar |
| R7 | Publication lags (AMFI +15 days, VIX +1 day, IPO heat by completed month, MTF same day) | Traced first appearance | Each lag | sound (MTF unverified) | weak | AMFI, VIX and IPO lags verified. MTF same-evening release not verified; if NSE publishes it next day, there is a one-session look-ahead on a 126-session change |
| R8 | Bull-and-bear “new issues” = 12-month IPO heat | — | Spec vs code | fix (log) | weak | Code uses ipo_heat (3-month window) (barometers_mood.py:126-129) |
| R9 | FG-H4a extreme fear → higher 63-session return (p = 0.0061, 14 episodes): the site’s main positive mood result | Circular block bootstrap, block 63 | Reproduction; block-length sensitivity | sound | fine | Reproduced at p = 0.0067. Block 126: p = 0.0073 (9 episodes). Block 252: p = 0.009 (6 episodes). Zone mean positive in all 11 years with zone sessions. Site-wide multiplicity in D3 |
| R10 | Stock universe liquidity floor in latest-CPI rupees | CPI of each bar’s month | Lag | sound | weak | CPI for the month is published about the 12th of the next month: a negligible look-ahead |
Portfolio lab v2 and SIP studies (lab/**)
| ID | What it claims | Method | Checks run | Verdict | Severity | Evidence |
|---|---|---|---|---|---|---|
| L1 | Headline CAGRs (gross, pre-tax, after tax), cost and tax drag | engine.simulate; years from calendar days | Independent recompute of 60/40, 2005-04 to 2026-09 | sound | fine | Gross 12.507% and after-cost pre-tax 11.912%; independent code matches to 3 decimals |
| L2 | Real after-tax CAGR | CPI ratio | Splice; end month | sound | fine | CPI ends a month before the window (negligible) |
| L3 | Volatility, Sharpe, Sortino, drawdown, Ulcer | Monthly excess over T-bill | Formulas | sound | weak | Inherits L4 for gold portfolios |
| L4 | Gold series in every gold portfolio; P2 “gold earned its place” | WGC INR price, forward-filled | Stale runs | fix (assets, lane 1 file) | misleading | Flat for 46 sessions from 2020-03-20 to 06-02, then +13.7% in one day. Also flat 22 sessions in 2016 and 20 in 2021. Gold looks smoothed through the 2020 crash; monthly rules and SIP buys used stale prices. Impact on results not yet measured |
| L5 | Synthetic 10-year G-sec | Par-bond reprice plus carry | Carry identity, duration, volatility, yield dates | sound | weak | Duration 7.1; correlation 0.84 with the 5-year index. The geometric fill between month-ends leaks about 1/21 of next month’s yield move into the first session (universe.py:131) |
| L6 | Yield-gap glide and yield-gap dynamic SIP | Expanding percentile of EY − 10y, one-month lag | P/E basis | fix | wrong number | Inherits V3. After April 2021 the glide holds 21 pp more equity on average than on the adjusted basis (up to 27 pp). The SIP multiplier differs in 70% of post-2021 months |
| L7 | CAPE glide and CAPE dynamic SIP | IIMA Sensex CAPE10, one-month lag | Vintage, lag | sound | weak | One 2026-08-31 vintage (revised data). IIMA’s real publication lag not verified |
| L8 | P1 “holds”; nine rules “significantly worse” than 60/40 | After-tax point gap, pre-tax bootstrap interval | Pre-tax vs after-tax labelling | fix | misleading | “Significant” is a pre-tax statement next to after-tax gaps. Dual momentum is +0.39 pp pre-tax and −0.37 after tax |
| L9 | P2 gold raised Sharpe “in every one of 15 pairs” | Paired stationary bootstrap | Independence of pairs | fix | misleading | The 15 pairs share one gold path, so this is one result. One lower bound is 0.00. Inherits L4 |
| L10 | P3 | Point comparison | — | sound | weak | No interval |
| L11 | P4: factor indices worse after launch | Annual excess over parent, back-test vs live | Sample | sound | weak | Live samples of 15–27 months for the six launched after 2020; no interval |
| L12 | P5: PBO 0.65 | CSCV, 16 blocks | Implementation | sound | fine | — |
| L13 | Deflated Sharpe, “none survives 136 trials” | Bailey and López de Prado | Formula; N | sound | fine | Conservative: N counts identical clones (60/40, three-fund, BL) |
| L14 | Bootstrap CAGR and “vs 60/40” intervals | Stationary bootstrap, block 12 | Labels | fix | misleading | Pre-tax intervals sit beside after-tax CAGRs without a pre-tax label (publish/lab_v2.py:35-37) |
| L15 | P6: tax drag | Pre-tax minus after-tax | — | sound | fine | Yield-gap row inherits L6 |
| L16 | “Annual rebalancing beat monthly, quarterly, bands” | One January phase | Timing luck | rethink | weak | Fixed phase (engine.py:233). Pre-2013 STT is mixed into the explanation |
| L17 | Walk-forward 8.2% vs 10.9% | Trailing 60 months | Look-ahead | sound | fine | — |
| L18 | Rolling 5-year windows, share_5y_beating_60_40 | Overlapping monthly windows | Disclosure | fix | misleading | The spec promised the independent-period count beside the share; not published (run.py:185) |
| L19 | SIP outcome distributions | Engine checkpoints, after-tax XIRR | Recompute | sound | fine | Nifty 500, 10-year from Jan 2010: XIRR 11.0254%, matches |
| L20 | SIP replication table | Nifty 50 TRI, no costs | Recompute | sound | fine | Matches the log (207 windows, minimum 3.85%, median 13.67%) |
| L21 | SIP paired tests (date of month, dynamic, pause, exemption): published Newey-West t | paired(), NW lags = 12h | Lags vs n; simulation | fix | wrong number | Lags ≥ number of windows gives t = 38.3 (15y, 78 windows) and 73.1 (20y, 18 windows). An AR(1) null simulation (ρ 0.97, n 138, lags 120) rejects 53% of the time at a nominal 5% |
| L22 | S1 fails (date of month matters) | Paired medians | Turn-of-month effect | rethink | weak | The effect is real on non-overlapping months (25th→1st return 0.85%, t 5.2). Per-horizon significance leans on L21. 6 dates × 12 portfolios × 6 horizons |
| L23 | S2 holds | Share of windows below the FD proxy | FD construction | fix | weak | The proxy accrues a floating T-bill yield, not the spec’s one-year locked rate. The 15-year result is one independent period |
| L24 | S3 fails (dynamic SIP lost: −0.26 CAPE, −1.3 yield gap) | Dynamic multiplier | Signal | fix | wrong number | The yield-gap leg inherits L6; t-values inherit L21 |
| L25 | S4, S5 | Shares over overlapping start months | — | sound | weak | Describe history, not probabilities |
IPO, BAF, MF stress
| ID | What it claims | Method | Checks run | Verdict | Severity | Evidence |
|---|---|---|---|---|---|---|
| I1 | IPO sample (mainboard) | Exclusion file: no issue price, excluded | Selection | fix | misleading | About 9 of the 17 mainboard issues dropped for “no issue price” are real IPOs with a known price band (NSDL, Sterling & Wilson Solar, DCX, …). The final price is almost always the top of the band |
| I2 | Mean buy-and-hold abnormal returns (1, 3, 5 years) | Means | Clustered bootstrap; trimming | sound | weak | 95% CIs: 1y −0.8% to +14.6%, 3y −2.4% to +42.5%, 5y −0.5% to +124%. Without the top 5 issues: 3.4%, 4.5%, 7.7%. The 5-year sample is the 2016–21 cohorts (2021 is 36 of 143). Intervals should be published |
| I3 | H3: hot markets (null) | Welch | Month-level NW; runs | sound | weak | NW t −2.23 vs Welch −1.91, but “hot” months come in only 6 runs |
| I4 | H1 calendar-time portfolio | Monthly portfolio, NW 3 lags | Gaps | sound | fine | 124 months, complete |
| I5 | “Only one mainboard issue stopped trading” | Price panel | ISIN chaining | fix | wrong number | Inherits R2: really zero |
| B1 | BAF net-equity headline (AUM-weighted) | Validation on screened schemes | Coverage | fix | misleading | Validation covers 56% of AUM; HDFC BAF (32% of AUM) is screened out of validation but carries the largest weight in the headline |
| B2 | Per-scheme valuation link | Non-overlapping 3-month changes | Independence | sound | weak | Median r −0.19 (25 of 32 negative); schemes correlate 0.35 with each other, so these are not 32 independent votes |
| B3 | Arbitrage control near zero | Quarterly medians | Drift | sound | weak | Latest 0.073 (highest); watch |
| S1 | Small-cap stress test: median days to liquidate fell from 8 to 5 (Dec 2024 → Aug 2026) | Category median | Composition | fix | misleading | Same 27 schemes: 8 → 8. The fall comes entirely from 9 new funds (median 1 day) |
| S2 | Stress data integrity | — | Duplicates, ordering | sound | fine | — |
Data layer (prices, breadth, TRI, assets, density, dossiers)
| ID | What it claims | Method | Checks run | Verdict | Severity | Evidence |
|---|---|---|---|---|---|---|
| D1 | Index TRIs are repaired against the price index | repair_tri replaces the TRI’s return with the price return on bad days | Re-ran the 11 repairs; synthetic bad-price test | fix | weak (latent) | All 11 repairs are genuine TRI errors (Smallcap 100, 2005–09). But a bad price print gets copied into a clean TRI: a reverting −20% price print puts a −19.9% day into the TRI (tri_check.py:50,65) |
| D2 | NSE TRI and price histories are clean | Used as published | Index days above 8% while Nifty 500 moved less than 2% (55 hits) | fix | wrong number | Nifty Metal on 2006-11-21: +31.6% in both price and TRI, while its constituents rose 2.4–2.9%. Published 2006 return 98.1%; should be about 50.5%. The other 54 hits match real events. Nifty Energy and PSE on 2002-04-01 are unexplained |
| D3 | Gold in rupees including import duty | WGC INR from 2005; USD × USD/INR before that | Duty premium by year; stale runs | fix | wrong number (volatility, drawdown, signals) | Same root cause as L4. INR price frozen 2020-03-23 → 06-02 (46 sessions) while USD gold rose 1605 → 1730; also 22 sessions in 2016 and 20 in 2021. Duty premium itself checks out |
| D4 | Cash = 91-day T-bill accrual, no look-ahead | Yield known the day before; y·dt/365 compounded daily | Annual return vs rollover | fix | weak | Lag correct. Daily compounding of a 91-day simple yield overstates by about 6 bp a year at 7%: 2024 cash 7.04% vs 6.92% from quarterly rollover (assets.py:135) |
| D5 | Asset panel on NSE sessions | Unlimited forward fill | Staleness detection | fix | weak | assets.py:150 has no fill limit, so the density snapshot’s stale-asset rule can never fire |
| D6 | Equity TRIs | NSE TRI | Days with |r| above 15% | sound | fine | Only 2009-05-18 (election day, genuine) |
| D7 | Survivorship-free adjusted prices | Bhavcopy, rename chains, corporate actions | Code; QA lists | sound, except R2 | fine | Corporate-action logic careful. Rename chaining gap: R2 |
| D8 | Breadth liquidity floor in constant rupees | CPI of the bar’s own month | Lag | sound | weak | One month of CPI look-ahead (about 0.5% on the floor) (breadth.py:63) |
| D9 | Index breadth, calendar-year tables, trend breadth “across NSE indices”; Nifty 500 from 1995 | Today’s catalogue with back-filled histories | Launch dates; disclosure | rethink | misleading | Indices count before their launch. Hindsight-built strategy indices vote in site_dossiers.trend_breadth (unlike highs_lows). Back-fill is not disclosed in indices.md, breadth.md or dossiers.md |
| D10 | Market value / GDP | Month-end market value ÷ latest complete FY nominal GDP | April steps; vintages | fix | misleading | Steps each April on the denominator alone: Mar-25 141.3% → Apr-25 132.5% while market value rose 2.6%. April–May use FY GDP published only about 31 May; revised vintages throughout; the 2011-12 and 2022-23 GDP bases are joined (density_mcap.py:117) |
| D11 | Flows since Dec 2023 | Validated sessions; dropped sessions count as zero | Dropped-session count | sound (disclosed) | weak | 34 sessions dropped; FY26 has 229 sessions; counts published |
| D12 | Index dashboard returns, trend, momentum | Completed months; stale for more than 10 days dropped | Code | sound | fine | Annualises only 3 years and longer |
| D13 | Risk dossier “what followed each volatility quintile” | Expanding quintile edges | Independence count | fix | weak | “Independent” = sessions ÷ 63 (site_dossiers.py:89), though quintile days cluster into episodes (Q5: 977 sessions labelled 15.5) |
| D14 | Dossier trend breadth | Share of indices above their 10-month average | Partial month | fix | weak | The last point is month-to-date (site_dossiers.py:124-127) |
| D15 | Fund-house league | Drops periods below 50% of the previous total | Completeness | suspicion | weak | Periods 50–99% complete pass (density_funds.py:145) |
| D16 | Macro, world, commodities, ETF density figures | Latest vintages | Gaps | sound | fine | — |
| D17 | Price QA lists | Review only | — | sound | fine | — |
Site-wide multiple testing (D3, audit_deepening_spec.md)
Every confirmatory test on the site with a p-value:
- p-values are as registered (one- or two-sided).
- Results graded by a deflated Sharpe ratio enter as 1 − DSR; the DSR already corrects within its own family.
- Not corrected, because they were graded by a descriptive bar: lab P1–P6, SIP S1–S5, TR-H1, MO-H2, C1’s sign prediction, BAF’s validations.
- Excluded as untestable: BB-H1 and BB-H2 (fewer than 8 episodes).
Method: Holm–Bonferroni at familywise 0.05 across 35 tests, and Benjamini–Hochberg at a 10% false-discovery rate.
| Rank | Spec | Test | p | Holm bar | Holm | BH (10%) |
|---|---|---|---|---|---|---|
| 1 | site_calculations | C3 returns predict DII flows | <1e-6 (p (<1e-6)) | 0.0014 | pass | pass |
| 2 | site_calculations | C3 returns predict FPI flows | <1e-6 (p (<1e-6)) | 0.0015 | pass | pass |
| 3 | momentum_barometer | MO-H1 blended stock momentum premium | 1.1e-06 (p) | 0.0015 | pass | pass |
| 4 | market_mood | FG-H4a extreme fear then higher 63-session return | 0.0061 (p) | 0.0016 | — | pass |
| 5 | trend_barometer | TR-H4 score adds to B0 for forward volatility | 0.024 (p) | 0.0016 | — | — |
| 6 | audit_deepening | D1 log_pb 5y (Nifty 50 real) | 0.025 (bootstrap p) | 0.0017 | — | — |
| 7 | market_mood | FG-H2 score adds to B0 for a 5% fall | 0.031 (p) | 0.0017 | — | — |
| 8 | momentum_barometer | MO-H3 crash flag (max of Welch and bootstrap p) | 0.035 (p) | 0.0018 | — | — |
| 9 | audit_deepening | D1 log_cape 5y (Nifty 50 real) | 0.055 (bootstrap p) | 0.0019 | — | — |
| 10 | audit_deepening | D1 dy 1y (Nifty 50 real) | 0.061 (bootstrap p) | 0.0019 | — | — |
| 11 | audit_deepening | D1 log_pb 3y (Nifty 50 real) | 0.064 (bootstrap p) | 0.0020 | — | — |
| 12 | audit_deepening | D1 dy 3y (Nifty 50 real) | 0.065 (bootstrap p) | 0.0021 | — | — |
| 13 | audit_deepening | D1 log_pb 1y (Nifty 50 real) | 0.093 (bootstrap p) | 0.0022 | — | — |
| 14 | site_calculations | C2 breadth thrust, 6-month forward return | 0.107 (p) | 0.0023 | — | — |
| 15 | market_mood | BB-H3 bull-bear adds to B0+FG for a 10% fall | 0.178 (p) | 0.0024 | — | — |
| 16 | audit_deepening | D1 log_cape 3y (Nifty 50 real) | 0.186 (bootstrap p) | 0.0025 | — | — |
| 17 | audit_deepening | D1 log_pe 1y (Nifty 50 real) | 0.215 (bootstrap p) | 0.0026 | — | — |
| 18 | audit_deepening | D1 log_cape 1y (Nifty 50 real) | 0.249 (bootstrap p) | 0.0028 | — | — |
| 19 | audit_deepening | D1 dy 5y (Nifty 50 real) | 0.267 (bootstrap p) | 0.0029 | — | — |
| 20 | ipo | H3 hot-month issues trail | 0.270 (p) | 0.0031 | — | — |
| 21 | audit_deepening | D1 log_pe 3y (Nifty 50 real) | 0.285 (bootstrap p) | 0.0033 | — | — |
| 22 | market_mood | FG-H3 score adds to B0 for a 10% fall | 0.287 (p) | 0.0036 | — | — |
| 23 | audit_deepening | D1 log_pe 5y (Nifty 50 real) | 0.318 (bootstrap p) | 0.0038 | — | — |
| 24 | site_calculations | C3 DII flows predict returns | 0.332 (p) | 0.0042 | — | — |
| 25 | trend_momentum | Sector momentum mom6_1 vs equal weight (deflated, 2 trials) | 0.345 (1 - DSR) | 0.0045 | — | — |
| 26 | site_calculations | C3 FPI flows predict returns | 0.354 (p) | 0.0050 | — | — |
| 27 | ipo | H1 IPO calendar-time alpha vs Nifty 500 | 0.370 (p (NW t 0.89)) | 0.0056 | — | — |
| 28 | momentum_barometer | MO-H5 de-duplicated sector momentum (deflated, 3 trials) | 0.433 (1 - DSR) | 0.0063 | — | — |
| 29 | trend_momentum | Sector momentum mom12_1 vs equal weight (deflated, 2 trials) | 0.674 (1 - DSR) | 0.0071 | — | — |
| 30 | market_mood | FG-H1 score adds to B0 for forward volatility | 0.744 (p) | 0.0083 | — | — |
| 31 | trend_barometer | TR-H4 score adds to B0 for a 10% fall | 0.767 (p) | 0.0100 | — | — |
| 32 | trend_momentum | Best trend rule vs buy-and-hold (35 trials, deflated) | 0.860 (1 - DSR) | 0.0125 | — | — |
| 33 | ipo | H2 listing gain vs 1-year BHAR | 0.950 (p) | 0.0167 | — | — |
| 34 | trend_barometer | TR-H3 ensemble beats buy-and-hold (deflated, 84 trials) | 0.984 (1 - DSR) | 0.0250 | — | — |
| 35 | market_mood | FG-H4b extreme greed then lower 63-session return | 0.993 (p) | 0.0500 | — | — |
Reading:
- Survive Holm: the stock-momentum premium (MO-H1), and returns predicting the next week’s DII and FPI flows.
- Survives only Benjamini–Hochberg: extreme fear followed by better 3-month returns (FG-H4a).
- Nothing else survives that forecasts market returns.
Summary of Phase 0
Wrong numbers to fix first:
- T1 (trend timing).
- V3 → L6, L24, E1 (the P/E basis break in earnings yield and yield gap).
- T5 (after-tax engine).
- L21 (SIP Newey-West t-values).
- T2 (partial month).
- T3 (annualisation).
- R2/I5 (rename chaining; lane 1 request).
- R4.
- The data-layer rows D2, D3 and D10 (other lanes’ modules: requests filed in
docs/status/quant.md).
Misleading:
- R1 (index attribution).
- E1/E3 (C1 inference and scatter).
- L8, L14 (pre-tax intervals beside after-tax numbers).
- L4 (gold stale prices).
- L9, L18, I1, B1, S1, T6.
Outcomes (Phase 1, 2026-10-03)
Numbers are from this worktree after the fixes (data to 1 Oct 2026). Old and new figures for each pre-registered result are in that spec’s results log.
| Rows | Outcome |
|---|---|
| T1, T2, T3, BT1 | Fixed. Trades at the close of t+1, completed month-ends only, calendar-time CAGR and costs; the barometer’s compensating lag removed. Midcap 150 sma200d 18.0 → 16.2% (buy-and-hold 15.0%); Smallcap 250 18.1 → 15.5% (13.4%). Verdicts unchanged. Logged in trend_momentum_spec.md and trend_barometer_spec.md |
| T5 | Fixed. The trend after-tax engine is the lab’s tax engine. Large caps after tax: behind buy-and-hold by 1.1–5.2 points (was 1.9–6.1) |
| T6 | Fixed: eligibility, equal-weight costs, drifted-weight turnover, DSR variance. DSR 0.41/0.73 → 0.33/0.66. Back-fill still open (needs launch dates; request to lane 1) |
| V3, V7 | Fixed. Earnings yield and yield gaps on the consolidated basis; the root-level valuation copies are maintained again. Lab yield-gap glide 10.74 → 10.32% after tax; S3’s yield-gap leg −1.30 → −1.57 points |
| E1 | Superseded by D1: no valuation predictor passes. The C1 “nine in ten subsamples” sentence is withdrawn |
| E2, E3, BT3 | Fixed (D2). 90% intervals and noise and “too few” flags on every what-next table; non-overlap flag on the C1 scatter |
| E4 | Fixed. C2 enters at the next close; 6-month p 0.09 → 0.11 |
| L8, L14, L18 | Fixed (labels and disclosure). interval_basis meta; independent_5y_periods |
| L21, L22, L24 | Fixed. SIP Newey-West t only where windows ≥ 3 × lag, with fixed-b critical values. S1 and S3 verdicts stand |
| R1 | Fixed. Attribution to Nifty200 Momentum 30 withdrawn (excess-return correlation 0.57); amendment logged |
| R4, R5 | Fixed. MO-H1 long-only 26.0% vs 13.7% on matched months; holiday-safe month completion |
| R8, R7, R9 | Logged in market_mood_spec.md |
| D4 | Fixed (lane 1 file, one-line clear bug, flagged): cash compounds per 91-day bill |
| D1–D3, D5, D8–D15, R2, I1, I5, B1, S1, L4 | Requested from the owning lanes, with failing tests (docs/status/quant.md, Requests) |
| L5, L7, L11, L16, L23 | Open, weak: noted in the specs; next revision |
Independent re-implementations: pipeline/tests/audit/ (13 tests, stdlib, numpy and DuckDB only). They cover:
- all 35 trend runs, and two deflated Sharpe ratios;
- Nifty 50 CAPE (in SQL) and the consolidated earnings yield;
- the C1 slope;
- sector momentum
mom12_1; - the cash index;
- the lab’s 60/40 pre-tax CAGR;
- one SIP XIRR (to the paisa).
All reconcile to their tolerances.