Data to 5 October 2026

Pre-registered research

Audit register: every calculation that produces a claim

Opened: 2026-10-03 by lane 2, for the model audit (docs/briefs/model_audit.md).

Phase 0 (this file, committed before any fix): one row per claim-producing calculation, with the method, the checks run, a verdict and a severity.

After Phase 0: fixes are logged in each spec’s results log (date, reason, old and new numbers) and in docs/status/quant.md. This register gains an “Outcome” column, and its rows are never deleted.

Key

Verdicts:

Severities:

Checklist used for every row:

Independent re-implementations of headline numbers live in pipeline/tests/audit/, and the rows say which ones are covered.

Who checked what:

Valuation (compute/valuation.py, publish/valuation.py, docs/methods/valuation.md)

IDWhat it claimsMethodChecks runVerdictSeverityEvidence
V1CAPE on a consistent consolidated basis (22 indices)Implied EPS = close ÷ P/E. Pre-2021-03-31 EPS scaled by the switch-day ratio. CPI-deflated 120-month mean, from the month beforeSplice robustness; CPI splice; deflator cancels; CPI lagsoundweakSplice: EPS moved more than 1% on a single day between 15 Mar and 15 Apr 2021, on 03-31 only, for Nifty 50 (×1.205), 500 (×1.200), Bank, IT, FMCG and Next 50. So the switch is a clean single break. Deflator: the latest-CPI level cancels in the ratio. CPI lag: the historical series deflates price by the current month’s CPI, published about the 12th of the next month. The look-ahead is at most one month’s inflation (about 0.5%) on the level; the latest point uses the last published CPI. Basis: the “constant consolidation gap” assumption is untestable and is stated
V2P/E and P/B percentile of today within own historyShare of history ≤ today, on the consolidated basisFull-sample vs expandingsoundfineToday’s position in history is legitimately full-sample. History starts 1999 (one regime), and the start date is shown
V3Earnings yield, nominal yield gap, CAPE yield minus real G-sec100 ÷ P/E; minus the 10-year yieldP/E basisfixwrong numbervaluation.py:174 uses raw pe_ratio, so the earnings yield steps up 17–20% on 2021-03-31 (Nifty 500 EY 2.31 → 2.79). The published earnings_yield_pct series and every consumer inherit the break: C1 (E1), the lab’s yield-gap glide and the dynamic SIP (L6)
V4Dividend yield from total-return over price-return250-session TRI growth ÷ price growth − 1Method consistency across 2021; TRI repairsoundfineOne method throughout. Depends on the TRI repair (see D rows)
V5Tipsheet CAPE vs IIMAMonth-end comparison—soundfineUnadjusted series: correlation 0.995 over 212 months
V7(found in Phase 1, 2026-10-03) Inputs of C1, the lab’s yield-gap glide and dynamic SIP, BAF and the site dossiersRead DERIVED/valuation_nifty_50/500.parquetWhich code writes the filefixwrong number (stale)Nothing wrote these files any more (the valuation step writes valuation_by_index/). The copies had stopped at 2026-09-28 and would have gone staler every day, in production too. This is also why V3’s fix did not reach the lab on the first rerun
V6Real 10-year G-sec yield (CPI and survey)Yield minus same-month CPI y/yLagsoundweakSame-month CPI is not yet published. Withheld anyway (bond yields not published)

Evidence chapters (compute/evidence.py, site_calculations_spec.md)

IDWhat it claimsMethodChecks runVerdictSeverityEvidence
E1C1: CAPE slope against 5- and 10-year real returns “negative”; “nine in ten non-overlapping subsamples negative”; NW t = −5.4 (5y)Overlapping monthly OLS, Newey-West with h lags; per-offset non-overlap slopesOverlap inference; predictor persistence; signal basisrethinkmisleadingNewey-West: 60 or 120 lags on 196–256 observations is badly undersized (Ang and Bekaert 2007), and the 5-year t of −5.4 is not credible either. Stambaugh bias: a persistent predictor (log CAPE) biases the slope in small samples; not addressed. Offsets: adjacent offsets share 59/60 of their data, so “90% of offsets negative” is not 90% of independent evidence. R²: overlapping R² (0.69/0.76) is inflated. Signal basis: IIMA CAPE is unadjusted for 2021 (affects the last 5 start months); the earnings-yield signal inherits V3. Replaced by D1 (audit_deepening_spec.md)
E2C1 terciles: 12.8% / 7.9% / 2.3% median real 5-year returnsFull-sample tercile boundsLook-ahead; uncertaintyfixweakFull-sample bounds are stated (“not a trading rule”). About 1.4 independent periods per tercile, with no interval. D2 adds intervals
E3C1 scatterAll overlapping monthly pointsPresentationfixmisleading256 points look like 256 observations; there are 5.3 independent periods. Mark the non-overlapping points in the bundle (nonoverlap flag)
E4C2 breadth thrusts: 12 signals; 6-month mean 15.5% vs 8.8%, p = 0.09; 12-month p = 0.02 (flagged exploratory)Zweig rule; circular-shift nullEntry timing; overlap; multiplicityfixweakTiming: the forward return starts at the signal day’s close, but breadth is known only after that close. Next-close entry: 1m mean 2.68 → 2.00, 3m 4.81 → 4.68, 6m 15.53 → 14.90, 12m 30.91 → 30.74. Null: the circular shift keeps the signal spacing, so overlap is handled. Multiplicity: 4 horizons; 12m already flagged exploratory
E5C3: returns predict flows; flows do not predict returnsWeekly Granger (BIC lag), F-test; cross-correlationsVariance; timing; samplesoundweakVariance: the homoskedastic F; the null is robust (p ≥ 0.29 in every sample). Timing: NSDL reporting dates trail trade dates by about a day. Sample: FPI covers only 5 years. All stated

Trend and momentum models (compute/trend.py, momentum.py, aftertax.py, trend_momentum_spec.md)

IDWhat it claimsMethodChecks runVerdictSeverityEvidence
T135 trend results (rule × asset): CAGR, Sharpe, drawdown, DSRbacktest(): exposure decided at close t earns close t → t+1Signal-to-trade timing vs spec (“position changes at the close of t+1”)fixwrong numberTrades at the signal’s own close. With the spec’s timing (exposure.shift(1) more): Midcap 150 sma200d 18.00 → 16.48%, Smallcap 250 18.13 → 15.83%, Nifty 500 14.24 → 12.38%, Nifty 50 11.51 → 9.87%. Daily rules lose most; G-sec unaffected. Published: models/trend/*, the sample post trend-rules.mdx (“18.1% … 15.4%”, “18.3% against 13.8%”), and the spec’s results log
T2Current state of the monthly rules (sma10m, tsmom12, blend)_month_ends = last date of each month presentPartial monthfixwrong numberThe latest close (e.g. 1 Oct, one day into the month) counts as a month-end. Nifty Next 50 sma10m reads “since 2026-10-01” although the September month-end was above its 10-month average (site Requests, 2026-10-02). /models/ carries a note about it
T3Every CAGR in trend, momentum (daily), the trend barometer and lab v1years = sessions ÷ 252; running cost = annual ÷ 252 per sessionSessions per year in the datafixwrong number (small)The panel has 248.8 sessions a year (Nifty 50 6,780 sessions over 27.25 years). Years are understated 1.3%, so every CAGR is overstated by about 1.3% of itself (12.3% → about 12.45%) and running costs are under-charged by 1.3%. The lab v2 engine already uses calendar days (L1)
T4Deflated Sharpe ratios (trend, 35 trials)Bailey and López de Prado 2014Formula; trial countsoundfineFormula checked term by term: per-period SR, raw kurtosis in (γ4 − 1)/4, SR0 from the variance of trial SRs with the Euler–Mascheroni weights. Trial count: 35, honest. Site-wide multiplicity in D3
T5After-tax CAGRs of trend rules (“trailed buy-and-hold on every large-cap index by 1.9–6.1 points”)aftertax.py: FIFO lots, own rate table, tax at each sale, no set-offRules against lab/tax.py (the lab’s audited engine)fixwrong numberNo set-off: losses never offset gains. Whipsaw losses are exactly what trend rules realise, so this biases the comparison against the rules. No indexation on pre-2023 debt and gold long-term gains: overstates tax on the cash leg and gold. Gold section 50AA: treated as ending 2024-07-23 instead of 2025-04-01. Debt sold after 2024-07-23: taxed at 20% after 3 years instead of 12.5% after 24 months. Pre-2014 debt: long-term threshold is 12 months, not 36. No cess. Two tax engines on one site disagree; use the lab’s
T6Sector momentum mom12_1/mom6_1 vs equal weight; DSR 0.41/0.73Top 3 of NSE sectoral TRIs by 12-1 / 6-1 returnBack-fill; eligibility; costs; DSRfixmisleadingBack-fill: 16 of 23 sector series start on a shared 2005-04-01 base date; most were launched years later, and which sectors exist is itself hindsight (stated generally in the spec, not quantified). Eligibility: peeks at whether next month’s return exists (minor). Costs: the equal-weight benchmark pays no rebalancing cost, though the spec says same costs. DSR: computed with trial-SR variance 0, so it is a probabilistic Sharpe ratio labelled as deflated
T7Sector momentum running cost 0.15% a yearSpec cost tableRealismsoundweakSector ETFs charge more (often 0.2–0.5%); pre-registered, so noted

Barometers (compute/barometers_*.py, three specs)

IDWhat it claimsMethodChecks runVerdictSeverityEvidence
BT1Trend barometer TR-H1 to H4 (ensemble Sharpe 0.56 vs 0.40, DSR 0.016, etc.)24-signal ensemble; _bt lags one extra session before calling backtestTiming; interaction with T1soundfineNumbers correct today. Coupled to T1: when backtest is fixed, _bt must drop its extra lag, or every barometer result shifts a further day. Annualisation inherits T3
BT2Risk tests (FG-H1–H3, TR-H4, BB-H3)OLS with HAC (lags = forward window); halves; out of sample from 2016; non-overlapping signLook-ahead in OOS training; lag choicesoundfineTraining rows whose forward window had not closed are excluded (BDay(lags) gap). Lags of 21/63 with n ≈ 4,500 are well-sized
BT3“What happened next” band tables (mood, bull-bear, trend)Band medians, IQR, hit rate, greedy independent countUncertaintyfixweakIndependent counts are shown, but there are no intervals on the medians. D2
R1Nifty200 Momentum 30 replica “passes the 0.90 gate”, so the timing-luck study is attributed to the indexCorrelation of monthly returnsWhat the gate can tell apartrethinkmisleadingNifty 500 alone correlates 0.892 with the index. Wrong-schedule replicas score 0.889–0.906 against the replica’s 0.912. Excess returns over Nifty 200 correlate only 0.57. CAGR 13.7% (replica) vs 18.8% (index) over the same 257 months; tracking error 11.2%. The gate can’t tell the replica from the market, so the study must not be attributed to the index
R2Delisted stocks exit at their last close (momentum, IPO)Adjusted price panelISIN continuityfix (lane 1 / prices)wrong number35 ISINs are split across two canonical symbols because renames are not chained (BURGERKING → RBA, DSML → DIL, DUDIGITAL → DUGLOBAL, AVANTIFEED, …). The IPO “only mainboard issue that stopped trading” (BURGERKING) kept trading for 1,145 sessions as RBA
R3Stock momentum robust to delistingStress testDelisting returns of −30%soundfine44 of 71,654 stock-months stop trading. All at −30% moves long-short from 1.827% to 1.809% a month
R4MO-H1: top quintile 25.7% vs universe 13.7% a yearlong_onlyMonth alignmentfixwrong number (small)Two partial months enter as 0% after costs for Q5 (260 vs 258 months). Matched months: 25.97% vs 13.73%. Both are price returns. Verdict unchanged (NW t 4.74, DSR 0.978)
R5Partial-month exclusionlast < last + BMonthEnd(0)HolidaysfixweakDrops a complete month whose last weekday is an NSE holiday. Same rule in ipo.py:321 and barometers_indexmom.py:34
R6MO-H3 Welch p = 0.035Normal approximationt vs normalsoundfineA t-distribution gives about 0.04; it still fails the 0.0167 bar
R7Publication lags (AMFI +15 days, VIX +1 day, IPO heat by completed month, MTF same day)Traced first appearanceEach lagsound (MTF unverified)weakAMFI, VIX and IPO lags verified. MTF same-evening release not verified; if NSE publishes it next day, there is a one-session look-ahead on a 126-session change
R8Bull-and-bear “new issues” = 12-month IPO heat—Spec vs codefix (log)weakCode uses ipo_heat (3-month window) (barometers_mood.py:126-129)
R9FG-H4a extreme fear → higher 63-session return (p = 0.0061, 14 episodes): the site’s main positive mood resultCircular block bootstrap, block 63Reproduction; block-length sensitivitysoundfineReproduced at p = 0.0067. Block 126: p = 0.0073 (9 episodes). Block 252: p = 0.009 (6 episodes). Zone mean positive in all 11 years with zone sessions. Site-wide multiplicity in D3
R10Stock universe liquidity floor in latest-CPI rupeesCPI of each bar’s monthLagsoundweakCPI for the month is published about the 12th of the next month: a negligible look-ahead

Portfolio lab v2 and SIP studies (lab/**)

IDWhat it claimsMethodChecks runVerdictSeverityEvidence
L1Headline CAGRs (gross, pre-tax, after tax), cost and tax dragengine.simulate; years from calendar daysIndependent recompute of 60/40, 2005-04 to 2026-09soundfineGross 12.507% and after-cost pre-tax 11.912%; independent code matches to 3 decimals
L2Real after-tax CAGRCPI ratioSplice; end monthsoundfineCPI ends a month before the window (negligible)
L3Volatility, Sharpe, Sortino, drawdown, UlcerMonthly excess over T-billFormulassoundweakInherits L4 for gold portfolios
L4Gold series in every gold portfolio; P2 “gold earned its place”WGC INR price, forward-filledStale runsfix (assets, lane 1 file)misleadingFlat for 46 sessions from 2020-03-20 to 06-02, then +13.7% in one day. Also flat 22 sessions in 2016 and 20 in 2021. Gold looks smoothed through the 2020 crash; monthly rules and SIP buys used stale prices. Impact on results not yet measured
L5Synthetic 10-year G-secPar-bond reprice plus carryCarry identity, duration, volatility, yield datessoundweakDuration 7.1; correlation 0.84 with the 5-year index. The geometric fill between month-ends leaks about 1/21 of next month’s yield move into the first session (universe.py:131)
L6Yield-gap glide and yield-gap dynamic SIPExpanding percentile of EY − 10y, one-month lagP/E basisfixwrong numberInherits V3. After April 2021 the glide holds 21 pp more equity on average than on the adjusted basis (up to 27 pp). The SIP multiplier differs in 70% of post-2021 months
L7CAPE glide and CAPE dynamic SIPIIMA Sensex CAPE10, one-month lagVintage, lagsoundweakOne 2026-08-31 vintage (revised data). IIMA’s real publication lag not verified
L8P1 “holds”; nine rules “significantly worse” than 60/40After-tax point gap, pre-tax bootstrap intervalPre-tax vs after-tax labellingfixmisleading“Significant” is a pre-tax statement next to after-tax gaps. Dual momentum is +0.39 pp pre-tax and −0.37 after tax
L9P2 gold raised Sharpe “in every one of 15 pairs”Paired stationary bootstrapIndependence of pairsfixmisleadingThe 15 pairs share one gold path, so this is one result. One lower bound is 0.00. Inherits L4
L10P3Point comparison—soundweakNo interval
L11P4: factor indices worse after launchAnnual excess over parent, back-test vs liveSamplesoundweakLive samples of 15–27 months for the six launched after 2020; no interval
L12P5: PBO 0.65CSCV, 16 blocksImplementationsoundfine—
L13Deflated Sharpe, “none survives 136 trials”Bailey and López de PradoFormula; NsoundfineConservative: N counts identical clones (60/40, three-fund, BL)
L14Bootstrap CAGR and “vs 60/40” intervalsStationary bootstrap, block 12LabelsfixmisleadingPre-tax intervals sit beside after-tax CAGRs without a pre-tax label (publish/lab_v2.py:35-37)
L15P6: tax dragPre-tax minus after-tax—soundfineYield-gap row inherits L6
L16“Annual rebalancing beat monthly, quarterly, bands”One January phaseTiming luckrethinkweakFixed phase (engine.py:233). Pre-2013 STT is mixed into the explanation
L17Walk-forward 8.2% vs 10.9%Trailing 60 monthsLook-aheadsoundfine—
L18Rolling 5-year windows, share_5y_beating_60_40Overlapping monthly windowsDisclosurefixmisleadingThe spec promised the independent-period count beside the share; not published (run.py:185)
L19SIP outcome distributionsEngine checkpoints, after-tax XIRRRecomputesoundfineNifty 500, 10-year from Jan 2010: XIRR 11.0254%, matches
L20SIP replication tableNifty 50 TRI, no costsRecomputesoundfineMatches the log (207 windows, minimum 3.85%, median 13.67%)
L21SIP paired tests (date of month, dynamic, pause, exemption): published Newey-West tpaired(), NW lags = 12hLags vs n; simulationfixwrong numberLags ≥ number of windows gives t = 38.3 (15y, 78 windows) and 73.1 (20y, 18 windows). An AR(1) null simulation (ρ 0.97, n 138, lags 120) rejects 53% of the time at a nominal 5%
L22S1 fails (date of month matters)Paired mediansTurn-of-month effectrethinkweakThe effect is real on non-overlapping months (25th→1st return 0.85%, t 5.2). Per-horizon significance leans on L21. 6 dates × 12 portfolios × 6 horizons
L23S2 holdsShare of windows below the FD proxyFD constructionfixweakThe proxy accrues a floating T-bill yield, not the spec’s one-year locked rate. The 15-year result is one independent period
L24S3 fails (dynamic SIP lost: −0.26 CAPE, −1.3 yield gap)Dynamic multiplierSignalfixwrong numberThe yield-gap leg inherits L6; t-values inherit L21
L25S4, S5Shares over overlapping start months—soundweakDescribe history, not probabilities

IPO, BAF, MF stress

IDWhat it claimsMethodChecks runVerdictSeverityEvidence
I1IPO sample (mainboard)Exclusion file: no issue price, excludedSelectionfixmisleadingAbout 9 of the 17 mainboard issues dropped for “no issue price” are real IPOs with a known price band (NSDL, Sterling & Wilson Solar, DCX, …). The final price is almost always the top of the band
I2Mean buy-and-hold abnormal returns (1, 3, 5 years)MeansClustered bootstrap; trimmingsoundweak95% CIs: 1y −0.8% to +14.6%, 3y −2.4% to +42.5%, 5y −0.5% to +124%. Without the top 5 issues: 3.4%, 4.5%, 7.7%. The 5-year sample is the 2016–21 cohorts (2021 is 36 of 143). Intervals should be published
I3H3: hot markets (null)WelchMonth-level NW; runssoundweakNW t −2.23 vs Welch −1.91, but “hot” months come in only 6 runs
I4H1 calendar-time portfolioMonthly portfolio, NW 3 lagsGapssoundfine124 months, complete
I5“Only one mainboard issue stopped trading”Price panelISIN chainingfixwrong numberInherits R2: really zero
B1BAF net-equity headline (AUM-weighted)Validation on screened schemesCoveragefixmisleadingValidation covers 56% of AUM; HDFC BAF (32% of AUM) is screened out of validation but carries the largest weight in the headline
B2Per-scheme valuation linkNon-overlapping 3-month changesIndependencesoundweakMedian r −0.19 (25 of 32 negative); schemes correlate 0.35 with each other, so these are not 32 independent votes
B3Arbitrage control near zeroQuarterly mediansDriftsoundweakLatest 0.073 (highest); watch
S1Small-cap stress test: median days to liquidate fell from 8 to 5 (Dec 2024 → Aug 2026)Category medianCompositionfixmisleadingSame 27 schemes: 8 → 8. The fall comes entirely from 9 new funds (median 1 day)
S2Stress data integrity—Duplicates, orderingsoundfine—

Data layer (prices, breadth, TRI, assets, density, dossiers)

IDWhat it claimsMethodChecks runVerdictSeverityEvidence
D1Index TRIs are repaired against the price indexrepair_tri replaces the TRI’s return with the price return on bad daysRe-ran the 11 repairs; synthetic bad-price testfixweak (latent)All 11 repairs are genuine TRI errors (Smallcap 100, 2005–09). But a bad price print gets copied into a clean TRI: a reverting −20% price print puts a −19.9% day into the TRI (tri_check.py:50,65)
D2NSE TRI and price histories are cleanUsed as publishedIndex days above 8% while Nifty 500 moved less than 2% (55 hits)fixwrong numberNifty Metal on 2006-11-21: +31.6% in both price and TRI, while its constituents rose 2.4–2.9%. Published 2006 return 98.1%; should be about 50.5%. The other 54 hits match real events. Nifty Energy and PSE on 2002-04-01 are unexplained
D3Gold in rupees including import dutyWGC INR from 2005; USD × USD/INR before thatDuty premium by year; stale runsfixwrong number (volatility, drawdown, signals)Same root cause as L4. INR price frozen 2020-03-23 → 06-02 (46 sessions) while USD gold rose 1605 → 1730; also 22 sessions in 2016 and 20 in 2021. Duty premium itself checks out
D4Cash = 91-day T-bill accrual, no look-aheadYield known the day before; y·dt/365 compounded dailyAnnual return vs rolloverfixweakLag correct. Daily compounding of a 91-day simple yield overstates by about 6 bp a year at 7%: 2024 cash 7.04% vs 6.92% from quarterly rollover (assets.py:135)
D5Asset panel on NSE sessionsUnlimited forward fillStaleness detectionfixweakassets.py:150 has no fill limit, so the density snapshot’s stale-asset rule can never fire
D6Equity TRIsNSE TRIDays with |r| above 15%soundfineOnly 2009-05-18 (election day, genuine)
D7Survivorship-free adjusted pricesBhavcopy, rename chains, corporate actionsCode; QA listssound, except R2fineCorporate-action logic careful. Rename chaining gap: R2
D8Breadth liquidity floor in constant rupeesCPI of the bar’s own monthLagsoundweakOne month of CPI look-ahead (about 0.5% on the floor) (breadth.py:63)
D9Index breadth, calendar-year tables, trend breadth “across NSE indices”; Nifty 500 from 1995Today’s catalogue with back-filled historiesLaunch dates; disclosurerethinkmisleadingIndices count before their launch. Hindsight-built strategy indices vote in site_dossiers.trend_breadth (unlike highs_lows). Back-fill is not disclosed in indices.md, breadth.md or dossiers.md
D10Market value / GDPMonth-end market value ÷ latest complete FY nominal GDPApril steps; vintagesfixmisleadingSteps each April on the denominator alone: Mar-25 141.3% → Apr-25 132.5% while market value rose 2.6%. April–May use FY GDP published only about 31 May; revised vintages throughout; the 2011-12 and 2022-23 GDP bases are joined (density_mcap.py:117)
D11Flows since Dec 2023Validated sessions; dropped sessions count as zeroDropped-session countsound (disclosed)weak34 sessions dropped; FY26 has 229 sessions; counts published
D12Index dashboard returns, trend, momentumCompleted months; stale for more than 10 days droppedCodesoundfineAnnualises only 3 years and longer
D13Risk dossier “what followed each volatility quintile”Expanding quintile edgesIndependence countfixweak“Independent” = sessions ÷ 63 (site_dossiers.py:89), though quintile days cluster into episodes (Q5: 977 sessions labelled 15.5)
D14Dossier trend breadthShare of indices above their 10-month averagePartial monthfixweakThe last point is month-to-date (site_dossiers.py:124-127)
D15Fund-house leagueDrops periods below 50% of the previous totalCompletenesssuspicionweakPeriods 50–99% complete pass (density_funds.py:145)
D16Macro, world, commodities, ETF density figuresLatest vintagesGapssoundfine—
D17Price QA listsReview only—soundfine—

Site-wide multiple testing (D3, audit_deepening_spec.md)

Every confirmatory test on the site with a p-value:

Method: Holm–Bonferroni at familywise 0.05 across 35 tests, and Benjamini–Hochberg at a 10% false-discovery rate.

RankSpecTestpHolm barHolmBH (10%)
1site_calculationsC3 returns predict DII flows<1e-6 (p (<1e-6))0.0014passpass
2site_calculationsC3 returns predict FPI flows<1e-6 (p (<1e-6))0.0015passpass
3momentum_barometerMO-H1 blended stock momentum premium1.1e-06 (p)0.0015passpass
4market_moodFG-H4a extreme fear then higher 63-session return0.0061 (p)0.0016—pass
5trend_barometerTR-H4 score adds to B0 for forward volatility0.024 (p)0.0016——
6audit_deepeningD1 log_pb 5y (Nifty 50 real)0.025 (bootstrap p)0.0017——
7market_moodFG-H2 score adds to B0 for a 5% fall0.031 (p)0.0017——
8momentum_barometerMO-H3 crash flag (max of Welch and bootstrap p)0.035 (p)0.0018——
9audit_deepeningD1 log_cape 5y (Nifty 50 real)0.055 (bootstrap p)0.0019——
10audit_deepeningD1 dy 1y (Nifty 50 real)0.061 (bootstrap p)0.0019——
11audit_deepeningD1 log_pb 3y (Nifty 50 real)0.064 (bootstrap p)0.0020——
12audit_deepeningD1 dy 3y (Nifty 50 real)0.065 (bootstrap p)0.0021——
13audit_deepeningD1 log_pb 1y (Nifty 50 real)0.093 (bootstrap p)0.0022——
14site_calculationsC2 breadth thrust, 6-month forward return0.107 (p)0.0023——
15market_moodBB-H3 bull-bear adds to B0+FG for a 10% fall0.178 (p)0.0024——
16audit_deepeningD1 log_cape 3y (Nifty 50 real)0.186 (bootstrap p)0.0025——
17audit_deepeningD1 log_pe 1y (Nifty 50 real)0.215 (bootstrap p)0.0026——
18audit_deepeningD1 log_cape 1y (Nifty 50 real)0.249 (bootstrap p)0.0028——
19audit_deepeningD1 dy 5y (Nifty 50 real)0.267 (bootstrap p)0.0029——
20ipoH3 hot-month issues trail0.270 (p)0.0031——
21audit_deepeningD1 log_pe 3y (Nifty 50 real)0.285 (bootstrap p)0.0033——
22market_moodFG-H3 score adds to B0 for a 10% fall0.287 (p)0.0036——
23audit_deepeningD1 log_pe 5y (Nifty 50 real)0.318 (bootstrap p)0.0038——
24site_calculationsC3 DII flows predict returns0.332 (p)0.0042——
25trend_momentumSector momentum mom6_1 vs equal weight (deflated, 2 trials)0.345 (1 - DSR)0.0045——
26site_calculationsC3 FPI flows predict returns0.354 (p)0.0050——
27ipoH1 IPO calendar-time alpha vs Nifty 5000.370 (p (NW t 0.89))0.0056——
28momentum_barometerMO-H5 de-duplicated sector momentum (deflated, 3 trials)0.433 (1 - DSR)0.0063——
29trend_momentumSector momentum mom12_1 vs equal weight (deflated, 2 trials)0.674 (1 - DSR)0.0071——
30market_moodFG-H1 score adds to B0 for forward volatility0.744 (p)0.0083——
31trend_barometerTR-H4 score adds to B0 for a 10% fall0.767 (p)0.0100——
32trend_momentumBest trend rule vs buy-and-hold (35 trials, deflated)0.860 (1 - DSR)0.0125——
33ipoH2 listing gain vs 1-year BHAR0.950 (p)0.0167——
34trend_barometerTR-H3 ensemble beats buy-and-hold (deflated, 84 trials)0.984 (1 - DSR)0.0250——
35market_moodFG-H4b extreme greed then lower 63-session return0.993 (p)0.0500——

Reading:

Summary of Phase 0

Wrong numbers to fix first:

Misleading:

Outcomes (Phase 1, 2026-10-03)

Numbers are from this worktree after the fixes (data to 1 Oct 2026). Old and new figures for each pre-registered result are in that spec’s results log.

RowsOutcome
T1, T2, T3, BT1Fixed. Trades at the close of t+1, completed month-ends only, calendar-time CAGR and costs; the barometer’s compensating lag removed. Midcap 150 sma200d 18.0 → 16.2% (buy-and-hold 15.0%); Smallcap 250 18.1 → 15.5% (13.4%). Verdicts unchanged. Logged in trend_momentum_spec.md and trend_barometer_spec.md
T5Fixed. The trend after-tax engine is the lab’s tax engine. Large caps after tax: behind buy-and-hold by 1.1–5.2 points (was 1.9–6.1)
T6Fixed: eligibility, equal-weight costs, drifted-weight turnover, DSR variance. DSR 0.41/0.73 → 0.33/0.66. Back-fill still open (needs launch dates; request to lane 1)
V3, V7Fixed. Earnings yield and yield gaps on the consolidated basis; the root-level valuation copies are maintained again. Lab yield-gap glide 10.74 → 10.32% after tax; S3’s yield-gap leg −1.30 → −1.57 points
E1Superseded by D1: no valuation predictor passes. The C1 “nine in ten subsamples” sentence is withdrawn
E2, E3, BT3Fixed (D2). 90% intervals and noise and “too few” flags on every what-next table; non-overlap flag on the C1 scatter
E4Fixed. C2 enters at the next close; 6-month p 0.09 → 0.11
L8, L14, L18Fixed (labels and disclosure). interval_basis meta; independent_5y_periods
L21, L22, L24Fixed. SIP Newey-West t only where windows ≥ 3 × lag, with fixed-b critical values. S1 and S3 verdicts stand
R1Fixed. Attribution to Nifty200 Momentum 30 withdrawn (excess-return correlation 0.57); amendment logged
R4, R5Fixed. MO-H1 long-only 26.0% vs 13.7% on matched months; holiday-safe month completion
R8, R7, R9Logged in market_mood_spec.md
D4Fixed (lane 1 file, one-line clear bug, flagged): cash compounds per 91-day bill
D1–D3, D5, D8–D15, R2, I1, I5, B1, S1, L4Requested from the owning lanes, with failing tests (docs/status/quant.md, Requests)
L5, L7, L11, L16, L23Open, weak: noted in the specs; next revision

Independent re-implementations: pipeline/tests/audit/ (13 tests, stdlib, numpy and DuckDB only). They cover:

All reconcile to their tolerances.

This is docs/research/audit.md. The specification was committed before any result was computed; changes after that are logged in it with dates and reasons.