Portfolio lab and SIP studies: methods
This note explains how the portfolio lab (v2) and the SIP studies are built. The full pre-registered rules are in docs/research/portfolio_lab_v2_spec.md and docs/research/sip_studies_spec.md; the code is in pipeline/tipsheet/lab/ (design: docs/research/portfolio_library.md).
What the lab is
A library of about 140 asset-allocation portfolios, each described by a short declarative entry: its family, its source in the literature, its construction and its rebalancing rule. Every portfolio is run through the same engine on the same window, with the same costs and the same Indian tax rules, so the comparison is fair. The parameters come from the papers that proposed each rule, not from fitting Indian data, and the whole set was written down before any result was computed.
Assets
All in rupees and with dividends or coupons reinvested:
- Indian equity: NSE total-return indices (Nifty 50, 100, 500, Next 50, mid and small cap, equal-weight, and NSE’s factor indices). Bad prints in the index files are repaired against the price index.
- Bonds: the NSE 5-year benchmark G-sec index, and a synthetic 10-year constant-maturity G-sec built from month-end yields (used only where a portfolio calls for long bonds).
- Gold: the domestic rupee price including import duty (what Indian gold ETFs track).
- Cash: the 91-day T-bill yield, accrued daily, standing in for a liquid fund.
International equity is not in the headline window: global index levels cannot be republished, and rupee international funds start only in 2011.
Window
From 1 April 2005 to the latest date, for every portfolio. Rules use earlier data to warm up. A longer window from October 2002 is reported for the portfolios whose assets and rules allow it.
Timing
Rules decide at month-end closes (static portfolios at the December close) and trade at the next session’s close. A rule only ever sees data up to its decision date; a test rebuilds every portfolio’s decisions on data cut at several dates and checks that nothing changes.
Costs
Each asset’s total-return index is reduced every day by the all-in annual cost of the cheapest widely available index fund or ETF of the time (direct plans from 2013). The cost schedule was checked against the actual tracking difference of every Indian index fund and ETF in AMFI’s NAV history, and replaced by the observed figure wherever the two differed by more than a quarter of a point. Trades pay a spread, stamp duty from July 2020, and securities transaction tax on equity-fund redemptions.
Tax
The tax engine applies Indian capital-gains rules as they stood on each sale date: the pre-2004 rules, the STT era (long-term equity gains exempt from October 2004 to March 2018), section 112A with the January 2018 grandfathering, indexation with the Cost Inflation Index for debt and gold until July 2024, the 2014 change to a 36-month holding period for debt funds, the 2023 “specified mutual fund” rule, the July 2024 rates, and the 2025 narrowing of the specified-fund definition. Every purchase is a separate tax lot, sold first-in first-out. Losses are set off and carried forward as the law allows. Tax is paid each April from the portfolio, and everything is sold at the end so that deferred tax is counted too.
The investor is a resident individual in the 30% slab with cess and no surcharge. The lump-sum results ignore the annual equity exemption (it depends on the investor’s other gains); the SIP results include it, because a Rs 10,000-a-month investor is exactly who it is for.
What the numbers mean
- CAGR after tax is the growth rate of a lump sum that is fully redeemed and taxed at the end.
- Before tax numbers are after costs.
- Risk numbers (volatility, Sharpe, Sortino) use month-end returns; drawdowns use daily values.
- SIP XIRR is the internal rate of return of Rs 10,000 instalments and the after-tax redemption value.
How we guard against luck
With about 140 portfolios, the best-looking one will look good partly by chance. We report:
- The deflated Sharpe ratio: the probability that a portfolio’s Sharpe ratio beats what the best of 140 random strategies would show.
- The probability of backtest overfitting: how often the in-sample winner falls below the median out of sample when the history is cut into halves every possible way.
- Block-bootstrap intervals: resampling the history in year-long blocks to show how uncertain each figure is.
- A walk-forward test: what you would have earned by picking last five years’ best mix each January.
Rolling and SIP windows overlap heavily. A SIP that starts in March and one that starts in April share almost all their months. So we state the number of independent periods next to every distribution and never treat overlapping windows as independent evidence.
Caveats
- Twenty-one years is three or four equity cycles; most differences between sensible portfolios are not statistically distinguishable.
- Mid-cap, small-cap and factor indices are back-tested by NSE before their launch dates; the lab marks the live date and reports the live record separately.
- Gold’s strong run over the window owes much to rupee depreciation and the 2012-13 import duty increases.
- The tactical rules would pay exit loads on some funds, which are not modelled.