SIP studies: pre-registered specification
Written 2026-10-02 by lane 2, before any SIP result was computed. The engine, costs and tax rules are those of the portfolio lab v2 (portfolio_lab_v2_spec.md, sections 3, 7 and 8 and Appendix B) and the library (portfolio_library.md). Changes after this date go under “Amendments” with the date and the reason.
1. Questions
- For each portfolio in the library, what range of after-tax returns did a monthly SIP deliver over 3, 5, 7, 10, 15 and 20 years, depending on when it started?
- How often did a SIP do worse than the same instalments put in bank fixed deposits, and worse than inflation?
- Do the popular SIP refinements (step-ups, a lucky date of the month, valuation-aware amounts, STPs) change the outcome, and what does stopping a SIP in a crash cost?
- Can we reproduce published Indian SIP studies, and where we differ, why?
2. Pre-registered predictions (graded in the results log)
- S1. The SIP date of the month makes no practical difference: the median after-tax XIRR across start months differs by less than 0.2 percentage points between any two dates, and no difference is significant once overlap is accounted for.
- S2. An all-equity (Nifty 500) SIP had a lower after-tax XIRR than the fixed-deposit SIP in some 5-year windows but in no 15-year window.
- S3. The valuation-aware dynamic SIP beats the plain SIP by less than 0.5 points of XIRR at the median 10-year window, with the interval for the difference including zero.
- S4. Pausing during drawdowns lowers the terminal value relative to an uninterrupted SIP in most windows that contain a drawdown of more than 20%.
- S5. Lump sum beats a 12-month STP in about two thirds of start months for an all-equity portfolio, in line with the US finding of Vanguard (2012, “Dollar-cost averaging just means taking risk later”).
3. The plain SIP
- Instalment: Rs 10,000 a month, bought at the close of the first NSE session on or after the 1st of each month (the date-of-month study varies this).
- Portfolio: every portfolio in the lab v2 library (sections 5.1 to 5.7 of its spec), with its own rebalancing rule. For multi-asset portfolios, each instalment goes first to the sleeves furthest below target (cash-flow rebalancing); the portfolio’s scheduled rebalancing still runs. For tactical rules, instalments follow the current target weights.
- Costs and tax: as in the lab spec, including stamp duty on each instalment from July 2020. Every instalment is its own tax lot (FIFO). Tax on gains realised by rebalancing is paid each financial year from the portfolio.
- Exit: the whole portfolio is redeemed at the end of the window, and that year’s tax is settled. The after-tax XIRR uses the instalments as outflows and the after-tax redemption value as the inflow. We also report the pre-tax XIRR and the value before redemption.
- Equity exemption: the annual long-term exemption (Rs 1 lakh from FY2018-19, Rs 1.25 lakh from FY2024-25) is on in the headline SIP results, because a Rs 10,000 SIP investor is the small investor the exemption is meant for, and off in a sensitivity run. (The lab’s lump-sum results, which are scale-free, keep it off.)
4. Windows and overlap
- Horizons: 3, 5, 7, 10, 15 and 20 years (36 to 240 instalments; the window ends one month after the last instalment).
- Start months: every month from the portfolio’s first usable month to the last month that leaves a full horizon. “First usable” means every asset it holds exists and its rule has finished warming up.
- Two views:
- Common starts from April 2005, for comparing portfolios on the same windows.
- Longest history for the portfolios that allow more (all-equity Nifty 50 from July 1999, Nifty 500 from 1995, and the n500 / G-sec / gold mixes from about 2002).
- Overlap: windows that start a month apart share almost all their months, so the many windows are far from independent. For each horizon we state the number of non-overlapping windows that fit in the sample (the effective number of independent periods) next to every statistic. Percentiles across start months are descriptions of history, not probabilities; we do not attach confidence intervals that treat overlapping windows as independent. Where we compare two SIP variants, we test the mean of the paired differences with a Newey-West standard error with lags equal to the horizon in months, and we also report the result on the non-overlapping subsample.
5. What is reported for each portfolio and horizon
- The distribution of after-tax XIRR across start months: 5th, 10th, 25th, 50th, 75th, 90th and 95th percentiles, the worst and best with their start months, and the mean.
- Share of windows below fixed deposits: the after-tax XIRR of the SIP against the after-tax XIRR of the same instalments placed in one-year fixed deposits rolled over each year, with interest taxed each year at 30% plus cess. The FD rate is the scheduled commercial banks’ deposit rate for more than one year (maximum, RBI, weekly) where available; until lane 1 adds it to the warehouse, and before November 2011 in any case, the 364-day T-bill primary yield, and the 91-day yield before the 364-day series starts. The proxy is stated wherever the result appears.
- Share of windows below inflation: the real after-tax XIRR (each cash flow deflated by the all-India CPI of its month) is below zero.
- What Rs 10,000 a month became: for each start month, the after-tax value at the end in rupees, and in rupees of the window’s final month (CPI-deflated), with the total invested.
- Pre-tax XIRR and the tax drag.
6. Variants
Run on a headline set of 12 portfolios: n50, n500, mid150, small250, sixty_forty, sixty_twenty_twenty, equal_thirds, permanent, all_weather, gtaa, dual_momentum, factor_four.
| Variant | Rule |
|---|---|
| SIP against lump sum | For each start month and horizon, the SIP’s XIRR against the CAGR of investing the same total at the start. This is the comparison people usually make, but the two are not alternatives for the same investor; the STP study below is the fair one. |
| Step-up SIP | Instalment grows 5% (and 10%) each year on the anniversary of the first instalment. |
| Date of month | First session on or after the 1st, 5th, 10th, 15th, 20th, 25th and 28th. Paired against the 1st. |
| Dynamic SIP (valuation-aware) | Cash-neutral: the investor sets aside Rs 10,000 a month. The amount invested is 1.5× when the Sensex CAPE10 (IIMA, previous month-end) is at or below its expanding 33rd percentile, 1× in the middle, and 0.5× at or above the 67th percentile. Money not invested waits in the liquid sleeve (taxed as a liquid fund), which funds the 1.5× months; if it runs short, only what is there is invested. Everything, including the buffer, counts at the end. A second version uses the Nifty 500 earnings yield minus the 10-year G-sec yield. Applied to n50 and n500 only. |
| STP from liquid to equity | A lump sum of Rs 12 lakh starts in the liquid sleeve and moves to the equity portfolio in 6 (and 12) equal monthly transfers, against investing it all on day one. Compared at 1, 3, 5 and 10 years. Equity portfolios in the headline set only. |
| Pausing in drawdowns (the behaviour gap) | The SIP stops when the Nifty 500 is more than 20% below its previous peak at a month-end and restarts when the drawdown is back above −10%. Two behaviours: (a) the skipped instalments stay in the liquid sleeve, never invested in equity; (b) the backlog is invested in one go on restart. Both against the uninterrupted SIP. |
| Exemption off | Section 3’s exemption switched off. |
7. Replication checks
The engine is run in a “replication mode” (no costs, no tax, the published index) and compared with these published studies:
| Study | What they report | Our closest reproduction | Expected differences |
|---|---|---|---|
| Freefincal (Pattabiraman 2020): Nifty 50 TRI, Jul 1999 to Jan 2020 | 10-year rolling SIP XIRR fell from above 20% (2009 endings) to about 10% (2020 endings) | Nifty 50 TRI, same dates, 10-year rolling XIRR by end month | instalment date conventions |
| WhiteOak Capital (2024): Sensex TRI, Sep 1996 to May 2024 | minimum SIP XIRR 4.6% at 10 years, 7.4% at 15; 79% / 90% of windows above 12% at 10 / 15 years | Nifty 50 TRI from Jul 1999 (we do not hold the Sensex TRI) | different index and a shorter sample: we expect the same shape, not the same numbers |
| Geojit (2024): Sensex, Sep 1996 to Jul 2024 | 10-year median 14.3%, minimum 4.1%; 15-year median 14.3%, minimum 7.0% | as above | as above, and Geojit does not say whether it used the TRI |
A reproduction counts as a match when the same index and dates give a median within 0.5 points and a minimum within 1 point. Where the index differs, we report both and explain the gap; we do not tune anything to close it. A study whose numbers we cannot reproduce on the same index is reported as such.
8. Outputs
Derived tables .cache/derived/sip_*.parquet and bundles under lab/sip/*, documented in FRONTEND_HANDOFF.md (lane 2 section): the outcome distributions by portfolio and horizon, the per-start-month XIRR series for the explorer, and one table per variant.
9. Known limits
- The FD comparison uses a T-bill proxy until deposit-rate data lands.
- 20-year windows exist only for starts up to about 2006, so the 20-year results rest on very few independent periods (one or two).
- Index funds on mid caps, small caps and factors did not exist for most of the sample; those SIPs are hypothetical before the funds launched.
Amendments
None yet.
Results log
2026-10-02: first run (starts April 2005 to September 2023; data to 2026-09-30)
136 portfolios × up to 222 start months × six horizons: 112,608 SIP windows, plus the variants and the longest-history view. Code: pipeline/tipsheet/lab/sip.py; tables .cache/derived/sip_*; bundles lab/sip/*. All XIRRs are after costs and tax, on full redemption at the end of the window, with the equity exemption on. The FD hurdle is the 364-day T-bill proxy (deposit-rate data not yet in the warehouse), after 30% tax and cess each year: its median after-tax yield was 4.5% to 4.9% depending on the horizon. “Independent periods” is the number of non-overlapping windows that fit in the sample; percentiles describe history and are not probabilities.
The plain SIP (Rs 10,000 a month), common starts from April 2005.
| Portfolio | Horizon | Independent periods | 5th pct | Median | Worst (start) | Below FD | Below inflation |
|---|---|---|---|---|---|---|---|
| Nifty 500 | 5 years | 4 | 3.9 | 13.5 | −6.1 (2015-04) | 8.1% | 20.2% |
| Nifty 500 | 10 years | 2 | 9.1 | 12.8 | 3.4 (2010-04) | 0.7% | 0.7% |
| Nifty 500 | 15 years | 1 | 8.9 | 12.9 | 6.1 (2005-04) | 0% | 1.3% |
| Nifty 50 | 10 years | 2 | 8.5 | 12.0 | 3.4 (2010-04) | 0.7% | 1.4% |
| Smallcap 250 | 5 years | 4 | −2.9 | 16.7 | −17.8 (2015-04) | 23.2% | 28.3% |
| Smallcap 250 | 10 years | 2 | 6.7 | 14.5 | −1.1 (2010-04) | 2.9% | 3.6% |
| 60/40 | 5 years | 4 | 5.5 | 11.2 | 0.6 (2015-04) | 3.5% | 19.7% |
| 60/40 | 10 years | 2 | 9.0 | 10.9 | 5.9 (2010-04) | 0% | 0% |
| 60/20/20 | 5 years | 4 | 7.4 | 11.4 | 1.1 (2015-04) | 1.5% | 11.1% |
| Equal thirds | 5 years | 4 | 6.6 | 10.5 | 4.5 (2015-04) | 0.5% | 7.6% |
| Permanent | 5 years | 4 | 6.6 | 9.3 | 4.7 (2015-04) | 0% | 7.1% |
| GTAA | 10 years | 2 | 5.7 | 7.8 | 5.4 (2009-08) | 0% | 0% |
| 100% G-sec | 5 years | 4 | 5.0 | 6.4 | 4.7 (2018-11) | 0% | 44.9% |
- Equity SIPs had a wide spread over 3 to 5 years and a narrow one at 10 to 15. The worst windows cluster on the same calendar: SIPs started in April 2010 and redeemed in the March-April 2020 crash were the worst 10-year outcome for every equity-heavy portfolio, and SIPs started in April 2015 and redeemed in 2020 the worst 5-year ones.
- Multi-asset mixes shrank the bad tail far more than the median. At five years, 60/20/20’s 5th percentile was 7.4% against the Nifty 500’s 3.9%, for a median 2 points lower. Equal thirds and the permanent portfolio never did worse than the FD proxy at 5 years or more except in one window, but beat inflation by less.
- Twenty-year results rest on 18 start months (April 2005 to September 2006): one independent period. They are shown, but say almost nothing about the next 20 years.
- Longest history (Nifty 500 from 1995, Nifty 50 from 1999): the 10-year Nifty 500 SIP had a median of 13.8% and a worst of 3.4% across 260 windows (three independent periods); the 20-year median was 14.0% with a worst of 10.7% (140 windows, one independent period).
Variants (headline set of 12 portfolios).
- Date of month: not a null. SIPs dated the 20th, 25th or 28th beat those dated the 1st in most windows, by a median of up to 0.75 points of XIRR at 3 years (mid and small caps), 0.2 to 0.45 at 5 years and under 0.21 at 10 years. The gain shrinks roughly as one over the horizon, which is what a fixed price advantage per instalment produces. The source is a turn-of-month effect in Indian indices: from the 25th to the next 1st, the Nifty 500 TRI rose 0.85% on average (2005-2026) against 0.26% expected from its average daily drift (t = 5.2; Midcap 150 1.16%, t = 7.1). Days 5 to 15 differ from the 1st by at most 0.14 points at any horizon.
- Step-ups (5% and 10% a year) barely move the XIRR (the Nifty 500 10-year median stays at 12.8%) but raise the final corpus, by about a fifth at 10 years for a 10% step-up, because more money is invested.
- SIP against lump sum. At 5 and 10 years, for equity portfolios the SIP’s XIRR beat the lump sum’s CAGR in 54% to 67% of windows (the range is wider, 22% to 89%, at 3, 15 and 20 years, where windows are fewer) (Nifty 500 at 10 years: median gap +0.6 points). For multi-asset portfolios it was a coin toss. As the spec says, this compares two different investors.
- STP against lump sum (Rs 12 lakh). Investing all of it on day one beat a 12-month STP in 59% to 70% of start months, and beat a 6-month STP in 58% to 64%. At the median a 12-month STP ended 3% to 7% poorer, depending on the asset and horizon.
- Valuation-aware dynamic SIP: it lost. Investing 1.5× when the CAPE was in its cheapest third and 0.5× in its dearest, with the difference parked in a liquid fund, lowered the median 10-year XIRR by 0.26 points (CAPE) and 1.3 points (yield gap) against the plain SIP, and ended with Rs 0.4 lakh to Rs 1.5 lakh less at 10 years. The CAPE was in its dearest third in 61% of months from 2014, so the rule kept money in the liquid fund through a rising market.
- Pausing in drawdowns cost money in almost every window. Stopping when the Nifty 500 fell 20% below its peak and restarting when it recovered to −10%, with the skipped money left in a liquid fund, cut the 10-year XIRR by 1.5 points at the median and the corpus by about Rs 1.8 lakh on Rs 12 lakh invested; the uninterrupted SIP won in more than 99% of 10-year windows. Investing the backlog on restart halved the damage but never removed it.
- Tax. Tax cost the plain SIP 0.4 to 0.6 points of XIRR at 10 years for buy-and-hold and static portfolios, and 1.1 to 1.4 for GTAA and dual momentum. The equity exemption was worth only 0.04 to 0.09 points at 10 years: the final redemption is one large sale in one year.
Replication (Nifty 50 TRI, no costs, no tax, instalment on the first session of each month from July 1999).
| Horizon | Windows | Min | Median | Share above 12% |
|---|---|---|---|---|
| 10 years | 207 | 3.9% | 13.7% | 73% |
| 15 years | 147 | 6.9% | 13.4% | 86% |
| 20 years | 87 | 10.9% | 13.6% | 91% |
- Freefincal (Nifty 50 TRI): reproduced. Our 10-year SIP ending December 2009 returned 21.5% and the one ending January 2020 11.6%; theirs fell from above 20% to about 10% over the same endings. The April 2020 ending gives 3.8%.
- WhiteOak (Sensex TRI from 1996; 10-year minimum 4.6%, 15-year 7.4%, 79% and 90% above 12%) and Geojit (Sensex; 10-year median 14.3%, minimum 4.1%; 15-year median 14.3%, minimum 7.0%): same shape, slightly lower numbers. The differences are what we expected: a different index, a sample that starts in 1999 instead of 1996 (missing the strong 1996-99 starts) and ends in 2026 instead of 2024 (adding the weak 2024-26 endings). We did not tune anything to close the gap. On the same index and dates we could not test them, so neither counts as a match under section 7.
Predictions.
- S1 fails. The date of the month mattered more than 0.2 points at horizons up to 5 years (late-month dates better), through a genuine turn-of-month effect. At 10 years and beyond it is under 0.21 points. Two cautions: we tested six dates on twelve portfolios and six horizons, and an effect this well known (Ariel 1987; Lakonishok and Smidt 1988) may weaken now it is visible. Fund houses also apply the NAV of the day the money is realised, which can shift a “25th” SIP by a day or two.
- S2 holds. The Nifty 500 SIP fell below the FD proxy in 8.1% of 5-year windows and in none of the 15-year windows.
- S3 fails in the other direction. The dynamic SIP did not beat the plain SIP at all; it lost 0.26 points (CAPE) and 1.3 points (yield gap) at the median 10-year window.
- S4 holds. Pausing lowered the terminal value in nearly every window that contained a 20% drawdown.
- S5 holds. Lump sum beat a 12-month STP in about two thirds of start months (59% to 70%).
2026-10-03: audit corrections (register rows L21–L24, V3, V7, D4); rules unchanged
- Newey-West t-values (L21): the paired tests used a Newey-West t with lags equal to the horizon in months even when there were fewer windows than lags. Values of 38 and 73 were published; with these sizes an AR(1) null rejects 53% of the time at a nominal 5%.
- Where n is too small: the t is now published only where the windows number at least three times the lag. At 7 years and beyond, the non-overlapping columns are the evidence.
- Where it is published: it comes with the fixed-b critical value
t_crit_fixed_b(Kiefer and Vogelsang 2005; 2.45 to 2.88 here), not 1.96.
- S1 (date of month): fails, unchanged. At 3 and 5 years, 50 of 144 date comparisons exceed their fixed-b critical value (62 exceeded 2 before). The effect is the turn-of-month return (late-month dates better by up to 0.48 points at 3 years). Beyond 5 years there is no valid t, and the non-overlapping differences are under 0.2 points.
- S3 (dynamic SIP): fails in the other direction, by more. The yield-gap signal now uses the consolidated P/E from a maintained table (V3, V7).
- Yield-gap dynamic SIP at the median 10-year window: −1.57 points (Nifty 500; was −1.30) and −1.43 (Nifty 50; was −1.25).
- CAPE version: −0.27 and −0.29 (were −0.26 and −0.26).
- The old t-values (−1.6 to −2.6) were invalid. With 2 non-overlapping 10-year windows, the honest statement is the median and the non-overlapping means (−1.0 to −1.4).
- S2, S4, S5: unchanged.
- FD proxy (L23), noted for the next revision: the proxy accrues a floating T-bill yield, not the one-year locked rate this spec describes. It stays labelled “FD proxy” until lane 1’s deposit-rate series arrive.