Trend barometer (an ensemble): pre-registered specification
Written 2026-10-02, before any result below was computed; the commit is the freeze. Changes after results are logged at the bottom with the date and reason. Brief: docs/briefs/barometers.md. Builds on trend_momentum_spec.md, where five single trend rules on seven assets failed to beat buy-and-hold after deflation.
Why an ensemble
- Specification risk. The best single lookback in a backtest is close to random out of sample. Averaging across signal types and horizons gives up the hope of picking the winner for a result near the middle of the pack, with a narrower spread.
- Timing luck. A rule checked once a month gives different results depending on which day it checks. Hoffstein, Faber and Braun (“Rebalance Timing Luck: The (Dumb) Luck of Smart Beta”, Newfound Research, 2020) measured gaps above 100 bp a year for monthly factor indices. Newfound’s remedy is tranching: split capital across overlapping sleeves that rebalance on staggered days (their ensemble trend indices use 20 daily-staggered tranches; “Ensemble Multi-Asset Momentum”, 2019).
- Signal types are near-equivalent. Levine and Pedersen (“Which Trend Is Your Friend?”, 2016) show time-series momentum and moving-average rules are weighted sums of past returns, so most of the diversification is across horizons.
- What an ensemble does not do: raise the expected return. The page says so. Niels Kaastrup-Larsen’s Trend Barometer (Top Traders Unplugged) is cited only as a reference reading of trend-environment strength; its method is not public and is not replicated.
Assets
Nifty 500, Nifty 50, Nifty Midcap 150, Nifty Smallcap 250 (total-return indices); every NSE sectoral index with at least 252 sessions of history; gold in rupees; the NSE 5-year G-sec index; the rupee (US dollars per rupee from ECB reference rates, from 2000, so up means a stronger rupee). Total-return assets are converted to excess-return series (asset total return over the 91-day T-bill cash index) before any signal is computed, so that a G-sec index rising on carry alone does not read as a trend. The rupee is used as a spot rate (no carry).
The grid (24 signals per asset)
Horizons h = 21, 42, 63, 126, 189, 252 sessions (about 1, 2, 3, 6, 9, 12 months). On session t, with P the excess-return series:
| Code | Signal type | “Up” when |
|---|---|---|
tsmom | Time-series momentum | P(t) > P(t−h) |
sma | Price against its average | P(t) > mean of P over the last h sessions |
xo | Moving-average crossover | mean of P over the last max(5, h/4) sessions > mean over the last h |
chan | Breakout channel (position in range) | P(t) > midpoint of the highest and lowest P over the last h sessions |
Score = share of the 24 signals voting up, × 100. Equal weight per horizon band (each band has four signals, so this is the plain mean). No fitted weights. The page shows the 4 × 6 grid as a heat table, so fragility is visible: a score of 50 made of a solid split by horizon reads very differently from a scatter.
Cross-asset trend environment: for each asset, strength = |2 × score − 100| (0 = signals split evenly, 100 = all agree). The environment reading is the mean strength across the broad-index, gold, G-sec and rupee series plus the median sector; direction is shown separately.
Strategies used to test it (long or cash, per asset)
- Signals at the close of t; position changes at the close of t+1. No look-ahead.
- Single specifications: 24, each long when its signal is up.
- Ensemble: exposure = score / 100, the rest in cash.
- Monthly-checked variants (for timing luck): each single spec and the ensemble evaluated only every 21 sessions, on each of 21 offsets; plus the tranched version, 21 sleeves each checked on its own offset (equivalent to averaging the offsets’ exposures).
- Costs (as
trend_momentum_spec.md): 0.10% of value traded each way; fund costs equity 0.15% a year, gold 0.50%, G-sec 0.20%, cash 0.20%. Pre-tax only; after-tax is not repeated here because the earlier run showed tax dominates for switching rules.
Hypotheses and measurements
Tested on Nifty 500 (main), with the other assets reported as the same table, not as separate hypotheses.
- TR-H1, robustness (pass bar fixed). Over rolling 5-year windows (stepping one month), rank the ensemble and the 24 single specs by Sharpe in each window. Passes if the ensemble’s worst window rank (as a percentile among the 25) is better than the median single spec’s worst window rank, and its full-sample Sharpe is at or above the median single spec’s. This is the claim: less regret, not more return.
- TR-H2, timing luck (measurement). For the monthly-checked rules: the spread (max − min) across the 21 offsets of CAGR, Sharpe and maximum drawdown, for each single spec, and for the tranched version (which has one outcome). Expected, not tested: spreads of 0.5–2 points of CAGR for the faster specs; the tranched version’s result sits near the offsets’ mean.
- TR-H3, returns (expected to fail). Ensemble against buy-and-hold: Sharpe difference and the deflated Sharpe ratio (Bailey and López de Prado 2014) of the ensemble’s excess over buy-and-hold. Trials counted honestly: 35 from the first run + 24 single specs + 1 ensemble + 24 monthly specs = 84. Passes if DSR > 0.95.
- TR-H4, risk (family with TR-H3; α = 0.05/2 = 0.025 for the regressions). The score adds to the baseline B0 of
market_mood_spec.md(HAR volatility, US VIX, 200-day dummy) in forecastingdd10andlog(rv_fwd)for Nifty 500. Two-sided. Same four pass conditions as the mood spec (full-sample HAC p, both halves, out of sample from 2016, non-overlapping sign). Both targets must pass for H4 to pass; each is reported. - TR-H5, cross-asset (description only). Today’s grid per asset and the environment reading with its history.
“What happened next”
Score bands 0–20, 20–40, 40–60, 60–80, 80–100 for Nifty 500: sessions, independent periods, forward 21/63/252-session returns (median, inter-quartile range, hit rate) and median worst fall, as in the mood spec. Descriptive.
Not tested, stated in advance
No other horizons, signal types, weightings or assets as hypotheses. A volatility-scaled ensemble is not run in v1. Anything tried later is exploratory.
Results log
Implementation notes, 2026-10-02, written before the first run
- Calendar: NSE equity sessions. Gold, the rupee (ECB holidays differ) and sector indices are carried forward at most five sessions.
- “The broad index” in the environment reading is Nifty 500; the other three broad indices are shown but not averaged in, so large caps are not counted four times.
- Sectors: NSE’s sectoral group, excluding indices with no print in the last 10 days or under 252 sessions of history.
- Monthly checks: “every 21 sessions” counts sessions from the first date with a full grid; offset o checks on sessions o, o+21, o+42…
- Found on the first run and fixed (no rule changed):
trend.backtest(lane 2) earns the return from close t to t+1 on the exposure decided at close t, i.e. it trades at the signal’s own close. This spec trades at the close of t+1, so the barometer lags the exposure one more session before calling it. On Nifty 500 the same-close version overstated the 21-session rules by 4–6 points of CAGR (index returns are autocorrelated). Reported to lane 2 indocs/status/site.md; the earliertrend_momentum_spec.mduses the same wording. - Rolling windows (TR-H1): 1,260 sessions, starting at each month’s first session; a strategy’s rank is its Sharpe percentile among the 25 in that window.
2026-10-02: first run (data to 1 Oct 2026), Nifty 500 from 1996-01-18
- TR-H1, robustness: pass. Over 309 rolling 5-year windows, the ensemble’s worst rank was the 24th percentile of the 25 strategies, against the 8th for the median single specification; its full-sample Sharpe 0.56 against the median single’s 0.53.
- TR-H2, timing luck (measured). Checked once every 21 sessions, the median single specification’s CAGR moved by 4.3 points a year depending on the day it checked; the widest spread was 11.5 points. The ensemble’s spread was 3.3 points; the tranched versions land near the offsets’ average.
- TR-H3, returns: fail. Ensemble 14.1% a year against buy-and-hold 14.1%, Sharpe 0.56 vs 0.40, worst fall -33% vs -67%. Deflated Sharpe of the excess over buy-and-hold 0.016 with 84 trials. Pre-tax.
- TR-H4, risk: fail.
dd10p = 0.77;rv_fwdp = 0.024 with opposite signs in the two halves; neither improves out of sample. - Reading: the ensemble behaves as the literature says: it gives up nothing in return against buy-and-hold here and halves the worst fall, and it is the least regrettable choice among the specifications, but it does not beat buy-and-hold after counting trials, and the score adds nothing to a risk forecast.
2026-10-03: rerun after the audit (register rows BT1, T3); rules unchanged
- Timing (BT1):
trend.backtestnow trades at the close of t+1 itself, so this module’s extra one-session lag was removed. The timing is the same as before, so nothing was lagged twice. - Annualisation (T3): CAGR and running costs now use calendar time (about 249 sessions a year, not 252).
- TR-H1: pass. Worst rank 0.28 against 0.06 for the median single (was 0.24 and 0.08). Sharpe 0.56 against 0.53 (unchanged).
- TR-H2: median single CAGR spread 4.2 points (was 4.3), widest 11.3 (was 11.5), ensemble 3.2 (was 3.3).
- TR-H3: fail. Ensemble 13.8% against buy-and-hold 13.8% (was 14.1% and 14.1%). Sharpe 0.56 vs 0.40. DSR 0.016 (unchanged).
- TR-H4: fail.
dd10p = 0.77;rv_fwdp = 0.024, with opposite signs in the two halves (unchanged). - D2 (
audit_deepening_spec.md): the “what happened next” table now carries 90% intervals and an “inside noise” flag.