Dossiers: how the questions are answered
Each question page (/q/<id>/) answers its question every day from the data. This note explains how. The rules were written down before they were first run on the data (docs/site/DOSSIERS_PLAN.md, with every later change logged there), and the code is apps/site/src/lib/dossier/. No sentence on a dossier is written by a model; every sentence is a template filled with numbers from the published data at build time.
The measures and their votes
Each dossier weighs a fixed list of measures. Every measure has a dated history and a rule that turns today’s value into one of three votes: for the question (+1: dear, up, strong, risky, heavy, cheap money, weakening rupee, a headwind), against it (−1), or in between (0).
- Most measures vote by percentile. Today’s value is compared with every earlier reading in its comparable history (the share of readings below it, ties counted half, the same calculation as the site’s readings). At the 70th percentile or above it votes +1; at the 30th or below, −1. Some measures are inverted: a high dividend yield, for example, counts as cheap.
- Some vote on fixed levels, where a level means something in itself: CPI inflation inside 2.5–5.5% (near the RBI’s 4% target) counts as strong and outside 2–6% as weak; the overnight call rate 10 basis points or more below the repo rate counts as cheap money; a trend score of 60 or more out of 100 counts as up. Each such measure states its levels on the page.
- Trend rules vote by state. The pre-registered 10-month rule is up when the last month-end is above the average of the last ten month-ends, and down otherwise.
- Context measures are shown but never vote, such as the unadjusted P/E or the broad dollar.
From votes to an answer
Primary measures count twice, supporting measures once. The score is the weighted average of the votes, from −1 to +1, and maps to five answers: 0.6 or more is the strong end (“Expensive”), 0.25 to 0.6 the leaning end (“Somewhat expensive”), between −0.25 and 0.25 the middle (“About usual”), and the same on the other side.
- Stale data does not vote. A measure is out of date when its latest value is older than 21 days (daily and weekly series), 75 days (monthly) or 200 days (quarterly), counted from the latest NSE session in the data. It is shown, marked “out of date”.
- No answer without enough data. If fewer than three measures can vote (or, where a dossier has fewer than three voting measures, if any of them cannot), or the measures that can vote carry less than half the weight of the primary measures, the page says it cannot answer today and lists what is missing.
- Disagreement is said out loud. When one measure votes +1 and another −1, the answer names the strongest of each, and gives the reason where we know it (for valuation, NSE’s 2021 switch to consolidated earnings and the fall in return on book value).
What has changed
The same rules are run as of one month and one year before the latest session, using only what was known then: each measure’s last value on or before that date, and its percentile within the history up to that date. The page gives the answer then and now, the measures whose votes changed, and the largest move in percentile terms over the year.
What would change the answer
For each primary measure the page gives the value at which its vote would change: the 30th or 70th percentile of its history, or its fixed level. It also counts the fewest single-step vote changes that would move the answer one level, and lists the relevant releases and results dates in the next 14 days.
What history says
Only our own tests and pre-registered studies are quoted, with their numbers read from the published results: evidence C1–C3 (valuation and later returns, breadth thrusts, flows and returns), the trend, momentum and mood barometers’ tests, the trend and momentum models and the IPO study. Failed and untestable results are stated as plainly as passed ones. Where we have not tested a link, the page says so.
Answers that are names
Three questions are not a scale. Who is buying? names the largest net buyer and seller over three months of exchange category flows (validated sessions only), says whether the latest month agrees, and places foreign and domestic institutions’ last 13 weeks against every 13-week span of their history. What is leading? names the leader by size (six-month return relative to the Nifty 500, and whether it is still ahead over one month), the sectors ahead over both six months and one month, and the best and worst asset over a year. Are the products doing what they promise? applies one fixed test to each promise (for example, the median equity ETF trades within ±0.5% of its NAV) and counts how many hold.
The risk dossier’s own series
dossiers/risk_weekly and dossiers/risk_next come from the site_dossiers pipeline step (pipeline/tipsheet/compute/site_dossiers.py):
- Volatility: the standard deviation of daily log total returns of the Nifty 500 over the last 21 or 63 sessions, annualised (× √252), from 1995.
- Fall from peak: the total-return index against its highest close so far.
- US VIX: published only as the share of all earlier daily closes below the day’s close (from 1990); the level is third-party copyright.
- What followed each volatility quintile: each session is placed in a fifth of 21-session volatility using only the volatility known up to that day, after three years of history. For each fifth the table gives the median volatility over the next 21 sessions, the share of sessions followed by a fall of 10% or more within 63 sessions, and the median 63-session return. Windows overlap, so the number of independent periods (sessions ÷ 63) is the honest sample size. This is descriptive: no hypothesis was registered.
The charts
Each dossier carries the charts its argument needs, in an order of reasoning: the main measure over its history, what is driving it, the cross-section today, and what followed similar readings in our own tests (docs/site/DOSSIERS_PLAN.md, Charts). Titles and reading lines are filled from the data at build time, so a title such as “Over the last three years prices rose 16% and earnings grew 34%” changes when the data do.
- Drawn at build time as SVG at three widths; the page shows the one that fits. Hover or drag on a time chart to read every series at a date; every chart has a table view.
- Percentile of own history (the five valuation measures, P/E by index): at each date, the share of readings up to that date below it, so no chart uses hindsight.
- Nifty earnings are not published by NSE Indices. They are implied from its published P/E (index value ÷ P/E, trailing twelve months), with earnings before 31 Mar 2021 scaled to the consolidated basis NSE switched to that day (
docs/methods/valuation.md). The “prices against earnings” chart rebases both to 100 in January 2006 and excludes dividends. - Derived series published by the
site_dossiersstep:dossiers/price_earnings_weekly(the above),dossiers/trend_breadth_monthly(share of NSE indices whose month-end is above their 10-month average; months with fewer than 20 indices left out; the current month is month to date),dossiers/sector_months(each sector index’s monthly total return minus the Nifty 500’s, last 12 months and the month to date),dossiers/assets_year(Nifty 500, gold, 5-year G-sec and cash, rebased a year ago). - Relative performance of size indices is the ratio of total-return indices rebased to 100 a year ago, so its end value is the relative return in per cent, not the difference of the two returns.
Caveats
- Percentiles describe where a reading sits in one market’s history; they are not forecasts.
- Histories differ in length (P/E from 1999, the margin book from 2017, the monthly labour survey from 2025). Each row says where its history starts.
- Thresholds of 30 and 70 are a convention, fixed in advance, not fitted to returns.
- Nothing here is a recommendation to buy or sell.