Mutual fund people, distributor payments and state comparisons: methods
These additions answer three questions on the existing fund pages: how many people invest, who receives distribution payments, and how large state fund assets are relative to population and economic output.
Sources and editions
- AMFI–Crisil Factbook 2025, printed p. 29: six March-end unique-investor observations, 2020–2025, in crore rounded to one decimal. These chart labels were transcribed and visually checked. The source does not spell out its deduplication method or split PAN and PEKRN.
- AMFI–Crisil Factbook 2024, printed p. 13: the five overlapping March 2020–2024 values match and the chart explicitly labels them Unique PAN & PEKRN. This captured definition reference was visually checked. PEKRN means PAN-exempt KYC reference number; the combined series must not be described as PAN alone. The 2025 chart does not repeat this definition, so no unprinted breakdown is assigned to its latest count.
- AMFI distributor commission disclosures, FY2023–24 and FY2024–25: 2,499 and 3,158 numbered distributor rows respectively. The complete tables are parsed from the original PDFs’ layout-preserving text.
- AMFI–Crisil Investor Trends, July 2026, printed pp. 16–17: the reported top and bottom ten for AUM per person and AUM/GSDP. Forty chart labels were transcribed and visually checked; the middle states are absent. These are AMFI’s ratios at printed precision, not new calculations from reconstructed denominators.
Original PDFs, extracted commission text, reviewed CSVs and SHA-256 checksums
are retained in docs/reference/amfi/. The mf_context step verifies their
integrity on every run. The build makes no network requests. Every bundle names
its source edition, source URL and checksum, and every panel labels its period.
The AMFI references are captured editions. The separate SEBI investor series below
is refreshed from Data Bank’s recurring monthly bulletin dataset.
People and accounts
Unique investors are compared with the AMFI Monthly Report’s conventional-MF grand-total folios on the same exact March-end dates. A missing monthly observation remains missing, with no nearby-month substitution. Headline folios exclude domestic fund-of-funds. The accounts-per-reported-investor ratio is an approximate comparison of industry aggregates, not the typical investor’s account count. Rounded unique counts limit its precision. No monthly unique investor history or new-investor count is inferred.
Monthly unique PAN plus PEKRN
SEBI Monthly Bulletin Table 1, “Indian Economy and Capital Markets at a Glance”, reports the unique number of investors in Mutual Funds (by PAN+PEKRN). The initial validated history covers April 2025–August 2026 (17 months), from the October 2025–September 2026 workbooks. This is industry-wide, not the CAMS-serviced universe. SEBI does not supply separate PAN-only and PEKRN-only components here. Monthly bulletin archive.
Data Bank owns the workbook downloads and extraction. A long section heading
formerly caused the generic parser to treat the following rows as notes. The
corrected footer rule retains section headings when numeric data follows.
The curated mutual_fund_investors / INDUSTRY / unique_pan_pekrn / investors
series accepts positive integer counts from the explicitly labelled row,
excludes fiscal-year summary columns and blank future months, and retains the
newest reporting issue per data month. Its existing first value, first issue,
vintage count, revision flag and source SHA are preserved in Tipsheet’s download.
Data months must precede their reporting issue; no guessed values are added.
Tipsheet dates the source month at month-end and joins headline folios only on that exact date. Counts remain exact integers; chart display in crore is rounded. Changes are calculated only between adjacent calendar months, with gaps left null. Negative changes are retained. These are net changes in the outstanding count, not gross first-time investors. The initial series includes a decline of 149,218 in May 2026, exactly as published; no explanation is inferred.
Annual Factbook observations remain a separate reference at their printed precision. Neither their rounded values nor quarterly CAMS counts are spliced into this monthly industry series. The site exposes a monthly chart, table, chart CSV and JSON table with source labels and revision metadata.
Distributor economics
Source money columns are lakh rupees, divided by 100 for crore. Commission and expense payments include GST and other taxes, cesses, charges and levies where applicable. Payments/AUM divides each reported payment by that distributor’s annual average AUM. It is not a contractual commission rate, a fund’s TER or an investor’s return drag. Net/gross inflows describes money flows, not retention, investor switching or measured churn.
A dash is missing, not zero. FY25 has two missing payout amounts. Payout shares use the sum of known payouts and are explicitly shares of reported payments within the disclosed universe. Neither edition represents every distributor or all industry commission. The universe changes between years, so totals are shown with their row counts rather than called same-distributor growth. Names are retained as printed; serial numbers are publication row numbers, not ARN identifiers. No cross-year identity match is inferred from names.
The first 15 recipients appear in a chart; the first 25 in a sortable table. The downloadable JSON contains all 3,158 FY25 rows and the source values. The source’s separate AUM/gross-inflow ratio does not consistently reproduce its money columns, so it is preserved in the download but not used in analysis.
The parser requires the source’s FY and lakh-unit headings, parses all numbered rows, and requires a contiguous, unique serial sequence matching the captured edition’s expected count. A truncated or malformed table fails the step. Net outflows are valid; negative payments, subscriptions or average AUM fail.
State penetration
These reported ratios use 31 July 2026 month-end AUM. The longer state history elsewhere on the site uses monthly average AUM. The two are not mixed. The snapshot retains the report’s administrative boundaries, with Delhi renamed from New Delhi and the combined Dadra/Nagar Haveli/Daman/Diu region labelled explicitly. It does not use the long history’s combined AP/Telangana or J&K/Ladakh geography.
Population is projected at 31 March 2025 from UIDAI’s Annual Report 2024–25, except Chandigarh, which uses December 2021. The report also specifies an older population estimate for Puducherry, which is absent from these top/bottom-ten charts. GSDP is FY2025 except Goa, Chandigarh, Gujarat, Sikkim, Nagaland, Manipur and Mizoram, which use FY2024. Gujarat’s GSDP comes from its Socio-economic Review 2024–25; others are attributed to MoSPI. Ladakh has no GSDP comparison.
These dates are included with every downloaded observation and the on-page tables. AMFI says its city/state master changed from March 2026. The snapshot is therefore kept separate from historical penetration comparisons; no smooth trend across that revision is inferred.
Corporate and institutional assets are included, so AUM per person does not describe the average resident’s investment. AUM/GSDP compares a stock of assets with annual economic output; corporate headquarters can lift a state’s ratio. Per-person or GDP-normalised figures alone do not establish household uptake.
Checks enforce distinct state/metric/date keys, positive finite values, both metrics, both top/bottom groups and ten rows per group. Absent states never receive zeros or inferred ratios.
Updating the references
Capture the next source PDF from AMFI’s page and record its URL, retrieval date
and SHA-256. For commissions, run pdftotext -layout and record the text checksum,
financial year and the final printed serial number. The parser supports variable
decimal precision, Indian comma grouping, negative flows and missing amounts.
For chart-only disclosures, add reviewed CSV observations with source file,
printed page, numerator date and denominator period; visually check every label
before updating the CSV checksum. A different state-report layout or coverage
must be modelled explicitly, not forced into the top/bottom-ten contract.
Run uv run python -m tipsheet.run --only mf_context and the focused tests,
then rebuild and check the site. New editions remain dated; never relabel an
old observation as current. Full state ratios recomputed from population and
GSDP inputs would be a separate dataset and must reconcile numerator definitions
before replacing this reported snapshot.