Survivorship bias in trading: how my backtest only saw the winners
By Arshad Ansari
Survivorship bias in trading sounds like a textbook warning. You test a strategy on today's stocks, the dead ones are missing, and the result looks better than it should. Everyone has heard it. Most people assume their data does not have it, because they never chose to drop anything.
Mine had it, and I had not chosen it either. My trading pipeline built its stock universe from a current list of listed companies. On one date I checked, it held 1,635 names where the exchange had traded 1,763. It was 7% short, and short of exactly the companies that later died. This post covers how that happened, how I rebuilt the universe from the exchange's own files, and what a fix can and cannot repair after the fact.
What survivorship bias in trading looks like
A backtest asks: if I had run this rule every day since 2023, what would I have held, and what would it have returned? To answer it, the backtest needs the list of stocks that were tradable on each of those days.
The easy source for that list is a stock screener or a data vendor's "all listed equities" endpoint. Both answer for today. A company that delisted in 2024 is not on today's list, so it is not in your 2023 universe either. Your backtest can never buy it, which means it can never lose money on it.
The effect is one-directional. Companies leave a market for many reasons, but the ones that leave because they failed fall a long way first. Dropping them removes losses and leaves gains in place. Returns go up, drawdowns go down, and nothing in the output tells you.
How it got into my pipeline
My pipeline (described in building a data platform solo; it has grown since) prices Indian equities from the NSE bhavcopy. That is the exchange's daily file of every security that traded. The equity master, the table that says which symbols are companies, came from screener.in's current list.
An earlier change made "has a row in the equity master" the rule for entering the universe. The goal was to keep ETFs out, because NSE lists ETFs in the same EQ series as shares. It did keep them out. It also took the dead with them.
The design record (ADR-0066 in the repo) measured the damage on the 2023-01-02 to 2026-09-04 archive:
- About 115 delisted or merged equities were missing. HDFC, which traded ₹1,105 crore a day until 2023-07-12, then PEL, IDFC, ISEC, JPASSOCIAT, TATAMTRDVR and the rest.
- The pre-rename half of 75 renamed companies was missing. Their history started on the day of the new ticker.
- On 2023-03-28 the universe held 1,635 names where the exchange traded 1,763 non-fund symbols.
The ADR says it plainly: "A model trained on that has seen only survivors."
The fix: rebuild the universe from the exchange's files
The rule now is: every EQ-series bhavcopy symbol, minus ETFs. The exchange file says what traded. It does not know which companies will die, so it cannot leave them out.
The hard part is the "minus ETFs" bit. You cannot infer a fund from a price series, so the pipeline names ETFs explicitly, from two lists:
- NSE's own ETF list,
eq_etfseclist.csv, saved every day at 02:20 IST. The universe excludes the union of every snapshot ever taken, so a fund that delists tomorrow stays out for good. The asset refuses to publish a list under 200 rows, and the price build refuses to run without one. A missing list would let every ETF back in, silently. delisted_etfs.csv: 101 funds that died before the first snapshot. 93 are proved by following ticker changes to a symbol NSE's current list still carries (DSPITETF → ITETFADD → ITADD). The other 8 are bond and index funds that died without a successor, each listed with its evidence.
The equity master is still joined for ISIN and sector, but a missing master row no longer drops the price row. An empty ISIN is the expected state for a delisted name.
Ticker renames are survivorship bias too
A rename looks like a death followed by a birth. The old ticker stops, a new one starts, and a universe keyed on ticker sees two short histories instead of one long one.
Nothing in the bhavcopy says "this is the same company", but the prices do. The old symbol's last close reappears, to the paisa, as the new symbol's first prev_close, within a session or two. That rule found 75 renames. A later change added a second rule: bhavcopy rows now carry the ISIN, and one ISIN under two tickers whose dates do not overlap is the exchange itself saying they are one security. The ISIN rule found 23 renames the carry-over rule could not see. The committed map, symbol_renames.csv, has 114 rows today.
The carry-over rule has a weak spot worth knowing. A round-number close can carry into two new symbols at once. AMBANIORG's last close of exactly 100.00 matched the first prev_close of two unrelated tickers, and the build script aborted until the ISIN rule broke the tie.
With the map in place, the pipeline republishes the old ticker's rows under the current one, chains included (CLNINDIA → HEUBACHIND → SUDARCOLOR). TMPV now has rows from 2019 instead of from 2025-10-24.
Delisted stocks in historical data can also be invented
Survivorship bias is the famous direction. The rebuild found the opposite error too.
27 symbols never appear in the bhavcopy at all, yet had about 900 continuous Yahoo "trading days" each from 2023 to 2026: RCOM, HDIL, EDUCOMP, SKIL, the IL&FS entities and others. These are NSE-suspended companies. They do not trade. The free price feed kept returning a price for them anyway, and the model had trained on about 24,000 rows of that series.
So the rule for delisted stocks in historical data cuts both ways. Keep the real history of names that died. Do not keep a vendor's continuation of a name that stopped trading. The exchange file answers both questions, because a day a stock did not trade has no row in it.
How to handle delisted stocks in a backtest
The checklist I would use now, with the parts I have built marked as such:
- Universe per date from the exchange's own files. Built. A symbol enters the universe on the dates it traded, whatever happened to it later.
- Stitch renames before you compute anything. Built. Otherwise a rename looks like a stock that vanished, and any per-symbol feature (a 200-day average, a 52-week high) restarts from nothing.
- Never fill after the last trading day. Built, from 2023 where my exchange archive starts. No forward-fill, no vendor rows on days the exchange has no row.
- Decide what a held position is worth when the stock stops. A merger pays the deal terms. A voluntary delisting usually has an exit offer. A bankruptcy pays close to nothing. CRSP records this as a delisting return for US stocks; with exchange files alone you have to set the rule yourself, and the honest default is the last tradable price, not the last price a vendor printed.
- Count one past date by hand. Compare your universe size with the number of symbols the exchange traded that day. This single count is what exposed my 7% gap.
Survivorship-bias-free stock data: where to get it
You need a source that keeps securities after they die.
- US equities: CRSP's US stock databases cover "over 32,000 active and inactive U.S. securities" and include delisting information. It is the academic reference and is usually reached through a university or a research subscription.
- Indian equities: the NSE's daily bhavcopy is survivorship-free by construction, because it lists what traded that day. What you have to add yourself is the archive, the ETF exclusion and the rename map.
- A free API keyed on today's tickers: fine for prices of names you already know, not for building a universe. Even a long history cannot include a ticker the API no longer recognises.
Not every attribute can be fixed afterwards
Prices can be rebuilt from an archive. Some attributes cannot, because nobody archived them.
My universe also carries a halal screen, a vendor verdict that the product filters on. That screen came from the current master on every date, which was survivorship bias and look-ahead at once. Measured on production on 2026-09-06: 20,840 forecast rows belonged to delisted names, and 0 of them were marked halal. A company with no row in today's master could never pass the screen. The halal backtest held 0 delisted names; the unrestricted one held 13, for 298 position-days.
There is no archive of that screen to rebuild from. So the pipeline now snapshots it daily (ADR-0067), and the halal backtest is point-in-time only from the first snapshot, 2026-09-05. Before that date every backtest run says, in its own metadata, that it used a current-status universe. One year of recording buys one year of point-in-time history. That is the cost of not starting in 2023.
So: is your backtest survivorship-free?
Ask where your stock list comes from. If the answer is a list of what is listed today, it is not, however clean the prices are. The fix is not a statistical correction. It is a data source that records what traded on each day, kept long enough to rebuild from.
Survivorship bias is one of the ways a backtest can see the future. The next one, information that was not public yet on the date you used it, is subtler; the checks that catch both kinds in a running pipeline are in data quality checks that catch real bugs. And before any of it touches money, a strategy should earn it on paper first: paper before real money covers the rules for that.
What made this fix possible was an archive of the exchange's raw daily files, never overwritten. Local-First Analytics has a chapter on exactly that layout: an append-only raw zone kept as delivered, with Parquet layers built from it that you can regenerate at any time. It does not cover backtesting. Chapter 1 is free to read; the rest is on Amazon.
More in this series: adjusted close vs close and look-ahead bias and point-in-time data.
Common questions
- What is survivorship bias in trading?
- Survivorship bias in trading is testing a strategy only on the stocks that still exist today. Companies that went bankrupt, were taken over or were delisted drop out of the data, so the backtest never holds a stock on its way to zero. Returns look higher and drawdowns look smaller than anything you could have traded. It usually enters through the stock list: if your universe is a current list of listed companies, every name on it is a survivor by construction.
- Where can I get survivorship-bias-free stock data?
- From a source that keeps dead securities. For US stocks the reference is CRSP, which covers over 32,000 active and inactive securities and includes delisting information. For Indian stocks, the exchange's own daily price files (the NSE bhavcopy) list every symbol that traded that day, including those that later delisted, so a universe rebuilt from the archive of those files is survivorship-free by construction. A free API that only answers for tickers listed today cannot give you this, however long its history is.
- How do you handle delisted stocks in a backtest?
- Keep their price history up to the last day they traded, let the strategy select them on every date they were listed, and decide in advance what a held position is worth when the stock stops trading: the takeover price for a merger, the last tradable price for a voluntary delisting, close to zero for a bankruptcy. Also stitch ticker renames, or a renamed company looks like one stock that died and a new one that appeared. Never fill the days after delisting with a price; a data vendor may still return one.
- How do you avoid survivorship bias in a backtest?
- Build the universe for each date from what was listed on that date, not from today's list. Use the exchange's own historical files or a vendor that keeps inactive securities. Check one past date by hand: count the symbols your universe holds and compare with what the exchange traded that day. In my pipeline that one count showed 1,635 names against 1,763, which was 7% short and short of exactly the names that later died.
- What is the difference between survivorship bias and selection bias?
- Survivorship bias is one kind of selection bias. Selection bias is any error that comes from studying a sample that is not representative of what you want to measure. Survivorship bias is the case where the filter is survival: you only see the things that lasted. In a backtest, picking stocks by today's index membership, today's market cap or today's screen result are all selection biases, and the survival filter is the most common of them because most stock lists are lists of what exists now.
Get new posts by email
Data engineering notes like this one — what breaks and what it costs, in production.
What breaks and what it costs — pipelines, warehouse bills, and the failures that only show up in production. A few a month, never padded to hit a schedule. No sequence, no pitch deck. Reply 'stop' once and you're off — it reaches me, not a queue.
Want the whole playbook?
If this was useful, the long version is my book. Local-First Analytics — 313 pages, runnable code for every chapter — is the full build: DuckDB, Parquet and Arrow, from install to production. On Amazon, and chapter 1 is free to read here.
Get the bookNot ready to buy? Read chapter 1 free — the whole chapter, no email required.
Rather talk it through? Book a free 30-minute call. No slot that suits your time zone? Email info@hikmahtech.in.