Adjusted close vs close: the formula, and why I stopped using Yahoo's
By Arshad Ansari
Adjusted close vs close looks like a settled question. Close is what traded. Adjusted close is close restated for later splits and dividends. Use adjusted for returns, and you are done.
The trouble is not the definition. It is that an adjusted price is a function of when you computed it, and most pipelines treat it as a fact. Mine spliced Yahoo's adjusted series onto raw exchange prices. That produced 1,499 fake price jumps of more than 50%, volatility features 5–6× too high for months, and training data that changed every time a company announced a split. This post covers the adjusted close formula, a runnable version in DuckDB and Polars, what yfinance actually returns, and what went wrong.
Adjusted close vs close: what each one is
Close is the last traded price of the day, as the exchange printed it. It never changes.
Adjusted close is that price multiplied by an adjustment factor, so it sits on the same basis as today's price. If a stock split 5:1 last month, every close before the split is divided by 5. The chart then shows no cliff on the split date, and a return computed across it is the real return.
That last point is why it exists. A 5:1 split moves the raw price from 1,010 to 204 in a day while nobody's holding changes in value. A volatility or momentum feature computed on raw close sees an 80% crash.
Yahoo says its adjusted close is "the closing price after adjustments for all applicable splits and dividend distributions", using multipliers that adhere to CRSP standards (Yahoo Help). Note the dividends. A dividend-adjusted series is rewritten every time a dividend goes ex, and that matters below.
The adjusted close price formula
adj_close(d) = close(d) × PRODUCT( f ) over every corporate action with ex_date > d
The factors my pipeline uses, from its design record (ADR-0066), each checked against the step the exchange actually printed on real ex-dates:
| action | factor f | checked on |
|---|---|---|
| Split, face value A to B | B / A | 163 ex-dates |
| Bonus, N new per O held | O / (O + N) | 151 ex-dates |
| Rights, N new per O at issue price P, cum price C | (O·C + N·P) / ((O+N)·C), capped at 1.0 | 124 ex-dates |
| Bonus of preference shares, debentures, warrants | 1.0 | 6 rows |
Several actions on the same day multiply. A rights factor is flagged as estimated, because it comes from the subscription price rather than an exact share ratio.
A rights example: 1 new share per 5 held, at ₹80, when the stock closed at ₹120 the day before. f = (5 × 120 + 1 × 80) / (6 × 120) = 680 / 720 ≈ 0.944. Every earlier close is multiplied by 0.944.
I left dividends out on purpose. The model ranks stocks over 5 to 21 days, and the median ratio of Yahoo's adjusted to raw close was 0.977–0.999. At that horizon dividends are noise, and leaving them out means history is not rewritten every week. For a long-horizon total-return backtest, you would add them.
Computing adjusted close in DuckDB
The whole thing is one window function. Join each day's factor onto the price rows, then take a running product over the dates after the current one, by walking the dates newest first.
import duckdb
con = duckdb.connect()
con.sql("""
CREATE TABLE prices AS SELECT * FROM (VALUES
('ACME', DATE '2024-03-04', 1000.0),
('ACME', DATE '2024-03-05', 1010.0),
('ACME', DATE '2024-03-06', 204.0), -- 5:1 split goes ex
('ACME', DATE '2024-03-07', 206.0),
('ACME', DATE '2024-03-08', 103.5), -- 1:1 bonus goes ex
('ACME', DATE '2024-03-11', 104.0)
) t(symbol, date, close);
CREATE TABLE corporate_actions AS SELECT * FROM (VALUES
('ACME', DATE '2024-03-06', 'SPLIT', 2.0 / 10.0), -- face value 10 -> 2: B / A
('ACME', DATE '2024-03-08', 'BONUS', 1.0 / 2.0) -- 1 new per 1 old: O / (O + N)
) t(symbol, ex_date, action, factor);
""")
print(con.sql("""
WITH day_factor AS ( -- same-day actions multiply
SELECT symbol, ex_date AS date, product(factor) AS f
FROM corporate_actions
GROUP BY ALL
)
SELECT p.symbol, p.date, p.close,
-- product of every factor with ex_date > date: walk the dates newest-first
-- and multiply everything strictly after the current row
coalesce(product(d.f) OVER (
PARTITION BY p.symbol ORDER BY p.date DESC
ROWS BETWEEN UNBOUNDED PRECEDING AND 1 PRECEDING), 1.0) AS adj_factor,
round(p.close * adj_factor, 2) AS adj_close
FROM prices p
LEFT JOIN day_factor d USING (symbol, date)
ORDER BY p.date
"""))
Output (DuckDB 1.5):
date close adj_factor adj_close
2024-03-04 1000.0 0.1 100.0
2024-03-05 1010.0 0.1 101.0
2024-03-06 204.0 0.5 102.0
2024-03-07 206.0 0.5 103.0
2024-03-08 103.5 1.0 103.5
2024-03-11 104.0 1.0 104.0
The raw close falls 80% and then 50%. The adjusted close climbs 1% a day, which is what a holder actually saw. Swap FROM prices for read_parquet('prices/*.parquet') and the same query runs over a whole archive.
One caution: this joins on ex_date = date, so an ex-date that is not a trading day in your price table is silently dropped. Count unmatched corporate-action rows as a check.
The same thing in Polars is a reversed cumulative product:
import polars as pl
# prices: symbol, date, close. actions: symbol, date (the ex-date), factor.
day_factor = actions.group_by("symbol", "date").agg(pl.col("factor").product().alias("f"))
out = (
prices.join(day_factor, on=["symbol", "date"], how="left")
.sort("symbol", "date")
.with_columns(
adj_factor=pl.col("f").fill_null(1.0)
.cum_prod(reverse=True) # product from this row to the newest
.shift(-1, fill_value=1.0) # ...excluding this row: ex_date > date
.over("symbol")
)
.with_columns(adj_close=pl.col("close") * pl.col("adj_factor"))
)
Both give the same six rows (Polars 1.42). Keep the factor itself as a column. You will need it again.
What yfinance auto_adjust does
In yfinance 1.7.0, the current release, auto_adjust defaults to True for both yf.download() and Ticker.history(). The source is short: it computes Adj Close / Close for each row, multiplies Open, High and Low by that ratio, replaces Close with Adj Close, and drops the Adj Close column.
So the column called Close in a default yfinance download is not the close. It is Yahoo's adjusted close, split- and dividend-adjusted as of the moment you downloaded it. Pass auto_adjust=False to get both columns.
In my data even the unadjusted Close was split-adjusted as of download. That is the detail that broke things.
What went wrong when I mixed the two
My price table took the raw NSE close where the exchange file had one and Yahoo's close where it did not. adj_close was Yahoo's Adj Close passed straight through, and it was empty for 145,472 rows. Every return, volatility and momentum feature then read "adjusted close, or close if that is missing".
That spliced two bases together. Measured on the live table:
- On pairs of adjacent days that were both exchange rows, raw close had 256 jumps over 50%. The mixed series had 1,499.
vol_fast, a 35-day average of that series, carried 19% of the model's gain. After a phantom 2× step, the blended volatility ran 5–6× too high for months. RELIANCE showed 0.050–0.080 against 0.011–0.014 realised, through July to September 2024. That number sized live positions.- Because Yahoo's adjustment is as-of-download, the features for a past date changed every time a new corporate action landed. The training data could not be reproduced. Train a model on Monday, retrain on the same dates on Friday, and the inputs differ.
Yahoo does not adjust for rights issues
When I measured it, all 120 rights ex-dates with Yahoo data on both sides showed a Yahoo step above 0.99. The exchange printed a median step of 0.936. The original plan was to read the rights factor from Yahoo's implied step. There was no step to read.
So the pipeline uses the formula above, with the premium parsed out of the announcement text. It can compute it for 124 of 126 observable rights ex-dates. Against the raw step, its median error is 0.033, against 0.066 if you ignore rights, and 113 of 124 land within 10%.
ETF unit splits have no corporate-action row
An Indian ETF can re-denominate its units the way a company splits its shares. The corporate-actions feed covers 2,212 symbols and none of them is an ETF. A scan of 451 ETFs found 36 unit splits, 34 of them 1:10 and two 1:100. Unadjusted, one of them booked a −90% day in the backtest: an 18% portfolio hit inside the 2023 figure of +3.6%. The fix is a reviewed register of those splits, read through one module, and a rebuild that fails on any unexplained one-day ratio under 0.5.
The rebuild, and the share-count trap
Now adj_close is computed in-house from the corporate-actions table and never copied from Yahoo. On a dry run over the real archive, 0 adjacent-trading-day moves over 50% remained in adj_close across 3,518,862 rows. RELIANCE is continuous across its 2024-10-28 bonus: 1327.85 to 1334.35, where the raw close halves. A jump guard moves any unexplained move over 50% to a quarantine table. 131 rows sit there, almost all ETF unit splits and bad ticks inside old Yahoo data.
One more trap, because it catches people computing market cap. ADANIPOWER split 5:1 on 2025-09-22. My fundamentals table paired the raw close with the share count from the last annual report. The day of the split, the stock's market cap fell from ₹13.68 lakh crore to ₹3.28 lakh crore and its PE from 105.7 to 25.4. Nothing had happened to the company. 64,287 rows across 321 symbols had the same defect. The raw close was right; the share count was on the old basis. The fix is the share count divided by the adjustment factor, which is why the factor is worth keeping as a column.
So: adjusted close or close?
Both, on purpose. Use an adjusted series for returns and anything built from them. Use the raw close for prices that must match what traded, and adjust share counts, not prices, when you join to fundamentals. Never put a raw price next to an adjusted one.
And compute the adjustment yourself. A vendor's adjusted close is correct as of the day you downloaded it, which is exactly what a backtest cannot use. An adjusted close you computed from a stored corporate-actions table is the same number every time you run it.
The check that caught these jumps, and others like it, is in data quality checks that catch real bugs. Why stored prices must never be quietly revised once a strategy is scored on them is in paper before real money.
The window-function part of this is ordinary DuckDB and Polars over local Parquet. Local-First Analytics covers window functions in both, and handing data between them, though not corporate actions or backtesting. Chapter 1 is free to read; the rest is on Amazon.
More in this series: survivorship bias in backtesting and look-ahead bias and point-in-time data.
Common questions
- What is the difference between adjusted close and close?
- Close is the price the stock actually traded at when the market shut that day. Adjusted close is that price restated so it can be compared with today's prices: every split, bonus issue and (depending on the source) dividend that happened later is applied to it. Close never changes once printed. Adjusted close changes every time a new corporate action lands, because each one restates the whole history before it.
- How is adjusted close calculated?
- Adjusted close is the close multiplied by the product of the adjustment factors of every corporate action with an ex-date after that day. A split from face value A to B has factor B/A. A bonus of N new shares per O held has factor O/(O+N). A rights issue of N new per O held at issue price P, with cum-rights price C, has factor (O·C + N·P) / ((O+N)·C). Several actions on one day multiply. In SQL it is one window function: a running product over the factors, ordered newest date first.
- What does yfinance auto_adjust do?
- In yfinance 1.7.0, auto_adjust defaults to True for both yf.download() and Ticker.history(). It computes the ratio Adj Close / Close for each row, multiplies Open, High and Low by it, replaces Close with Adj Close and drops the separate Adj Close column. So the "Close" you get back is Yahoo's adjusted close, adjusted for splits and dividends, as of the moment you downloaded it. Pass auto_adjust=False to get Close and Adj Close separately.
- Should a backtest use adjusted close or close?
- Returns and anything built from them (volatility, momentum, moving averages) need an adjusted series, or a 2:1 split reads as a 50% crash. Order prices, position sizing in shares, and anything joined to a share count need the raw close, or the numbers stop matching what traded. The trap is mixing them: a raw price next to an adjusted one, or a raw price divided by a share count from before a split. Compute your own factor table so the adjustment does not change under you between runs.
- How do rights issues affect adjusted prices?
- A rights issue lets existing holders buy new shares below the market price, so the price drops on the ex-date without anyone losing value. The adjustment factor is the theoretical ex-rights price divided by the cum-rights price: (O·C + N·P) / ((O+N)·C) for N new shares per O held at issue price P. In my data, Yahoo did not adjust for rights at all: all 120 rights ex-dates with Yahoo data on both sides showed a Yahoo step above 0.99, while the exchange printed a median step of 0.936.
Get new posts by email
Data engineering notes like this one — what breaks and what it costs, in production.
What breaks and what it costs — pipelines, warehouse bills, and the failures that only show up in production. A few a month, never padded to hit a schedule. No sequence, no pitch deck. Reply 'stop' once and you're off — it reaches me, not a queue.
Want the whole playbook?
If this was useful, the long version is my book. Local-First Analytics — 313 pages, runnable code for every chapter — is the full build: DuckDB, Parquet and Arrow, from install to production. On Amazon, and chapter 1 is free to read here.
Get the bookNot ready to buy? Read chapter 1 free — the whole chapter, no email required.
Rather talk it through? Book a free 30-minute call. No slot that suits your time zone? Email info@hikmahtech.in.