Dagster vs Airflow: which to choose for a small data team
By Arshad Ansari
"Dagster vs Airflow" sounds like a feature contest. It mostly isn't. Both schedule Python, both retry failures, both have a UI, both are free to self-host. The real difference is what each one thinks it is scheduling. Airflow schedules tasks. Dagster schedules the things those tasks produce.
For a big team with an Airflow platform already running, that difference rarely justifies a migration. For a small data team starting from nothing, it decides how much you can hold in your head. I run a production Dagster pipeline alone: 112 assets, 93 asset checks, 57 schedules. This post is about why that model fit, what still hurt, and when Airflow is the better answer. I have not run Airflow in production, so the Airflow side comes from its documentation, not from my scars.
Tasks or assets: the one difference that matters
In Airflow you write a DAG: task A, then task B, then task C. Airflow knows whether each task succeeded. It does not, by default, know what the task wrote.
In Dagster you write an asset: a function that produces a table, a file or a model, and names the assets it depends on. The graph comes from those dependencies. Dagster records each materialization, so it knows which partitions of which table exist, when they were built, and whether the checks on them passed.
from dagster import asset, AssetIn
@asset
def daily_prices(): ...
@asset(ins={"prices": AssetIn("daily_prices")})
def sector_statistics(prices): ...
That is the whole pitch. If your job is "keep these 100 tables correct and current", the asset model matches the job. If your job is "call these 40 systems in this order", Airflow's task model matches it at least as well.
Why the asset model fit a 112-asset pipeline run by one person
My pipeline ingests market data, builds features, trains models and writes to ClickHouse. The counts today: 112 assets, 36 of them partitioned (34 daily, 2 static), 100 jobs, 57 schedules, 3 sensors, 93 asset checks. It has 913 commits since November 2025 and about 3,000 tests. The teardown of the whole platform gives the architecture; it has grown since that was written.
Three things made the asset model worth it for one person:
- Partitions are data, not a loop. Dagster tracks the status of every daily partition. When a scrape fails for one date, I see one red square and rerun one square. I don't write backfill scripts that remember which dates they have done.
- Checks sit next to the thing they check. 93 asset checks run against the assets they guard and show up on the asset's page. When something is wrong, the check and the data are in one place.
- Granularity is a decision, not a default. I could choose to drop partitions where they cost more than they saved. One sector-statistics asset rebuilt 1,882 dates, 109,152 rows, in about 45 seconds as a single unpartitioned asset. The per-partition version would have meant over 2,000 partitions and 5,000 to 20,000 queries. The first phase of that change materialised 9 of 10 such "full" assets in about six minutes.
None of this is impossible in Airflow. Airflow 3 now has assets too (more below). But in Dagster the asset is the centre of the tool, and with one maintainer, the centre of the tool is what you get for free.
What still hurt
The asset model did not remove the hard parts. It moved them. Three that cost me real time:
Code said "stopped"; the database said "running". I set default_status=STOPPED on sensors in code. Four kept running. Dagster keeps a sensor's on/off state in its own database, and that stored state wins over the default in code. One of those sensors rewrote all the NSE price data twice a night. The only off switch I trusted in the end was deleting the sensor.
A bad partition key ran everything. I launched a run with a partition key that was not in the partition set. It did not error. Dagster ran all partitions. The symptom surfaced much later as a confusing "multiple partitions" error from the pickled-object IO manager, nowhere near the cause. Validate keys yourself before you launch.
Clocks are not dependencies. An index-consolidation asset was scheduled at 02:03 and 02:15. It failed five or six times a night for a week, on correct data, because the calendar data it checks against only extends at 02:30. The fix was to stop scheduling by clock and declare the real dependency with AssetDep, so the asset runs when its input is ready. This is the asset model's argument made by a bug: if you schedule outputs by clock, you have rebuilt Airflow with extra steps.
Each of these is an operations lesson more than a product flaw. But they are the kind of thing a comparison table does not show you.
Airflow 3 closed some of the gap
Airflow 3.0 shipped on 22 April 2025. From its release notes, the changes that matter for this comparison:
- DAG versioning. The grid and graph views know which version of a DAG ran, and versioning applies on clear, rerun and backfill.
- Assets. What Airflow 2 called Datasets are now Assets. An
@assetdecorator creates an asset, a DAG and a task in one declaration, and asset updates can trigger downstream DAGs. - Scheduler-managed backfills, instead of a separate command-line process.
- A Task SDK and Task Execution API, which separates task code from Airflow's core.
- A rewritten React UI, and SubDAGs removed.
So "Airflow has no idea what data exists" is now out of date. What hasn't changed is the centre of gravity. In Airflow 3 an asset is something a DAG produces and something that can trigger another DAG. In Dagster the asset is the unit, with partitions, checks and materialization history attached. If you are on Airflow 2 and the asset idea appeals to you, upgrade to 3 before you consider switching tools.
Where Airflow is the better choice
Be fair to Airflow. It has real advantages for a small team:
- The ecosystem. Airflow has a provider package for almost every system you will touch. If your pipeline is mostly "move data between SaaS tools and a warehouse", that matters more than an asset graph.
- The hiring pool. Far more data engineers have used Airflow than Dagster. If you plan to hire your second and third engineer, they will probably arrive knowing Airflow.
- Managed hosting from the clouds. Amazon MWAA, Google Cloud Composer and Astronomer all run it for you. If your company already buys AWS or GCP, managed Airflow may be one procurement form away. Dagster's managed product comes from one vendor.
- An existing setup. If you already have Airflow working, the migration cost almost always beats the benefit. Rewriting 50 working DAGs gives you the same tables you have today.
Dagster vs Airflow cost
Both cores are open source under Apache 2.0 and free to self-host. The real cost of self-hosting either is a Postgres database, a server, and your time keeping them healthy. I covered what that time costs in what a data pipeline actually costs.
The managed prices differ in shape:
- Dagster+: when I checked dagster.io/pricing on 1 October 2026, the Solo plan was $10 a month (one user, one code location) and Starter was $100 a month (up to three users, five code locations). Both add pay-as-you-go credits at $0.040 and $0.035 per credit respectively, plus $0.010 a minute for serverless compute. A credit is one asset materialization or one op execution. Pro is "contact sales".
- Managed Airflow: priced per environment or per worker by each vendor. I have not verified current numbers, so check the MWAA, Composer and Astronomer pricing pages for your region.
The credit model is worth a minute with a spreadsheet. My pipeline materializes 112 assets, 36 of them daily-partitioned, and the partitioned ones are where the count grows. Count your expected materializations per month before you assume the small plan fits.
Dagster vs Airflow for dbt
If dbt is the centre of your stack, this question changes. Dagster's dagster-dbt reads the dbt manifest and turns every model into an asset and every dbt test into an asset check. Your Python assets and your dbt models end up in one lineage graph. On Airflow, the common route is Astronomer Cosmos, an Apache-2.0 library that renders a dbt project as Airflow DAGs and task groups.
Both work. If you are still choosing the SQL layer, SQLMesh vs dbt covers that decision.
A short decision checklist
Answer these in order. The first "yes" usually decides it.
- Do you already run Airflow, and does it work? Stay. Upgrade to Airflow 3 if you want assets.
- Will you hire several engineers in the next year? Lean Airflow, for the hiring pool.
- Is your pipeline mostly calls to other systems, not tables you own? Airflow fits at least as well.
- Are you one to five people building tables, features or models from scratch? Lean Dagster.
- Do you need daily partitions with per-date status and reruns? Lean Dagster.
- Must the managed service come from your existing cloud vendor? Airflow.
So: Dagster or Airflow?
For a small team building a data platform from nothing, I would pick Dagster, and I did. The asset model let one person keep 112 assets, 36 partitioned, in my head. It did not make operations free: stored sensor state, silent partition behaviour and clock-based schedules each cost me days.
For a team with a working Airflow setup, a hiring plan, or a cloud contract that includes managed Airflow, Airflow 3 is the sensible choice, and it is closer to Dagster than it was.
Local-First Analytics covers the storage side of a small-team platform: DuckDB, Parquet and Arrow on hardware you already have. Chapter 1 is free to read, the rest is on Amazon. Related reading: building a data platform solo, and what a data pipeline actually costs.
If you are choosing an orchestrator and want a second opinion on your workload, a Data Platform Audit is a week and a written roadmap you keep. If you already have a working Airflow setup and no specific pain, you probably don't need me, and I'd rather say so here.
Common questions
- When would you choose Dagster over Airflow?
- Choose Dagster when your pipeline is mostly about producing tables and files that depend on each other, and you want the orchestrator to know about those outputs: which partitions exist, which are stale, which checks passed. That fits a small team building a warehouse or a feature store from scratch. Choose Airflow when your work is mostly running tasks against other systems, when you already have Airflow and people who know it, or when you want the widest choice of managed hosting and the largest hiring pool.
- How does Dagster compare to Airflow?
- Airflow's unit is the task inside a DAG: it schedules work. Dagster's unit is the asset: a table, file or model that a function produces. Both are open source under Apache 2.0, both run Python, both have managed offerings. Airflow has the bigger ecosystem, more providers and more people who already know it. Dagster gives you partitions, lineage and data-quality checks as first-class objects, which suits teams whose job is the data rather than the jobs. Airflow 3 (April 2025) moved towards assets, but its core is still the DAG of tasks.
- Dagster vs Airflow cost: which is cheaper?
- Both are free to self-host; the real cost is the server and your time. Managed Airflow comes from several vendors (Amazon MWAA, Google Cloud Composer, Astronomer), each priced differently, so compare their current pricing pages. Dagster's managed product, Dagster+, listed a Solo plan at $10 a month and a Starter plan at $100 a month plus per-credit charges when I checked its pricing page on 1 October 2026, where a credit is one asset materialization or one op execution. Check the page before you budget; prices change.
- Dagster vs Airflow for dbt: which is better?
- Dagster's dagster-dbt library reads the dbt manifest and turns each dbt model into a Dagster asset and each dbt test into an asset check, so dbt lineage and Python lineage live in one graph. On Airflow the common route is Astronomer Cosmos, an open-source library that renders a dbt project as Airflow DAGs and task groups. Both work. Dagster's model is closer to how dbt already thinks about models; Cosmos is the better choice if the rest of your platform is already on Airflow.
- What changed in Airflow 3?
- Airflow 3.0 shipped on 22 April 2025. The main changes: DAG versioning, so the UI and reruns know which version of a DAG ran; Datasets renamed to Assets, with an @asset decorator that creates an asset, a DAG and a task in one go; backfills managed by the scheduler; a new Task SDK and Task Execution API that separates task code from the core; a rewritten React UI; and SubDAGs removed. It narrows the gap with Dagster, but Airflow is still built around the DAG of tasks.
- Is Dagster better than Airflow?
- Not in general. Dagster is better when the outputs are the point and the team is small enough to adopt a newer model without retraining a department. Airflow is better when you need its integrations, its hiring pool, or a managed service your cloud already sells you. For a small team starting a data platform from nothing, I would pick Dagster; for a team with a working Airflow setup, I would rarely migrate.
Get new posts by email
Data engineering notes like this one — what breaks and what it costs, in production.
What breaks and what it costs — pipelines, warehouse bills, and the failures that only show up in production. A few a month, never padded to hit a schedule. No sequence, no pitch deck. Reply 'stop' once and you're off — it reaches me, not a queue.
Want the whole playbook?
If this was useful, the long version is my book. Local-First Analytics — 314 pages, runnable code for every chapter — is the full build: DuckDB, Parquet and Arrow, from install to production. On Amazon, and chapter 1 is free to read here.
Get the bookNot ready to buy? Read chapter 1 free — the whole chapter, no email required.
Rather talk it through? Book a free 30-minute call. No slot that suits your time zone? Email info@hikmahtech.in.