Dagster and dbt: do you need both?
Dagster with dbt: what each tool does, how dagster-dbt runs dbt models as assets and dbt tests as asset checks, how partitions work, and when Dagster alone is enough.
Blog
Pipelines, analytics infrastructure, ML systems, and lessons from building them in production.
Start here: your Snowflake warehouse may never be sleeping — the finding, the arithmetic and the SQL to check your own account in ten minutes. Then the 18-check platform teardown, the questions I work through in a $3,000 audit, published in full. No signup for either.
I wrote a book — Local-First Analytics: warehouse-class analytics on your laptop with DuckDB, Parquet and Arrow. The book →
Posts carrying a number I measured myself — a production log, a teardown of something I run, a bill I took apart. Everything else is below.
Dagster with dbt: what each tool does, how dagster-dbt runs dbt models as assets and dbt tests as asset checks, how partitions work, and when Dagster alone is enough.
Dagster vs Airflow for a team of one to five: what each model is good at, what Airflow 3 changed, what each costs, and what still hurt after running a 112-asset Dagster pipeline alone.
An Iceberg catalog keeps one pointer per table, to its current metadata file, and swaps it atomically on commit. What that looks like on disk, the catalog types, which Iceberg catalog a small team should pick, and when you need none.
MCP security best practices for servers that touch data: when MCP authentication applies (HTTP, not stdio), how MCP OAuth works under the 2026-07-28 spec, the token passthrough ban, CIMD versus dynamic registration, and the scoped tokens and approval gate I use.
Postgres and Iceberg meet in four ways: query the lake, write Iceberg tables, act as the catalog, or replicate into Iceberg. What pg_lake, pg_duckdb and the managed options each do, which are dormant, and a hands-on test of Postgres as the catalog.
How to build an MCP server in Python with FastMCP 4: a runnable read-only DuckDB tool with a row cap, a timeout that really stops the query, and a call log, plus test output. Then the production version: how my own MCP server gates every tool.
Apache Iceberg vs Parquet is not a choice between two file formats. Iceberg is a table format whose data files are Parquet. What it adds, what it costs in files and commits, and when plain Parquet is enough, with numbers from a laptop test.
The data quality checks worth writing, as runnable SQL: nulls, duplicates, ranges, relationships, freshness, volume, jumps and gaps. How to automate them in a pipeline, how the data quality tools compare, and which checks caught real bugs in mine.
A self-hosted LLM keeps prompts and documents on your own hardware. What it needs in VRAM, how to serve it with Ollama, llama.cpp or vLLM, whether local RAG works, and the job I moved off my own GPU.
Which DuckDB MCP server to run: MotherDuck's local server, its remote MotherDuck MCP, the duckdb_mcp extension, or your own. I sent the same five probes to two of them, covering writes, file reads, row caps and timeouts. Here is what got through.
Lakehouse vs data warehouse for a small team: what each one is, the data lakehouse architecture in four parts, and the minimum DuckDB lakehouse on a laptop and a bucket, with measured load and query times.
A Postgres MCP server lets Claude Code or Cursor query your database. The official one is archived and had a read-only bypass; here is what to run instead, and the database settings that make it safe.
Look-ahead bias in backtesting, from five real leaks in a production trading pipeline: fundamentals joined 45 days early, a validation split by alphabet, a live bar saved as a close. How to avoid look-ahead bias with point-in-time data and an as-of join, with runnable DuckDB and Polars code.
Set up DuckLake with a Postgres catalog and S3 object storage: the secrets, the ATTACH, what lands in Postgres and what lands in the bucket, how inlining avoids small files, and what two writers did. Every command run on DuckDB 1.5.6.
Adjusted close vs close, with the adjusted close price formula for splits, bonuses and rights issues, a runnable DuckDB and Polars example, what yfinance auto_adjust does, and the 1,499 fake price jumps that mixing the two caused in my pipeline.
A three-state hidden Markov model runs daily inside my trading system. Three features, four indices, a monthly retrain, and one rule that decides whether the regime is allowed to touch a position at all. This is what it does, and what it is not allowed to do.
DuckDB works as an LLM response cache when one process owns it, and the SQL over the cache is the reason to pick it. The three kinds of cache, the single-writer rule, the key that stops you serving a retired model's answers, and measured numbers for exact, brute-force and HNSW lookups.
DuckLake vs Iceberg, measured: the same 50-commit workload on both, with file counts, metadata size and commit times. What DuckLake is, whether it is production ready, and when Iceberg is still the right call.
Survivorship bias in trading, measured on a real NSE pipeline: a universe built from today's stock list missed about 115 delisted names. How to get survivorship-bias-free stock data from the exchange's own files, and how to handle delisted stocks in a backtest.
What drives the build cost of a data pipeline, what it costs to run and maintain afterwards, how managed-ETL pricing models compare with building your own, and the shapes I see by funding stage.
Four LLM agents, 51 Temporal workflows, every risky action behind a human approval gate. An architecture teardown of AEGIS — and what it teaches about production AI automation.
NSE and crypto ingestion, ClickHouse, LightGBM, backtests and a Prolog compliance layer — a self-hosted data platform built and run by one person.
When does running an open-weight model on your own hardware beat paying per token? The answer depends on volume, frequency, and whether you have real privacy constraints.
Batching, checkpointing, and idempotency are what make local inference scale. Get these right and a single GPU chews through millions of rows overnight.
DuckDB vs Spark comes down to one number: the bytes a job touches, set against what one machine can hold. The sizing math, where a Spark cluster earns its cost, and where it doesn't.
Maou, the finance agent in my AEGIS platform, now keeps my books and runs a paper trading desk on real prices. Here is how the desk works, the rules that make a paper result mean something, and what has to be true before it touches real money.
I barely type code any more. My hours go on finding out what a client means, deciding what correct looks like and checking the result against reality. That work is the job now, and it is what you should be paying for.
Snowflake warehouse size sets the credits you burn per hour: X-Small is 1, and each size up doubles it. The full table, what a credit costs by edition, worked monthly figures, and how to pick a size.
dbt-duckdb pairs DuckDB's engine with dbt's models, tests and lineage. The project setup that works, run end to end, plus the limits worth knowing first.
dbt Labs ships an Apache 2.0 MCP server for dbt. Setup takes ten minutes; the real decision is which tool groups you leave switched off.
I ran my personal AI system's thinking on an 8GB GPU at home for months, then moved it to Bedrock. Not because local models are bad — because my GPU is small and the work is bursty. Here are the production numbers, and why the embeddings stayed home.
Define a metric once in YAML and MetricFlow compiles it to SQL. Free to define under dbt Core; the APIs and BI integrations need a paid dbt plan.
Eight open tasks for one failing probe, and thirteen dedupe schemes that disagreed. How I rebuilt AEGIS's alerting so a problem has a record of its own before any ticket exists — and the three things production changed about it afterwards.
My research agent had indexed 20,400 documents. Prompts had used 78 of 10,284 PDFs. Here's what measuring retrieval instead of ingestion changed — and why the fix wasn't a better filter, it was storing less.
My finance agent read 200 characters of each email and called it accounting. Rebuilding it on a plain-text double-entry journal, with Postgres demoted to an index, fixed the numbers, cut the lane's model calls by 71% and caught a bug that kept paid bills open forever.
dbt is the default and the ecosystem. SQLMesh is better engineering for dev environments, incremental state and real SQL parsing. Who should actually switch.
Lazy execution, no hidden copies, and parallelism by default. What each one changes in practice, why the speed claims are softer than the marketing, and when DuckDB is the better answer than either.
Parquet is smaller and faster than CSV for three mechanical reasons, not because it is magic. What the compression ratio depends on, why the speed gap collapses on a warm cache, and when CSV is still the right answer.
DuckDB runs inside one program on batch data. ClickHouse is a server for many clients querying fresh data. How to choose, from running both, and why teams use both.
dbt Core is free and does the transformation. dbt Cloud sells scheduling, CI, an IDE, docs and governance — so compare it with Core plus what you self-host.
A buyer's guide from someone who sells it: the three engagement shapes, what each should cost and deliver, the questions to ask, and when to hire nobody.
They are in the same speed class, so speed should not decide it. The real split is SQL and a database file against Python expressions inside your pipeline.
Keep Postgres as the system of record. Add DuckDB when wide analytical scans slow your app down. A side-by-side table and three ways to run both together.
Postgres is the system of record. ClickHouse is the analytics engine. When Postgres is still enough, what bites when you move, and how to run both together.
Use SQLite for an app's own reads and writes. Use DuckDB for analytics over many rows. A side-by-side table, and how DuckDB queries a SQLite file in two lines.
How to set up DuckDB on a server you control for faster queries — memory and temp-dir settings, Parquet layout, the read-only fan-out pattern, containers, and the mistakes that make a single node look slow.
DuckDB is production-ready for a specific shape of workload and genuinely unsafe for another. The single-writer model, durability semantics, memory behaviour and version compatibility — what actually bites, and how to design around it.
For a small SaaS, a single-node engine over Parquet handles more than founders expect. The three questions that decide whether it fits, what customer-facing analytics needs, and the specific point where you outgrow it.
The full rate card — audit, build, AI workflow, fractional — with the numbers in USD, plus why builds are scoped ranges instead of a day rate.
DuckDB wins when the data a query touches fits on one machine and few people query at once. Snowflake wins on concurrency and governance. Four questions decide it.
The real fully-loaded cost of a full-time data engineer, when a consultant is the better call, and how to tell which one your situation actually needs.
The four ways to wire an LLM to your data — MCP, text-to-SQL, a semantic layer, function calling — and the guardrails that stop the demo becoming an incident.
ClickHouse wins on steady, sub-second dashboards if you run it yourself. Snowflake wins on ad-hoc SQL with nothing to operate. How the two cost models differ.
What it takes to make natural-language querying actually useful for non-technical teams — the schema semantics, guardrails and verification the demos skip.
A teardown of a macro data product: 171 countries scored daily from 10 free public data sources, point-in-time correct, with the licensing and cost work done.
Three months ago I wrote about my personal AI orchestration system and ended with a section titled 'Why It's Not Open Source Yet.' That's fixed. AEGIS is on GitHub, MIT-licensed — here's what it actually took to get there.
The biggest refactor of the AEGIS open-sourcing sprint: removing every branch on an agent's identity so behavior lives in the database as capability tags — resolved at runtime, edited from a UI, and safe even when a tag has no owner.
Scheduling a social post doesn't need a new UI. A to-do already has copy, a time, and labels — so in AEGIS, a Todoist task with a publish label is a scheduled post, and nothing goes out until a card in chat gets a human tap.
For a self-hosted system that's meant to be forked, there's exactly one honest way to handle infrastructure credentials: the user brings their own, the system stores them encrypted, and nothing in the code assumes a vendor. Here's how AEGIS got there — and what it cost me to cut my own vendors out.
The day Next and Someday stopped being Todoist projects and became labels — and why that small modelling decision is what made AEGIS's GTD layer, and its bidirectional sync, actually work.
AEGIS is going open source, and today the first public commits landed. Day one of turning a private system into a shippable one: twenty-seven migrations squashed into a baseline, credentials evicted from the build, and an admin panel redesigned around decisions.
If you're paying thousands a month to query tens of gigabytes, you're renting a freight train to carry a backpack. Here's how to tell, and what to do instead.
AEGIS watches my homelab and GitHub, but the interesting part is what happens between an alert firing and me hearing about it. Most of the time, the correct answer is nothing.
Is DuckDB production-ready? Where it shines, the workloads it fits, and the deployment patterns that work — from someone who ships it.
A short, honest decision framework for the question every data team eventually faces — before you sign a five-figure annual contract for infrastructure your data might not need.
A small SQL-tuned model running on your machine can convert plain English questions into correct SQL. No API keys, no data leaving your laptop, no vector database.
A concrete walkthrough of the local-first pattern: query Parquet directly with DuckDB, no warehouse, no server, no per-query bill — from a notebook or the browser.
Run a model directly in SQL to classify rows, extract text fields, or summarize data without exporting. Here's where it works, what it costs, and when not to bother.
Most analytics stacks are cloud round-trips solving problems that fit on a laptop. Local-first analytics is the case for bringing the compute back to the data — and to the user.
Parse PDFs to markdown, extract structured data with a local model and strict validation schema, no API needed.
Use the Model Context Protocol to let Claude explore your data in plain English. The whole game is permissions: read-only role, standard server, 90% of the value with almost none of the risk.
Most teams reach for Pinecone or Weaviate the moment they hear semantic search. For a team-sized knowledge base, you don't need one. Your database already has what you need.
Claude Code can run dbt in a loop, read errors, and fix them. It needs you to steer grain, business logic, and naming — and to review the tests it writes.
Structural tests catch NOT NULLs and schema errors. Semantic problems—does this description match its category?—need a judge. How to use one responsibly.
An LLM is probabilistic, slow, and costly. Here's how to decide whether it belongs in your pipeline — and where it clearly doesn't.
How to harden a data pipeline by feeding it carefully generated edge cases, using a local LLM when simpler generators aren't enough.
Local-first, file-based data has a real ceiling. Most teams never reach it. When you do, here's what actually changes.
The design decision in AEGIS I'm proudest of isn't an agent or a model — it's a single interactions primitive. How one Postgres table, five kinds of card, and a Temporal workflow replaced every per-domain approval pattern.
AEGIS is my personal operating layer: four agents that watch my inbox, repos, alerts and feeds, and bring a proposed next step instead of a notification.
Claude Code made my AEGIS v3 rewrite fast. The honest part is what got harder: drift, verification and knowing when to stop. The developer stays hands-on.
Prolog and a knowledge graph take routing, classification and fact lookup off the LLM, so most of a personal AI system runs on a cheap local model.
Fit a 2-state hidden Markov model (hmmlearn GaussianHMM) to S&P 500 returns to find market regimes, then compare it with Wasserstein clustering. Python code for both.
What breaks and what it costs — pipelines, warehouse bills, and the failures that only show up in production.
What breaks and what it costs — pipelines, warehouse bills, and the failures that only show up in production. A few a month, never padded to hit a schedule. No sequence, no pitch deck. Reply 'stop' once and you're off — it reaches me, not a queue.