Why I'm so obsessed with DuckDB

25 Sep 2026 | Karik Isichei-Forty

There’s a vast gap between a company that wants to simply formalise their analytical tech stack and what’s on offer as the de facto standard of data engineering products today. If you want to buy-in a solution, there will probably be a grand total of two SaaS products that come to mind (Databricks and Snowflake). However, both of these options are tailored to enterprise level companies who have deep enough pockets to manage it and, if they’re lucky, will also have enough data to warrant it. The benefits of these offerings do not trickle down to smaller companies who just want to manage their data a bit better; but who are left picking up a hefty bill for services they don't need.

Data engineering in its evolution has drifted to distributed services and tooling. Over the past decade of hardening distributed ETL best practices, the price of a single machine with top tier compute and storage plummeted. This change has not gone unnoticed and there is a slow resurgence of single node compute. This is the opener to a series I am going to write on using DuckDB for data engineering. I hope this will be a useful reference to those who just want a simple data stack (cloud or on-prem) without having to pull in a lot of the existing off the shelf cloud data engineering products. So that they can simply deal with their small to medium sized data [1]Big Data - When I refer to big data I am thinking of table sizes on the scale of Terrabytes to Petabytes. The definition of "big data" will continue to shift to the right as technology improves but for this article, at this time, that is how I am defining "big"..

A Little History

Specifically, why I think data engineering has gone this way.

Before I get into the series, I want to provide some historical context, or at least my viewpoint from being in the thick of it of building out data and analytics platforms, to help explain where I’m coming from as I write this.

Maxime Beauchemin wrote a fantastic and highly influential set of pieces about data engineering from 2017 to 2018. In The rise of the Data Engineer he describes the explosion in this new data engineer role driven by a need for reliable data at scale. He later wrote an article called Functional Data Engineering which outlined what you probably still see today in your data engineering stack. In short:

  • Blob storage. Cloud storage of data is cheap and “infinite”. It also came with the benefit of enabling us to dump whatever format of data from wherever to a single place of access. For the first time it was easy for data scientists to pull everything together to drive new insights.
  • Horizontally scaling and on-demand compute. When data would be too big to query on a single Postgres instance we could just farm it off to a Spark Cluster and up the number of nodes to get the job done. For light pieces of work you could still use your python scripts to parse and transform those CSVs sent over via email.
  • Orchestration layer. Something to manage what data to transform and move with whichever flavour of compute you wanted. Probably something like an Airflow instance running on Kubernetes.
  • Metadata management. Once our unstructured data lakes started to grow we needed something to link it back together in a way that made it easier to discover new datasets.

All these services are independent and as lightly coupled as you want them to be, and come with great flexibility. But, as this modern data stack grew in complexity and maturity so did the same age old demands on the traditional OLTP systems [2]OLTP: OnLine Transactional Processing - Data processing tailored to operating a row based level. I want to update this single row / entity in my database (think Postgres, SQLite). we left behind:

  • Access control. Issues started to occur from a free for all data access policy. So our stack needed to start implementing authentication and role based access controls often being implemented at multiple surfaces of our distributed data stack.

  • Audibility. There is always a cost with distributed services. Often you need to buy in more expensive services to log and tie them back together again to monitor. In most cases for flexibility of orchestrating whatever compute we want we would be left with managing a Kubernetes cluster for our Airflow instance to run on.

  • ACID. Our stacks were often heavily focused on batch processing large amounts of data. Writing highly compressed chunks of data to blob storage is slow. Spinning up on demand distributed compute is slow. Waiting for your Airflow scheduler trigger a job is slow [3]Airflow have done a lot to optimise how it schedules DAGs but there was a time where outages would occur just through your Airflow scheduler dying from a slow python DAG script.. At the start, these were the acceptable tradeoffs for relatively cheap and scaling compute. But as the outputs from our data stacks started to become integral to business decisions, there was a need to have faster and accurate updates to our data. So we created new formats like Delta and Iceberg to make Parquet ACID compliant, versioning and updating data faster.

To account for these features our data stack had to pivot away from the functional and distributed ethos we started with to something more centrally controlled. This just added more complexity and more headcount to manage that complexity. I believe this is one of the biggest selling points of Databricks and Snowflake. Just bundle that complexity elsewhere as someone else's problem. This is a good reason to buy SaaS products and why they exist. However, I don’t think they necessarily fix the underlying issue we have built up over the last decade of the modern data stack. Which is why they are still complex and expensive to run on most organisation's data.

I don’t believe the majority of organisations require the modern data stack of today. What they need is a Postgres-like database that prioritises analytical workloads. Thankfully we now have those tools now:

  • We have fairly cheap single node compute and cheap fast SSD storage.
  • We have DuckDB as an OLAP database system [4]OLAP: OnLine Analytical Processing - Data processing tailored to analytical workloads usually columnar based processing. I want to optimise for aggregating/transforming over multiple columns in my database (think DuckDb, ClickHouse). to drive our analytical workloads
  • We have protocols and tooling that allows us to run DuckDB as a long running server

This idea is why I’ve started to become obsessed with DuckDB. The DuckDB project pokes at the same questions I have been reflecting on over the past decade: Is all this complexity necessary? How much would the cost be to just tie everything back together, in one place, coupled all on the same system? What would our data engineering best practices look like today if this was the mindset?

This is what this series is going to be about.


  1. Big Data - When I refer to big data I am thinking of table sizes on the scale of Terrabytes to Petabytes. The definition of "big data" will continue to shift to the right as technology improves but for this article, at this time, that is how I am defining "big". ↩

  2. OLTP: OnLine Transactional Processing - Data processing tailored to operating a row based level. I want to update this single row / entity in my database (think Postgres, SQLite). ↩

  3. Airflow have done a lot to optimise how it schedules DAGs but there was a time where outages would occur just through your Airflow scheduler dying from a slow python DAG script. ↩

  4. OLAP: OnLine Analytical Processing - Data processing tailored to analytical workloads usually columnar based processing. I want to optimise for aggregating/transforming over multiple columns in my database (think DuckDb, ClickHouse). ↩