AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

DuckDB 2.0’s alpha adds asynchronous read-ahead for external files, allowing S3 downloads and Parquet decoding to happen at the same time. A MotherDuck report measured faster reads in its tests, while also describing a recursive CTE rewrite and a claimed graph-reachability speedup; the supplied results do not establish general performance gains.

DuckDB 2.0’s alpha changes how the database reads data from external storage, letting download threads fetch data while query workers process previously fetched data. In MotherDuck’s tests on one laptop, that design cut several S3 read times, though the results are preliminary and do not show how the release will perform across machines or workloads.

The reported change centers on asynchronous I/O. In DuckDB 1.5.5, query workers take turns downloading and decoding Parquet row groups. When a worker waits for network data, its CPU work pauses. In the 2.0 alpha, a separate download pool fetches row groups ahead of the workers, which can decode while more downloads continue. The setting read_ahead_depth controls how far downloads run ahead; it defaults to -1, which the report says selects an automatic value, while 0 restores the earlier behavior.

MotherDuck’s author tested the same queries on an M5 laptop against AWS S3. Reading one column from a 2.2 GB Parquet file took 18.8 seconds with DuckDB 1.5.5 and 7.7 seconds with the 2.0 alpha. A test across 23 Parquet files totaling 13.6 GB took 11.8 seconds versus 3.9 seconds. A 1.7 GB CSV read took 116 seconds versus 55 seconds. These are single-machine results, not independent benchmark findings or a guarantee of equivalent gains elsewhere.

The report also describes a rewritten recursive CTE engine, used for queries that repeatedly follow relationships such as manager-to-employee links. It says the DuckDB team claims a 40-fold improvement on graph reachability. The supplied source excerpt begins explaining a hierarchy example but does not include benchmark conditions or results supporting that figure, so its scope cannot be evaluated here.

At a glance
reportWhen: DuckDB 2.0 alpha reported ahead of a pl…
The developmentMotherDuck’s report on the DuckDB 2.0 alpha describes asynchronous I/O and a rewritten recursive CTE engine as sources of potential speed improvements.

Faster Reads From Object Storage

For users who query data directly in Amazon S3, overlapping network waits with CPU work may shorten runtimes without requiring changes to SQL. That could matter for analytics pipelines that repeatedly scan large files, particularly when the workload has enough row groups to keep downloads and processing moving in parallel.

The measurements also show why the improvement is not universal. The report found little difference for 30 small Parquet files, each about 1 MB: the combined read took 3.7 seconds in version 1.5.5 and 3.3 seconds in the alpha. The author attributes the limited change to per-file network round trips that read-ahead cannot eliminate. File layout and storage location remain relevant to performance; version 2.0 does not remove those constraints.

For teams evaluating an upgrade, the practical point is to test representative queries and datasets. A faster read in this setup may not predict performance on a different cloud region, network, file format, or workload. The recursive CTE claim may matter to users traversing hierarchical or graph-like data, but the excerpt provides too little detail to establish what queries benefit or how the result was measured.

Amazon

Amazon S3 compatible data storage

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Changed in the Alpha

The report presents DuckDB 2.0 as an upcoming release, with an alpha available at the time of writing and a planned release “this fall.” It does not specify the year in the supplied material. Its focus is on changes relevant to people building tables and data pipelines, rather than a full account of the release.

In the S3 example, the source file contains 228 million rows in 2,268 row groups. The query reads one of four columns and counts votes by type. The author disabled DuckDB’s external file cache so repeated runs would access S3, and noted that home internet slowed both versions. The report says cloud compute near the storage could produce faster absolute times, but gives no results from such a setup.

Another test read 30 small Parquet files, where the author saw no meaningful improvement. That contrast is relevant because the new mechanism addresses overlapping downloads and processing, not every source of query latency. The supplied material mentions other “hidden gems” in commit logs but does not describe them, so they are not included here.

“DuckDB 2.0 is faster.”

— MotherDuck report author

Amazon

high performance external storage SSD

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Performance Evidence

The figures are from one computer and one network connection, and the source does not provide repeated-run variation, hardware configuration details beyond the M5 laptop, or results from other cloud regions. It is also unclear how performance changes across different thread counts, file layouts, query types, and storage providers.

The report’s stated 40-fold graph-reachability gain is attributed to the DuckDB team, but the supplied material does not provide the workload, baseline, measurement method, or detailed results. It should be treated as a reported claim, not a general benchmark conclusion. The release timing is described only as “this fall,” without a year or a final release date.

Amazon

Parquet file reader for Amazon S3

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing Before the Stable Release

DuckDB 2.0 was described as being in alpha, with a stable release expected in the fall according to the report. Users considering the change can compare the alpha with their current version using their own data and queries, paying particular attention to network location, file sizes, and the number of row groups.

Further evidence would be needed to establish how broadly the measured gains hold, and to clarify the recursive CTE benchmark behind the reported 40-fold figure. The source does not provide a firm release date or additional benchmark schedule.

Amazon

database query acceleration tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is making DuckDB 2.0 faster on S3 reads?

The alpha uses a separate download pool to fetch data ahead of query workers. Workers can decode available data while downloads continue, reducing time spent waiting for network responses in some workloads.

How much faster was the S3 test?

In MotherDuck’s test on one M5 laptop, reading one column from a 2.2 GB Parquet file took 7.7 seconds in the 2.0 alpha and 18.8 seconds in version 1.5.5. The result is specific to that setup and is not a universal benchmark.

Does every dataset get the same speedup?

No. The report measured only a small change when reading 30 Parquet files of about 1 MB each: 3.7 seconds in version 1.5.5 compared with 3.3 seconds in the alpha. File layout and network overhead affect results.

What is the recursive CTE improvement?

The report says DuckDB rewrote its recursive CTE engine and attributes a 40-fold graph-reachability improvement to the DuckDB team. The supplied excerpt does not include enough benchmark detail to assess the claim’s scope.

When will DuckDB 2.0 be released?

The report says the alpha is out and the release is planned for “this fall,” but the supplied material does not identify the year or give a specific release date.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What SemiAnalysis Found About The 5X And AI Subscription Prices

SemiAnalysis estimates Claude subscriptions offer about 5.4–5.6 times ChatGPT’s API-equivalent value on mid-tier models, amid changing limits and prices.

OpenAI Connects The Dots

A Platformer columnist reports that OpenAI’s Dots handled work tasks in early testing, while access, pricing and reliability remain open questions.

A Look At Gewerkton’s Two-Day Relaunch And Its AI Agents

Gewerkton says its app, Studio and 529-page website moved to HORIZON on Oct. 2-3, with AI agents doing implementation under human review.

Show HN: Bigwords.page – Turn Any Screen Into A Sign. The URL Is The App

Bigwords.page is a browser-based tool for displaying signs, messages, slides and countdowns from a link, without an account or app.