01 September 2026
55 new posts across 9 database and data engineering sources — 9 worth your time.
How one connection kills a database
Analyzes a classic database outage scenario where an uncommitted transaction holding an exclusive lock on a table blocks a pending schema change. This block cascades, preventing all subsequent queries from executing and ultimately taking down the database.
PlanetScale · 01 September 2026 · Read the original →
Andrew Atkinson: PostgreSQL 18: 23x Faster Inserts With UUID V7
Shares a real-world migration from UUID v1 and v4 primary keys to UUID v7, resulting in 23x faster inserts. It discusses the performance benefits of sequential keys and the locking trade-offs of altering column defaults.
Planet PostgreSQL · 01 September 2026 · Read the original →
Radim Marek: Read your own writes, off the primary
Addresses the common issue of replication lag where a user writes to the primary database but immediately reads stale data from an asynchronous replica. It explores architectural patterns to ensure read-your-own-writes consistency.
Planet PostgreSQL · 01 September 2026 · Read the original →
semab tariq: How to Calculate Fillfactor for Your Tables
Provides a practical guide and mathematical formula to calculate the optimal fillfactor for PostgreSQL tables. Adjusting this setting leaves free space on pages to accommodate in-place (HOT) updates, reducing write amplification.
Planet PostgreSQL · 01 September 2026 · Read the original →
How DuckDB Runs Recursive CTEs Faster
Deep dives into DuckDB's query engine internals to show how optimizing the scope of reusable runtime state speeds up recursive CTEs. By keeping state active across iterations instead of rebuilding it, the engine avoids redundant evaluation overhead.
DuckDB · 01 September 2026 · Read the original →
Ensuring reliable OpenTelemetry ingestion at scale
Explains the architecture and engineering decisions behind ClickHouse Cloud's ingestion pipeline, which processes 50 million OpenTelemetry events per second. It focuses on maintaining reliability and performance at extreme scale.
ClickHouse · 01 September 2026 · Read the original →
Problems with large tables in Postgres
Examines how tables that grow excessively large, wide, or 'fat' (containing oversized out-of-line values) degrade database performance in distinct, predictable ways. It outlines the specific operational bottlenecks associated with each type of growth.
PlanetScale · 01 September 2026 · Read the original →
GPU-Accelerated Spark on EMR for Faster Analytics
Amazon EMR on EKS now runs Apache Spark workloads up to 3.7x faster and 31% cheaper on Amazon EC2 G7 instances with NVIDIA RTX PRO 4500 GPUs compared to comparable CPU instances. In TPC-DS 3 TB benchmarks, G7 instances completed the 103-query power run in 4.7 minutes, saving 750 seconds per run.
aws.amazon.com · 01 September 2026 · Read the original →
Oracle XStream CDC to Kafka for High-Volume Streaming
The Oracle XStream CDC Source connector for Confluent Platform captures changes from Oracle databases to Apache Kafka topics, supporting high-volume scenarios of over 150 GB of CDC data per hour using Oracle's XStream API. It offers two topologies, including downstream capture to offload capture workload from the source OLTP database, ensuring production performance is not affected.
Confluent · 01 September 2026 · Read the original →
We will do absolutely anything to avoid just buying a larger database instance.
9 stories, every Tuesday
Published here every week. Follow by RSS to get it as it lands.