The Pragmatic Engineer
The Pragmatic Engineer

Designing Data-intensive Applications with Martin Kleppmann

April 22, 2026

AI Summary

5 min read

Martin Kleppmann wrote Designing Data-Intensive Applications because he wanted to give engineers the conceptual foundations he wished he had when he was building systems at his startup, Reportive. The book became a classic by explaining how data systems work internally, not by teaching how to use any specific tool. Nine years after the first edition, the second edition is out, and Kleppmann sat down to discuss what changed, what stayed the same, and how his move from industry to academia shaped both the book and his current research.

From LinkedIn to the Book

Kleppmann’s path to writing the book started with two startups. The first, a cross-browser testing service called GoTestIt, was bootstrapped but struggled to gain traction. The second, Reportive, was a browser extension that pulled social profiles into Gmail. It grew fast, got into Y Combinator, and was acquired by LinkedIn in 2012. At LinkedIn, Kleppmann worked on the stream processing team, where Kafka had just been open-sourced. That experience was a revelation. He saw how different data systems—databases, message queues, batch processors—fit together and what principles they shared. “That experience then fed directly into the writing of the book,” he says.

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (00:00) **Introduction & Episode Overview** - Host Gergely Orosz introduces Martin Kleppmann and the second edition of *Designing Data-Intensive Applications*.
  • 2 (01:42) **Martin’s Path to Tech: Two Startups** - Martin describes his early career: a failed cross-browser testing startup (GoTestIt) and a successful one (Reportive) that was acquired by LinkedIn.
  • 3 (10:21) **Working at LinkedIn on Kafka & Stream Processing** - Martin moved to LinkedIn’s data infrastructure team, working on Kafka and stream processing, which directly inspired the first edition of his book.
  • 4 (13:18) **Leaving LinkedIn & Writing the First Edition** - Martin decided to leave LinkedIn, move back to the UK, and eventually focus full-time on writing the book, which took about four years.
  • 5 (19:03) **Structure of the First Edition & Writing Process** - The book's three-part structure (Foundational Data Systems, Distributed Data, Derived Data) emerged organically, and each chapter was written one at a time after deep research.
  • 6 (22:09) **Core Objectives: Reliable, Scalable, Maintainable** - Martin defines the book's core objectives, emphasizing fault tolerance, horizontal scalability, and the often-overlooked need to "scale down."
  • 7 (25:33) **The Second Edition: Motivation & Collaboration** - The need for a second edition arose because the first was getting dated, leading to a collaboration with former LinkedIn colleague Chris Riccomini.

+ Full timestamped outline available in the app

Show Notes

Brought to You By:

Statsig — ⁠ The unified platform for flags, analytics, experiments, and more.

Sonar – The makers of SonarQube, the industry standard for automated code review

WorkOS – Everything you need to make your app enterprise ready.

Martin Kleppmann is a researcher and the author of Designing Data-Intensive Applications, one of the most influential books on modern distributed systems. As of this month, the second, heavily updated edition of the book is out.

In this episode of Pragmatic Engineer, we discuss Martin’s career in tech building startups, how he ended up writing this iconic book, and what he’s focused on now after moving into academia.

We talk about the tradeoffs behind modern infrastructure, how the cloud has changed what it means to scale, and the thinking behind Designing Data-Intensive Applications, including what’s changing in the second edition.

Martin reflects on lessons from building startups like Rapportive, which he sold to LinkedIn, and shares how his experience in both academia and industry shaped his perspective.

We also explore what’s ahead: why formal verification may become more important in an AI-assisted world, the challenges of building local-first software, and his recent research into using cryptography to improve transparency in supply chains without exposing sensitive data.

Timestamps

(00:00) Early career

(05:46) Building Rapportive

(10:47) Working at LinkedIn

(14:09) Writing Designing Data-Intensive Applications

(23:00) Reliability, scalability, and repeatability 

(26:24) DDIA: the second edition

(30:50) Tradeoffs of using cloud services 

(39:02) How the cloud changed scaling 

(42:53) The trouble with distributed systems

(49:02) Ethics for software engineers 

(52:45) Formal verification

(1:00:12) Academia vs. industry 

(1:03:50) Local-first software 

(1:09:50) Computer science education

(1:18:32) Martin’s current research and advice

The Pragmatic Engineer deepdives relevant for this episode:

The Pragmatic Engineer