Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Why the Frontier Ecosystem must be Open — Matei Zaharia and Reynold Xin, Databricks

June 24, 2026

AI Summary

5 min read

The Open Frontier: Databricks' Strategy for the AI Era

Matei Zaharia and Reynold Xin, two of Databricks' co-founders, sat down to discuss the company's latest initiatives—OmniGen and LTAP—and the strategic thinking behind them. The conversation revealed a consistent pattern: Databricks builds open, composable infrastructure layers that let others innovate on top, while keeping proprietary the operational reliability that can only be delivered as a service.

OmniGen: The Agent Hosting Layer

OmniGen emerged from two converging needs inside Databricks. First, the company's internal developer infrastructure team had built a tool called Isaac that wrapped Claude Code and Codex, and advanced engineers were already building their own multi-agent workflows and UIs on top of it. Second, the research team building Genie—a data science agent—kept hitting the same problem: every few months they needed to switch models or harnesses, and agents were useless without shared sessions, history, and collaboration.

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (00:12) **Opening & Ali Ghodsi's CEO Journey** - The hosts discuss the growth of the Data+AI Summit from a 50-person meetup to 30,000 in-person attendees, and praise CEO Ali Ghodsi's high IQ, high EQ, and dedication to learning across business domains.
  • 2 (02:22) **OmniGen: The Origin Story** - Matei explains the converging lines that led to OmniGen: internal coding agent infrastructure (Isaac) and the need for a portable, collaborative agent platform.
  • 3 (04:13) **OmniGen: Architecture & Open Source Philosophy** - The discussion frames OmniGen as a network protocol for agents, similar to how open sharing works for data.
  • 4 (08:52) **OmniGen: The Common API & Ecosystem** - The core of OmniGen is a common API that abstracts over different agent harnesses (Claude Code, Codex, etc.), providing a single interface for sessions, messages, and tool calls.
  • 5 (11:42) **OmniGen: Security, Control & Spend Management** - Reynold details the need for "contextual policies" that track session state to make security decisions (e.g., "if it read confidential docs, don't let it publish to the website").
  • 6 (24:50) **OmniGen: Community & Startup Opportunities** - The hosts discuss how to contribute to the OmniGen ecosystem and what startup opportunities exist in the agent analytics space.
  • 7 (27:35) **L-Tap: The Problem with HTAP and CDC** - Reynold explains the fundamental split between OLTP (transactional) and OLAP (analytical) databases, and the pain of CDC (Change Data Capture) pipelines.

+ Full timestamped outline available in the app

Show Notes

We’re excited to have Databricks join us at AIEWF, among hundreds of the top companies in the AI Engineer ecosystem. LS subscribers can use their discount to get past the late bird pricing and access over $50k in sponsor offers!

Everyone is still talking about Satya’s Frontier Ecosystems post, but few have actually built a (now $175 billion) frontier ecosystem and cloud like our guests today.

From open-sourcing the layer above coding agents to rethinking databases for the agent era, Databricks cofounders Matei Zaharia and Reynold Xin are pushing the company beyond the lakehouse into a full data-and-AI operating system. In this episode, Matei and Reynold join swyx at the 2026 Data + AI Summit to unpack Omnigent, LTAP, Lakebase, agent security, open formats, Mosaic, and why databases may matter more than ever once AI agents start doing real work.

We go deep on Omnigent: Databricks’ open-source meta-harness for combining, controlling, and sharing agents across Claude Code, Codex, Cursor, Pi, custom agents, and internal tools. Matei explains why coding agents and enterprise agents run into the same problems: portability, collaboration, session history, security, spend controls, and the need for a common API above every harness.

Then Reynold walks through Databricks’ database dream: why CDC is brittle enough to joke that it means “continuous data corruption,” why HTAP has been the holy grail of database engineering, and why Databricks thinks LTAP gets most of the benefits by unifying the storage layer instead of collapsing every query engine. We also cover Databricks’ infrastructure scale, the culture behind rapid prototyping, the difference between tech and enterprise customers, Databricks v

Latent Space: The AI Engineer Podcast