Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week

October 1, 2026

AI Summary

5 min read

Computer Use Agents: What Actually Changed at OpenAI Dev Day

When OpenAI's Ari and Nikunj sat down for this podcast, they had just come off stage from a keynote that included Dots (personal AI assistants with their own Linux virtual computers), GPT-6.1 (a model that costs one-fifth of Astra for general use and one-seventh for computer use specifically), a new Decisions API, and computer use now built into the Agents API. But the most revealing moment came when they addressed a claim from a prominent AI podcaster that "computer use hasn't advanced in the last two years." Ari's response was direct: "Computer use is like 180 degrees different than that."

The Three Levers That Changed Computer Use

Ari, who leads product and engineering for computer use agents and previously worked on automation at Apple and his own startup Sky (which OpenAI acquired), described the through-line of his career as "helping people automate tasks so you can save time and focus on things more important than operating a computer very intricately." The biggest shift he's seen in the last year isn't just model capability—it's that models have gone from being able to "reliably start tasks" to being "really good at debugging, really good at trying again, introspecting what isn't working."

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (00:04) **Opening & Dev Day Recap** - Ari leads product/engineering for computer use agents; Nikunj leads product for the API team. They recap OpenAI Dev Day announcements including Dots, GPT-6.1 Sole, the Agents API, and the Decisions API.
  • 2 (03:27) **What Computer Use Unlocks** - The core insight: agents can now do anything a human can on a computer because all software was designed for humans.
  • 3 (05:00) **Addressing the "No Progress" Claim** - Directly responds to a prominent AI figure's claim that computer use hasn't advanced in two years.
  • 4 (07:07) **The Three Technical Levers of Progress** - The improvements come from model, harness, and new modalities working together.
  • 5 (09:10) **The Pareto Frontier of Speed vs. Capability** - The harness and model are co-optimized, and the biggest "aha" moments come from introducing new modalities.
  • 6 (11:27) **App Shots Explained** - A deep dive into the feature that pulls in full app metadata, not just a screenshot.
  • 7 (12:17) **The Next Frontier: Superhuman Speed** - Computer use is now faster than average humans; the goal is to be faster than expert users.

+ Full timestamped outline available in the app

Show Notes

Three months ago Dwarkesh, who has been posting incredible blogs and episodes about RL, posted a framing question for his video essay on RLVR which upset a lot of Computer Use folks:

We are no strangers to learning in public and are no strangers to the stress of getting things wrong when you have a big platform. However, we were at Anthropic for the Computer Use launch, there for Claude Cowork with the first big podcast on it, organized the first Computer Use track at AIE presenting the state of the art, and were close to the OpenAI-Sky Software acquisition that now powers the complete domination of computer use that Codex enjoys today. This is why we’re excited to bring you today’s first guest, Ari Weinstein, cofounder of Sky and now leading all the amazing CUA progress that casuals might miss:

Ari explains why Computer Use is now “180 degrees different” from where it was months ago, how agents are learning to debug and recover from failures, why combining screenshots with accessibility data, the DOM, Playwright, and generated code changes the speed equation, and why the next frontier is making agents literally superhuman at using software.

OpenAI clones Jev

In the second half, Nikunj Handa from OpenAI’s API team breaks down the new developer stack: async tool calling, mid-turn steering, WebSockets, UltraFast inference, the Decisions API, prompt caching, pre-warming, compaction, and the Agents API. Given that we were the first Jev podcast, we particularly focus on the unusually fast sprint on the Decisions API:

And why it is just a Luna wrapper for now but the team is motivated and egoless enough to clone what they consider to be good patterns.

We discuss:

* Why OpenAI thinks Computer Use has changed dramatically in just the last few months

* Dots and what changes when every agent gets its own Linux computer

* Why Computer Use can now complete some tasks faster than the average human

* The path from human-level to “literally superhuma

Latent Space: The AI Engineer Podcast