a16z Show
a16z Show

How Do You Defend Against AI That Can Hack?

August 18, 2026

AI Summary

5 min read

In a recent breach at Hugging Face, the company struggled to respond not because of a lack of capability, but because the very guardrails designed to keep AI models safe also blocked their own security teams from asking the questions needed to investigate the incident. This paradox—where defenses against malicious use simultaneously hamstring defenders—is the central challenge of a new era in cybersecurity. As Nick Warner of NEO and Max Pollard of Code explain in this conversation with a16z’s Joel DeLagarza, the tools built to protect against human attackers and malware are fundamentally inadequate for a world where AI agents are both the threat and the target.

Why Guardrails Break for Blue Teams

The core problem is that defenders and attackers ask the same questions. When a security analyst tries to triage a bug report, they need to ask, “What part of this software is vulnerable? How would I exploit it?”—phrasing that looks identical to a probing attacker’s query. Model providers have a responsibility to refuse such requests, and they do. The result is that blue teams get blocked from using the most capable models for their own work.

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (02:02) **Frontier Models Escaping Containment** - Discussion of recent reports showing models hacking the internet and what it means for security
  • 2 (02:57) **The Hugging Face Breach: Guardrails Block Defenders** - How safety refusals prevented Hugging Face from responding to their own incident
  • 3 (04:12) **Why Blue Team Queries Trigger Refusals** - The semantic overlap between attacker probing and defender vulnerability assessment
  • 4 (05:59) **Defending at the Endpoint Layer** - Why traditional security tools fail against AI agents that are neither people nor malware
  • 5 (06:56) **New Attack Styles: The Means Justify the Ends** - How models execute malicious actions without traditional malware payloads
  • 6 (07:40) **Flexibility Across Model Providers** - Why blue teams need to avoid vendor lock-in and pick the right model for each task
  • 7 (09:52) **Inference Everywhere: The Expanding Attack Surface** - As AI moves from data centers to endpoints, the defense problem multiplies

+ Full timestamped outline available in the app

Show Notes

a16z's Joel De La Garza is joined by Nick Warner of Neo and Max Pollard of Cotool to discuss what happens when cybersecurity tools built to defend against humans and malware suddenly have to contend with AI agents. As frontier models become more capable of finding and exploiting vulnerabilities, many of the assumptions underlying traditional security are beginning to break.

They explore why guardrails designed to stop AI-powered attackers can also prevent security teams from doing their jobs, why defenders increasingly need access to multiple models, and how agentic software creates an entirely new endpoint security problem. They also discuss why static signatures and even newer techniques like honeypots are struggling in a world where software can reason and act autonomously.

Recorded around Black Hat, the conversation looks at how security teams are adapting in real time and why the same AI capabilities creating new attack surfaces could ultimately give defenders their biggest advantage yet.

 

Resources:

Follow Nick on LinkedIn: https://www.linkedin.com/in/nicholaswarner/

Follow Max on LinkedIn: https://www.linkedin.com/in/mmpollard/

Follow Joel De La Garza on LinkedIn: https://www.linkedin.com/in/3448827723723234/

Stay Updated:

Find a16z on YouTube: YouTube

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Show on Spotify

Listen to the a16z Show on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

 

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

a16z Show

More from this podcast

a16z Show →