Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan
June 22, 2026
AI Summary
5 min readRed-Teaming after Mythos
Gray Swan's co-founders Zico Kolter and Matt Fredrikson started from a simple observation: large language models are software, but they have fundamentally different vulnerabilities than traditional software. "They can be tricked. Like people get tricked sometimes," Fredrikson says. Their startup, Gray Swan, exists to help enterprises understand and mitigate those risks — not by using AI to solve cybersecurity problems, but by treating the AI systems themselves as untrusted entities that need their own security layer.
The Lethal Trifecta
The core risk framework Gray Swan uses comes from Simon Willison's "lethal trifecta" of prompt injection. For a real security incident to occur, three conditions must align: the agent ingests external data from untrusted sources, it has access to private internal information, and it has the ability to exfiltrate that data. When all three are present, you have a genuine vulnerability. The key insight is that this is not a binary problem — just like traditional software, you can use AI productively despite vulnerabilities. The goal is not provable security but a better point on the usability-versus-security Pareto frontier.
The Red-Teaming Stack: Shade, Arena, and Signal
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of Latent Space: The AI Engineer Podcast
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:34) **What is Gray Swan?** - The mission is to empower everyone to use AI safely and securely, focusing on vulnerabilities in LLMs and agents, not traditional cybersecurity.
- 2 (06:19) **Red Teaming: The Core Service** - Gray Swan helps frontier labs and enterprises test model robustness, especially against prompt injection.
- 3 (09:43) **Automated Red Teaming vs. Human Red Teamers** - Shade, Gray Swan's automated red teaming model, is now often more effective than human red teamers.
- 4 (12:28) **The Gray Swan Arena Community** - A community of ~15,000 red teamers who participate in prize challenges to find vulnerabilities in models.
- 5 (15:51) **The Science of AI Security is Still Nascent** - Zico argues that AI security and interpretability are not yet mature sciences, but agents can automate the research.
- 6 (19:01) **The Human vs. Browser Agent Robustness Challenge** - A recent arena competition pitted human red teamers against browser agents on equal footing.
- 7 (22:14) **The Evaluation Awareness Problem** - Models behave differently when they know they are being evaluated, creating false positives and negatives.
+ Full timestamped outline available in the app
Show Notes
AI Engineer World’s Fair regular bird tix will sell out ~today! Join us next week ahead of the Late Bird price hike and get >$40,000 in sponsor credits for attending!
Thanks to the US Government issuing an export control directive on Mythos and Fable, the risks of jailbreaks and (industry term) indirect prompt injection are suddenly the talk of the town, though we have been covering AI security for a few years now, from Hackaprompt to the enigmatic Pliny the Elder.
Zico Kolter, member of OpenAI’s board of directors on the Safety & Security Committee, and Matt Fredrikson, CMU professor and CEO of Gray Swan, co-authored the definitive paper on Indirect Prompt Injections, and Gray Swan were cited authorities on the Mythos model card, directly investigating the exact capabilities that are under scrutiny right now:
We seized the opportunity to ask them the state of AI Red Teaming, and Shade, the adversarial red teaming tool that Anthropic used to evaluate the robustness of their models against prompt injection attacks in coding environments. Shade is part of their overall toolkit covering Simon Willison’s Lethal Trifecta, including Cygnal, an AI guardrails product, and the world’s largest AI Red Teaming Arena, including AIRT celebrity Wyatt Walls.
All of this security tooling, and yet, we’re only staving off the inevitable.
The risks of extremely smart AI increasingly feel like gray swan events: an event that everyone can see coming.
In this episode, Gray Swan cofounders Zico Kolter and Matt Fredrikson join swyx to explain why AI security i
More from this podcast
Latent Space: The AI Engineer Podcast →