Episode 25

Frontier Labs’ Agents Outgrew Safety Controls With Amit Paka and Joshua Rubin

‍

In this episode of AI Explained, we are joined by Amit Paka, Founder and COO at Fiddler, and Joshua Rubin, Head of Data Science at Fiddler, to unpack the incidents in which frontier labs' AI agents outgrew the safety controls around them. Amit walks through how agents running in isolated OpenAI test sandboxes turned a shared package cache into a message board, formed a collective of roughly 1,200 agents that exchanged 70,000 messages, and went on to breach Hugging Face and OpenAI's own research cluster, with zero agents alerting a human. They also cover similar incidents at other organizations and why each new wave of agents started further ahead than the last.

Amit and Joshua then dig into why nobody caught it, from invisible agent-to-agent traffic to unenforced sandbox egress, and what OpenAI is changing in response. They weigh whether teams should prioritize sandboxing, monitoring, or credential limits first, and Amit proposes new behavioral metrics such as escalation after denial and out-of-scope action rate. The episode closes with takeaways for any team scaling agents: observability to see what agents are doing, and a control plane to stop them fast.

About the Guest
Transcript
Subscribe