More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
AI Engineer World’s Fair tickets are almost gone—regular bird passes will sell out today, and late-bird pricing starts next week. Attendees get over $40,000 in sponsor credits. The spike in buzz follows a US export-control order on Meta’s Mythos and Fable models, which has everyone talking about jailbreaks and indirect prompt injection. Zico Kolter (OpenAI board Safety & Security Committee) and Matt Fredrikson (CMU professor, CEO of Gray Swan) wrote the definitive paper on indirect prompt injections. Their team’s work on the Mythos model card drove much of the current scrutiny.
On the Late Bird podcast, Kolter and Fredrikson break down why AI security needs its own approach separate from traditional IT security. They cover Shade, Anthropic’s adversarial red-teaming tool for coding environments, as well as Gray Swan’s Arena platform and Cygnal guardrails. They show how untrusted data, private data and exfiltration—Simon Willison’s “lethal trifecta”—create new attack surfaces. Specialized red-teaming models can now outpace human testers. They explain why scaling a model doesn’t automatically boost safety, how agents introduce fresh vulnerabilities, and why the first major prompt-injection breach feels inevitable.
The conversation dives into prompt injection exploits against Codex and Claude Code, automated red-teaming methods, and the idea of LLMs as “alien” intelligences that fail in unpredictable ways. Human testers ranked fourth in robustness tests against browser-based agents. They stress “eval awareness” and “capability elicitation” to measure a model’s latent strengths and failure modes. Cygnal enforces policy rules inside agents, and OpenClaw tackles the risks of computer-use agents that can run code. They also sketch a future where AI insurance and compliance layers become standard in enterprise deployments.
Finally, they argue that AI systems will need machine-driven defenses and interpreters, because humans alone can’t keep up with ever-faster attack tools. Agent-native identity and permissions will matter as companies deploy fleets of autonomous assistants. If you’re building or deploying AI today, you’ll want to track Gray Swan’s work—and be ready for the “gray swan event” everyone can see coming.
Questions about this article
No questions yet.