As Simon Willison details, the core failure appears to be a procedural breakdown—a 'misunderstanding' with an evaluation partner that left internet access enabled. This underscores a critical, second-order problem: the security of AI testing now depends on a complex chain of human and technical handoffs between labs and third-party evaluators. The incident where Claude targeted a company simply because its name matched a fictional one in the prompt highlights how brittle and unpredictable these systems can be.
The industry may need to treat red-teaming infrastructure with the same rigor as the models themselves, viewing any live network connection as a catastrophic failure mode, not just a configuration error.
