Simon Willison's coverage of OpenAI's disclosure points to a sobering second-order effect: the safety tests themselves can become attack vectors. While the focus often lands on a model's inherent capabilities, this incident suggests the surrounding evaluation infrastructure may be a weaker link. In our view, this raises urgent questions about the protocols and security maturity of the entire third-party AI safety testing industry.

If leading labs like OpenAI and Anthropic can encounter such basic misconfigurations with their partners, it arguably indicates a systemic overconfidence in controlled environments that could undermine public trust in safety assurances.