As Simon Willison notes, the primary interest here may lie less in the benchmark scores and more in the apparent behavioral quirk across reasoning levels. If this effect proves reproducible and significant, it could suggest a more granular control mechanism within the model's architecture than typically seen. However, the community's reliance on unofficial data pipelines for performance metrics highlights a persistent transparency gap in the industry, where crucial evaluation data remains siloed in private channels. This pattern arguably makes independent verification difficult and slows collective understanding of model capabilities.