As Simon Willison notes, the primary interest here may lie less in the benchmark scores and more in the apparent behavioral quirk across reasoning levels. If this effect proves reproducible and significant, it could suggest a more granular control mechanism within the model's architecture than typically seen. However, the community's reliance on unofficial data pipelines for performance metrics highlights a persistent transparency gap in the industry, where crucial evaluation data remains siloed in private channels. This pattern arguably makes independent verification difficult and slows collective understanding of model capabilities.
DeepSeek V4 Pro model appears to show reasoning level differences
A new version of DeepSeek's model reportedly displays varying outputs across reasoning settings, according to an independent analysis.
AIpressr commentary on an article originally published by Simon Willison.
For informational purposes only. AI-assisted commentary may contain errors. full disclaimer ↓hide ↑
This is AIpressr's editorial commentary on a report originally published by another outlet — it is opinion, not the original reporting, and not an endorsement by or affiliation with that outlet. Follow the linked source for the underlying facts. Editorial & AI disclosure.
Editor's Take
Simon Willison's examination of DeepSeek V4 Pro 0813 reveals what appears to be a notable characteristic: the model reportedly generates distinct visual outputs at different reasoning levels. While this observation is intriguing, it raises questions about whether this represents a meaningful technical advancement or simply an artifact of the testing methodology. In our view, the fragmented nature of the benchmark data—which reportedly passed through multiple unofficial channels before being posted—makes it difficult to assess the model's actual capabilities with confidence.
“Interestingly I got very different looking pelicans for the three different reasoning levels of low, medium, and high. I've not noticed this kind of difference from any other model.”
Our analysis
Have AI news to share?
Submit your release →Publisher or subject of this story? Object to this commentary or request a correction →
