One Useful Thing's analysis highlights a crucial, if nascent, trend: the move from prompting to managing. However, the piece arguably glosses over the significant operational and security hurdles that come with deploying autonomous agents at scale. The vision of non-experts becoming effective 'managers' of AI systems presupposes a level of reliability and interpretability that frontier models have not yet consistently demonstrated.

The real test won't be in benchmark scores or internal studies from AI labs, but in whether these systems can operate safely and predictably outside highly curated environments, where their 'jagged' capabilities and unpredictable failures could prove costly.