Simon Willison's smevals tool, as described in his announcement, addresses a genuine pain point—the friction of setting up reproducible evaluations—but its success hinges on community adoption beyond his own projects. The architecture, separating runs, graders, and checks, is sensible, yet the ecosystem for AI evals is already crowded with both heavyweight industry benchmarks and other lightweight frameworks. The tool's utility may be greatest not as a new standard, but as a pragmatic, personal workflow accelerator that others can fork and adapt.
In our view, the more interesting second-order effect is the continued fragmentation of the evaluation layer, suggesting that no single approach has yet captured the full complexity of judging model performance, especially for creative or agentic tasks.
