Simon Willison's hands-on testing reveals a critical, second-order issue for the open-source AI ecosystem: developer ergonomics. While benchmarks for a model like Qwen 3.8 27B might show impressive gains, its default behavior could render it nearly unusable for many practical applications without manual tuning. This serves as a reminder that model performance is not just a function of parameter count or benchmark scores, but of the entire user experience stack, including sensible presets.

In our view, this misstep may hinder adoption among the very developers—those running models on local hardware—that open-weight models are meant to empower, pushing them back toward more polished, closed APIs where such defaults are carefully managed.