Sonnet 5 review: I ran 64 generations to find out if it's worth it
NT
Nick Test
@nick-test
I’ve been testing every major frontier model release since the start of the year, and when Anthropic dropped Sonnet 5, I wanted more than a vibe check. I got tired of one-off tests I couldn’t repeat or compare over time, so I built something better: the How I AI Bench, a repeatable eval harness I constructed live using Claude Code while recording this episode. I ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The results were not what I expected.
- Uploaded
- Uploaded Jul 1, 2026
- Queried
- Queried 0 times
No preview text is available for this document yet.
Want to learn more?
Ask a question