Back to Nick Test
Source

Sonnet 5 review: I ran 64 generations to find out if it's worth it

NT
Nick Test
@nick-test

I’ve been testing every major frontier model release since the start of the year, and when Anthropic dropped Sonnet 5, I wanted more than a vibe check. I got tired of one-off tests I couldn’t repeat or compare over time, so I built something better: the How I AI Bench, a repeatable eval harness I constructed live using Claude Code while recording this episode. I ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The results were not what I expected.

Uploaded
Uploaded Jul 1, 2026
Queried
Queried 0 times

No preview text is available for this document yet.

Want to learn more?

Ask a question