Skip to content
DurkBench

Claude Code

Sonnet 5

Inside Claude Code we ask for claude-sonnet-5 every hour. Everything below comes from those runs on our lab machine.

Claude Codeclaude-sonnet-5Think setting highLast round 2h agoOther models

Write speed, last 24 hours

32.1wall tok/s

Tokens per second across the whole command — you press enter, the app starts, thinks, writes, finishes. Higher is better. This is the middle value of 32 clean rounds out of 34 tried in the last 24 hours.

Time to first word
2.88s
How long before the first visible word shows up. Lower is better.
Middle wall time
7.69s
How long the whole command took, start to finish, including the app booting. Lower is faster.
Last ping
The newest tiny “reply ok” check. It measures wake-up time, not writing speed.

What stands out

Four facts pulled straight from the saved runs.

Board position
#4 of 10
Just ahead: Grok 4.6 at 41.8 wall tok/s. Just behind: Opus 5 at 31.8 wall tok/s.
Rounds that worked
32/34
Clean rounds out of rounds tried, in the last 24 hours.
Served as
claude-sonnet-5
The model name the app reported back. If it differs from the pin, we say so.
Round-to-round swing
1.5×
How much the fastest round beat the slowest one. A big swing means a noisy day.
Rate limit warnings
2
Rounds the vendor served but flagged the plan's usage window. They worked, but we keep them out of the medians to be safe.

Recent rounds

Newest first. Fails stay in the list with their reason.

WhenResultWall tok/s
1m agorate limit warningclean34.9
1h agorate limit warningclean36.7
2h agookclean32.2
3h agookclean36.4
4h agookclean35.6
5h agookclean30.6
6h agookclean35.7
7h agookclean35.3

All rounds are in Logs.

Read next

Compare with GPT-5.6 Terra

Its closest rival in Codex — same window, same units, with the leader marked on every row.

Compare with Haiku 4.5

The pin next to it in Claude Code. See what the tier jump buys inside one app.

More models

How to read this

Every number here comes from our lab machine running the paid app, one try per hourly round, on test kit v6. Our measurements, not the vendor’s.