Skip to content
DurkBench

Real tests of the AI coding tools you pay for.

Claude Code, Codex, and Grok Build — timed on the same paid plans you buy. How fast they write, and whether the answer is right. For people on a subscription, not an API key.

SpeedWall tok/s

Live33m ago

  1. Grok 4.641.59
  2. Opus 531.83
  3. Fable 530.30
  4. GPT-5.6 Sol22.54
Measured
10:02 PM
Timezone
ET
Test kit
v6
Models measured
10
Every round
Rounds a day
24
One every hour
Clean rounds
99%
318 of 320 in 24 h
Rate limits
0
In the last 24 h

Highlights

Today, in two numbers

The same two questions we ask every round, for each lab's leading model: how fast it wrote, and how often it finished a clean round.

Speed

Middle wall tokens per second, last 24 h · higher is faster

  1. Grok 4.632/32 clean41.59
  2. Opus 532/32 clean31.83
  3. Fable 532/32 clean30.30
  4. GPT-5.6 Sol32/32 clean22.54
041.59 tok/s

4 of 10 models

Other models

Finished the round

Clean rounds out of rounds tried, last 24 h · higher is better

100%

All 4 tied at 32/32 clean. Nothing to rank in this window.

4 of 10 models

Lab status

Both boards cover the last 24 hours. Each board has its own scale · test kit v6. 4 of 10 pinned models shown · the rest live on Other models.

Speed

Speed, ranked

The 4 models we lead with, ranked on the middle wall tok/s of their clean rounds in the last 24 hours. Longer bar is faster.

Other models

Tokens per second, start to finish. Higher is faster.

The 4 models on the home boards, ranked on median wall tokens per second over the last 24 hours.
RankModelBar, on one shared scaleWall tok/sCleanRounds
1Grok 4.6xAI · effort xhigh41.59100%32/32
2Opus 5Anthropic · effort high31.83100%32/32
3Fable 5Anthropic · effort high30.30100%32/32
4GPT-5.6 SolOpenAI · effort high22.54100%32/32

Window 24 hours · 4 of 4 models scored · n = 128 rounds tried · test kit v6. Day-over-day change appears here once the lab has two full days of rounds in the log. 4 of 10 pinned models shown · the rest live on Other models.

24 hours

The last 24 hours, round by round

Each line is one model's middle speed over an hour of rounds. The faint dots behind it are the single rounds, so the swings stay in plain sight.

Open today's report

Wall tokens per second · higher is faster

nothing recorded before 12:00am · the plot starts there

tok/s020406080
12am6am12pm6pmnow
Aug 20

Line: the middle wall tok/s of each 4 hours of testing — up to 4 rounds. Dots: those single rounds. Nothing is filled in, so a stretch with no clean round breaks the line, or ends it early.

Grok 4.6median41.5923/24 cleanOpus 5median31.8323/24 cleanFable 5median30.3023/24 cleanGPT-5.6 Solmedian22.5423/24 clean
Full record — every round, every number
Wall tokens per second for every round in the last 24 hours. A dash means that round had no clean result.
RoundGrok 4.6Opus 5Fable 5GPT-5.6 Sol
Aug 20 10:00pm41.2132.3628.6615.38
Aug 20 9:00pm40.7028.7730.4922.34
Aug 20 8:00pm39.5135.1231.8126.26
Aug 20 7:00pm49.4130.2732.0423.67
Aug 20 6:00pm38.3233.5231.4326.76
Aug 20 5:00pm49.1941.5430.1123.22
Aug 20 4:00pm43.2832.9529.6417.36
Aug 20 3:00pm47.8036.5732.9024.96
Aug 20 2:00pm37.0333.5027.2824.98
Aug 20 1:00pm48.4334.5832.4716.95
Aug 20 12:00pm50.0331.2231.6621.82
Aug 20 11:00am34.3933.3526.6419.65
Aug 20 10:00am45.2928.5931.8022.56
Aug 20 9:00am53.9426.7730.0422.23
Aug 20 8:00am43.6332.0832.1019.56
Aug 20 7:00am39.6133.7629.8726.41
Aug 20 6:00am5.5533.4524.6524.46
Aug 20 5:00am40.4831.5734.6021.38
Aug 20 4:00am45.8631.2429.7522.57
Aug 20 3:00am32.5031.9423.8624.68
Aug 20 2:00am44.1229.3932.6523.79
Aug 20 1:00am41.2725.9829.0120.34
Aug 20 12:00am17.6733.5718.9822.52
Aug 19 11:00pm

Window 24 hours · one round every 60 minutes · 24 rounds on the clock · n = 92 clean readings drawn · test kit v6. The lab recorded nothing in the first 1 hours of this window, so the plot starts at 12:00am rather than drawing empty space. Those rounds are still in the full record above, as dashes. 4 of 10 pinned models shown · the rest live on Other models.

We also ping each app every round to check it answers at all. Those checks, round by round, live on Lab status.

Method

How we get these numbers

We run the real paid apps on our own lab machine, not the hidden APIs. Every number traces back to a saved run.

Open How we test
Cadence
Every hour
Rounds a day
24
Test kit
v6
Window
24hours
Test kit v6
Old rows stay in Logs. They do not mix into these boards.
Changelog
Paid apps
We run Claude Code, Codex, and Grok Build ourselves. Not the hidden APIs.
Status
Graded answers
A round counts only when the answer is right. Not “it printed text.”
How we test
Saved proof
A public number traces back to a saved run in our archive.
Logs

Last round 35m ago · 483 results saved in the last 24 hours, 5 failed · Lab status.