Skip to content
DurkBench

Start here

New here? This page teaches you to read our boards. It takes about a minute.

We pay for Claude Code, Codex, and Grok Build — the same paid coding apps you type into every day. A machine in our lab gives each app the same short jobs every hour and times what comes back. You only read the results: no sign-up, nothing to set up.

Anatomy

One row, taken apart

A leaderboard row with example numbers, not live data. Five parts. Read one row and you can read every board on the site.

  1. 1st place

    Fable 5Fastest

    claude-fable-5 · high

    61.4

    wall tok/s

    Sets the pace

  2. 2nd place

    GPT-5.6 Sol

    gpt-5.6-sol · high

    54.9

    wall tok/s

    11% behind

Example numbers for teaching. The live board is on the home page.

  1. Rank

    The board sorts on wall tok/s, fastest first. 1st place sits on top and its chip fills with the site color. A dash means no clean run yet.

  2. Swatch and name

    The color square names the app: orange is Claude, violet is Codex, green is Grok. The name links to that model's page.

  3. Pin and effort

    The small line under the name: the exact model we ask for, then the thinking setting we pinned. Each model runs at the top of its own scale.

  4. Wall tok/s

    The ranking number and the bar. Tokens per second from enter to done. Higher is better, and it is fair across all three apps. Bars share one scale, so a longer bar means a faster model.

  5. The pace line

    The small line under the number. The leader says “Sets the pace”. Every other row says how far behind it sits, like “11% behind”.

Model pages carry the detail numbers: stream speed, time to first word, clean runs, and change since yesterday.

Open the live board and read along.

The words on the boards

The seven terms you will meet on the boards, in plain English.

Token
A token is a small chunk of text, roughly four letters. Models read and write text one token at a time.
Wall tok/s
Tokens per second across the whole command — you press enter, the app starts, thinks, writes, and finishes. Higher is better. It is the only speed number that is fair across all three tools.
Ping
A tiny check that asks for the word “ok”. It shows how long the app takes to wake up and reply. It is not a writing-speed test.
Round
A one-hour slot. Our lab runs one set of checks per round, at the top of each hour. We try once. A fail stays a fail.
Usual
The middle value of a column, not the average. Line the samples up smallest to largest and take the one in the middle. One freak run cannot drag it around.
Clean round
A round where the tool answered and the answer was right. A plain script checks it. Only clean rounds set a speed number, so a fast wrong answer earns nothing.
An empty cell
No clean run for that model inside the window yet. It does not mean zero and it does not mean broken. We leave it blank on purpose.

One rule above all: rank on wall tok/s. Every other number is context.

Questions people ask first

Is my tool slow right now, or is it just me?
Open the live board. If our lab is slow on that tool in the round that just finished, it is not you. If our numbers look normal, the trouble is closer to your machine, your network, or your prompt.
Why is a cell empty?
That model has not finished a clean run inside the window yet. We would rather show a blank than a guess, so empty means empty.
Why is there no single winner?
Write speed and the wake-up ping answer different questions. A tool that writes fast but takes ages to start has not won anything, so we never blend the two into one score.

Next: open the Live board, see Today’s card, or read For teams if you buy seats.

How to read this

Every number here comes from one lab machine, on one connection, with the app versions listed on Status. These are our own measurements, not official vendor numbers.