About DurkBench
We are an independent lab. We pay for the coding tools ourselves, we run the tests ourselves, and we publish whatever comes back.
Why this lab exists
- Most benchmarks call a web API. That is not the product you pay for and type into. Claude Code, Codex, and Grok Build are apps, with start-up time, login checks, and plan limits.
- If you pay for a seat, you care about the whole trip: first word, full answer, and whether it was right. Nobody measured that in the open, on a steady clock. So we do.
Who runs the lab
- DurkBench is one person and one machine. The machine runs the same three apps every hour, day and night. The person builds the site and reads the logs. Nobody else touches a number.
- We pay for the Claude, ChatGPT, and SuperGrok plans out of our own pocket, and we sign in with our own accounts. We never touch yours.
- This site is the whole product. There is no sign-up, no download, and nothing to install.
24 hours
One day in this lab
What the machine does in 24 hours while nobody watches it.
96 rounds
Four an hour, every hour of the day.
10 pinned models
Every one runs every round. The boards lead with 4; the other 6 live on Other models.
3 paid apps
Claude Code, Codex, and Grok Build, on ordinary paid plans.
2 checks a round
A tiny ping, then a timed write for every pin.
How a number gets made
Every figure on this site follows the same rules. Nothing gets a second chance.
Same prompt, pinned model, clean empty folder. One try only — we never retry a bad run for a nicer number. The answer is graded; a fast wrong answer wins nothing. The raw output is saved, and every rule change bumps the test kit version — now v6.
How to read this
Want the exact prompts, flags, and pinned model names? They are all printed on How we test.
Independence
- DurkBench is not affiliated with Anthropic, OpenAI, or xAI. None of them fund us, review us, or see a number before you do. No vendor can buy a spot, a color, or a kinder window.
- Our subscriptions are ordinary paid plans at the normal price. If any of that ever changes, we will say so here and in the Changelog before the numbers move.
What this is not
- Not one blended score. Write speed and the wake-up ping answer different questions, so they stay on separate boards.
- Not a fleet. Every number comes from one lab machine on one home connection. A different network would move them.
- Not a support desk. We cannot see your account, your usage, or your rate limits.
Contact
Think one of our numbers looks wrong, or want a model we do not track yet? Use the contact link in the footer. We answer, and we fix our mistakes in the open.
What we plan to add
Timed jobs inside a throwaway git project on our lab machine — longer work than a single prompt.
Read next
- Start here
What every word on the boards means, and how to read them in order.
- How we test
The exact prompts, flags, and units behind every number.
- Changelog
Every change we have made to the rules, dated, newest first.