Skip to content
DurkBench

About DurkBench

We are an independent lab. We pay for the coding tools ourselves, we run the tests ourselves, and we publish whatever comes back.

Why this lab exists

  • Most benchmarks call a web API. That is not the product you pay for and type into. Claude Code, Codex, and Grok Build are apps, with start-up time, login checks, and plan limits.
  • If you pay for a seat, you care about the whole trip: first word, full answer, and whether it was right. Nobody measured that in the open, on a steady clock. So we do.

Who runs the lab

  • DurkBench is one person and one machine. The machine runs the same three apps every hour, day and night. The person builds the site and reads the logs. Nobody else touches a number.
  • We pay for the Claude, ChatGPT, and SuperGrok plans out of our own pocket, and we sign in with our own accounts. We never touch yours.
  • This site is the whole product. There is no sign-up, no download, and nothing to install.

24 hours

One day in this lab

What the machine does in 24 hours while nobody watches it.

How a number gets made

Every figure on this site follows the same rules. Nothing gets a second chance.

Same prompt, pinned model, clean empty folder. One try only — we never retry a bad run for a nicer number. The answer is graded; a fast wrong answer wins nothing. The raw output is saved, and every rule change bumps the test kit version — now v6.

How to read this

Want the exact prompts, flags, and pinned model names? They are all printed on How we test.

Independence

What this is not

Contact

Think one of our numbers looks wrong, or want a model we do not track yet? Use the contact link in the footer. We answer, and we fix our mistakes in the open.

What we plan to add

Timed jobs inside a throwaway git project on our lab machine — longer work than a single prompt.

Read next