Skip to content
BetaDurkBench v0.1 is in Beta.Numbers are real, but pages and rules can still change.See what changed

Side by side

Compare two models

Put two models side by side. Same tests, same units. Each row names its leader.

The short answer · test kit v6

Grok 4.5 writes 2.0× as fast as GPT-5.6 Terra right now.

Grok 4.5
46.4wall tok/s
GPT-5.6 Luna
25.9wall tok/s
GPT-5.6 Terra
23.8wall tok/s

Speed is one question. Check the clean-rounds row too — a fast model that keeps failing is not the faster tool.

Pick the models

Your picks live in the link. Copy it and the matchup travels with it.

Swap the two columns

Row by row

The marked cell leads its row. An empty cell means no clean run in this window — not zero.

RowGrok 4.5GPT-5.6 LunaGPT-5.6 Terra
Write speedwall tok/s · every token written over the whole command · higher is better46.4leadswall tok/s25.9wall tok/s23.8wall tok/s
Time to first wordhow long the screen stays blank · Codex does not stream, so its cell stays empty · lower is better3.10s
Typical waitstart to finish for the whole command · lower is better6.09sleads9.84s10.2s
Clean roundsclean rounds out of rounds tried in 24 hours · higher is better10/10 (100%)leads10/10 (100%)9/10 (90%)

Every row covers the last 24 hours. Test kit v6 rows only.

How to read this

Add a third model to the web address with ?c=sonnet. Every number follows the same rules as the home board: current test kit only, right answers only, one try per round. Units are spelled out on How we test.

Matchups worth opening

One click loads the whole card.

Read next

Grok 4.5 pageGPT-5.6 Luna pageThe model boardHow we test