Side by side
Compare two models
Put two models side by side. Same tests, same units. Each row names its leader.
The short answer · test kit v6
Grok 4.6 writes 1.3× as fast as Opus 5 right now.
- Grok 4.6
- 41.3wall tok/s
- Sonnet 5
- 32.8wall tok/s
- Opus 5
- 32.5wall tok/s
Speed is one question. Check the clean-rounds row too — a fast model that keeps failing is not the faster tool.
Pick the models
Your picks live in the link. Copy it and the matchup travels with it.
Left column
Right column
Third column (optional)
Row by row
The marked cell leads its row. An empty cell means no clean run in this window — not zero.
| Row | Opus 5 | Sonnet 5 | Grok 4.6 |
|---|---|---|---|
| Write speedwall tok/s · every token written over the whole command · higher is better | 32.5wall tok/s | 32.8wall tok/s | 41.3leadswall tok/s |
| Time to first wordhow long the screen stays blank · Codex does not stream, so its cell stays empty · lower is better | 2.44sleads | 2.84s | 4.39s |
| Typical waitstart to finish for the whole command · lower is better | 7.43s | 7.62s | 7.14sleads |
| Clean roundsclean rounds out of rounds tried in 24 hours · higher is better | 26/29 (90%) | 26/29 (90%) | 28/28 (100%)leads |
Every row covers the last 24 hours. Test kit v6 rows only.
How to read this
Add a third model to the web address with ?c=sonnet. Every number follows the same rules as the home board: current test kit only, right answers only, one try per round. Units are spelled out on How we test.
Matchups worth opening
One click loads the whole card.
Best of each lab
Fable 5GPT-5.6 SolGrok 4.6
The top model from Anthropic, OpenAI, and xAI — Fable 5, GPT-5.6 Sol, and Grok 4.6 — on one card.
Fable 5 vs Opus 5
Fable 5Opus 5
Anthropic’s two flagships against each other. Same app, same think setting.
More matchups
The quick ones, the middle weights, and Grok old vs new live with the other pins.