Skip to content
DurkBench

Codex

GPT-5.6 Luna

Inside Codex we ask for gpt-5.6-luna every hour. Everything below comes from those runs on our lab machine.

Codexgpt-5.6-lunaThink setting highLast round just nowOther models

Write speed, last 24 hours

22.6wall tok/s

Tokens per second across the whole command — you press enter, the app starts, thinks, writes, finishes. Higher is better. This is the middle value of 34 clean rounds out of 34 tried in the last 24 hours.

Time to first word
This app does not stream text, so we cannot time the first word.
Middle wall time
11.4s
How long the whole command took, start to finish, including the app booting. Lower is faster.
Last ping
8.59s
The newest tiny “reply ok” check. It measures wake-up time, not writing speed.

What stands out

Four facts pulled straight from the saved runs.

Board position
#9 of 10
Just ahead: GPT-5.6 Terra at 22.7 wall tok/s. Just behind: GPT-5.6 Sol at 22.6 wall tok/s.
Rounds that worked
34/34
Clean rounds out of rounds tried, in the last 24 hours.
Served as
not echoed
This app does not echo a model name back, so we cannot confirm it.
Round-to-round swing
1.8×
How much the fastest round beat the slowest one. A big swing means a noisy day.

Recent rounds

Newest first. Fails stay in the list with their reason.

WhenResultWall tok/s
just nowokno tool items23.5
1h agookno tool items17.6
2h agookno tool items17.1
3h agookno tool items18.6
4h agookno tool items18.3
5h agookno tool items18.8
6h agookno tool items24.7
7h agookno tool items21.3

All rounds are in Logs.

Read next

Compare with Haiku 4.5

Its closest rival in Claude Code — same window, same units, with the leader marked on every row.

Compare with GPT-5.6 Sol

The pin next to it in Codex. See what the tier jump buys inside one app.

More models

How to read this

Every number here comes from our lab machine running the paid app, one try per hourly round, on test kit v6. Our measurements, not the vendor’s.