Raw data
Logs
Every check our lab has saved: the good ones, the failed ones, and the tokens they cost us. This is the ledger behind every board on the site.
What our lab spent
Tokens the paid apps reported back to us for our own test runs. A token is a small chunk of text, about four letters. Total is new input + cache + output.
Pick a time window
Every count on this page follows the window you pick here. The default is the last 24 hours. 30d is as far back as we serve.
Total tokens · last 5 hours
0
0 samples · — at published API list prices
How to read this
Two dollar columns, and they are not the same. Cost is the figure the app itself sent back with the answer — Claude and Grok send one, Codex sends none. API $ is our own recomputation of the same tokens at the vendor's published API list price, so all three apps sit on one scale. Neither is our bill: we pay flat monthly plans. A blank means not measured, never $0. The same API dollars, rolled up to a whole usage window, are what the quota board prices.
No samples in this window yet
The lab starts its next round within the hour. Widen the window above to reach older samples.
Test runs
0 launches in this window · 0 failed samples. Open a run to see every model it touched and what that cost.
No runs in this window
No runs in this window yet. Our lab starts the next round within the hour. Older samples still show below.
Each row: start time · kind · trigger (cron means the timer, anything else means by hand) · wall time · rec / skip / fail · total tokens. Skipped means it was not that target's turn, or a rate limit stopped the app. Failed is a real error, and we keep it.
Sample log
Every measurement we saved, newest first. Filter by check or by app. Rows from old test kits stay here too — the boards ignore them, but we never delete them. The counts above always cover the full window, not the filter.
The status words
- ok
- The app answered and the answer passed our grader.
- rate limit
- The vendor refused the request. We stop the rest of that app for the round. Busy plan, not a slow model.
- rate limit warning
- The vendor served the request but flagged that the plan's usage window is getting full. The run worked, but we keep it out of the medians to be safe. It never stops the round.
- timeout
- The app never finished in the time we allow.
- auth
- The sign-in was not accepted. We skip that app until it is fixed.
- used tools
- The app reached for a tool even though we said not to. The row does not count.
- wrong answer
- Text came back, but it was not what we asked for. Speed with a wrong answer never counts.
- remapped
- We asked for one model and the app served another. We save both names and mark the row, and the boards treat it with care.
- parse
- The app’s own output was not readable, so we could not trust the timings.
- overloaded / credits / failed
- The vendor said it was busy, out of credit, or gave a plain error.
No samples match
No samples match this filter. Clear the chips to see the whole window.
Showing 0 of 0 filtered samples · 0 in the full window. Same rows as JSON: /api/probes · /api/speed · /api/runs
Read next
- How we test
What each check asks, and what every unit means.
- Quota
The same API dollars, rolled up to a whole week.
- Lab health
Is the lab fresh right now, and is it still signed in?
- Developers
The same data as JSON. No key, no login.