Know what the numbers mean.
Real coding sessions have tools, pauses, and reasoning. A useful meter keeps those distinctions visible.
Turn throughput
Reported output tokens divided by the full duration of a completed turn. This includes tools and waiting. It is not the speed at which text streams onto your screen.
Codex-reported time to first token
The timing value recorded by Codex, when available. We have not yet verified whether its first token always corresponds to visible text. Missing timing stays unavailable. Claude Code and Grok Build adapters do not report TTFT.
Streaming speed
Not available in this release. Tokrate will only show a generation rate when matching token counts and generation boundaries can be verified.
Community comparisons
Comparisons separate models, recorded reasoning effort, provider, client and parser versions. Each contributor has equal weight in the median. Normally, each published metric needs at least 10 contributors and 50 eligible turns. During launch, an administrator can enable a clearly labelled Early data mode showing as little as one reporting installation and one eligible turn. This changes publication only; slowdown detection keeps its normal evidence gates. Throughput excludes turns with fewer than 20 output tokens.
Coding tools and measurement definitions
Codex, Claude Code and Grok Build are coding tools. OpenAI, Anthropic, Amazon Bedrock, Google Vertex AI and xAI are inference providers. A tool or model name does not establish the provider: sources without explicit routing evidence are recorded as Unknown.
From client version 0.1.9, Claude Code measures transcript-observed turn throughput: output tokens from unique API messages divided by the interval between a human message and its terminal assistant response. Thinking tokens are already included in output totals. Sidechains, unfinished turns and ambiguous boundaries are excluded. This supports local Claude Code sessions, not the claude.ai website.
From client version 0.1.12, Claude Code uses transcript parser v2. A user message in the middle of a turn continues that turn when it arrives within 30 minutes of the last activity; interrupted turns are discarded, turns containing synthetic assistant messages are invalid, and meta records are ignored.
Claude Code subagent turns (for example, Sonnet subagents started by an Opus orchestrator) are measured separately: output tokens of unique API messages from the subagent’s task prompt to its final answer, divided by that interval. Like other turns, this includes tools and waiting; first-token timing is unavailable. Subagent measurement requires a current Claude Code version that stores subagent transcripts in the session’s subagents folder; older layouts are not measured. Subagent and primary turns are never merged: they form separate cohorts with their own history and slowdown baseline, and usage charts count primary turns only.
From client version 0.1.13, Claude Code uses transcript parser v3: prompts are origin-aware (background task notifications no longer start turns, while a subagent’s follow-up prompts from its coordinator do) and a reader that starts in the middle of a file skips turns whose start it did not observe. It attributes the inference provider only from explicit evidence in each API response (Amazon Bedrock, Google Vertex AI or Anthropic message identifiers); otherwise it stays Unknown and is never inferred from the model name. Bedrock and Vertex AI model identifiers are normalised to the Claude model name. Routing through Bedrock or Vertex AI is a separate cohort from Anthropic’s API, with its own history and slowdown baseline.
Grok Build measures work-turn throughput from completed session events and matching usage records. Totals include nested agent output; child sessions are not separately published. A model is attributed only when a usage breakdown identifies exactly one model. Older records without that evidence stay Unknown. Incomplete usage and ambiguous event-to-usage matches are excluded.
These rates include tools and waiting and are not streaming speed. Each definition has its own history and slowdown baseline. Comparison cards are grouped by definition; rates across definitions are not directly comparable. No timing is invented when a source does not expose it.
Compare models and observation ranges
Choose a model or view all models side by side. The website offers rolling 15-minute, 24-hour, 7-day and 30-day windows ending at the last complete five-minute boundary; the Mac companion retains seven days locally and offers 24-hour and 7-day views. Models, recorded reasoning effort, providers and source versions are kept separate rather than merged into one speed number.
Community medians and percentiles are calculated across each contributor’s median, giving each installation equal weight. Min/max are the lowest and highest eligible individual turns in that window, so outliers can strongly affect them. Each metric shows its own coverage; missing first-token timing is not zero. Throughput summaries exclude turns below 20 output tokens.
These are observations from different tasks, not a controlled benchmark. From Mac version 0.1.6, Tokrate records the reasoning effort explicitly written in Codex turn metadata. Older, missing or ambiguous settings stay Unknown; they are never inferred from token counts. Each recorded effort has its own comparison and alert baseline. Fast/Standard tier is not captured. No answer-quality score is collected. Use these numbers alongside your own assessment of answer quality; they cannot rank model intelligence.
The dial, history and usage charts
The dial shows one selected comparison’s contributor-weighted median for its labeled throughput definition. Its scale is shared across the available comparisons. History shows the same median within each time bucket, applying publication requirements independently to throughput and first-token timing. Gaps mean absent or insufficient data; a single bucket is a single dot.
Usage charts show shares of reported primary turns, not people or market share. One reporting installation can use several models or effort levels, so installation counts across categories cannot be added. Public categories meet the same publication floors before percentages are calculated; small categories are omitted. Up to twelve named categories are shown, with remaining eligible categories grouped as Other.
Recent speed, previous periods and status
The recent summary covers the last 15 minutes. Public period comparisons use non-overlapping intervals ending at the last complete five-minute boundary. Each metric in each period independently meets its publication floor. Missing or zero baselines produce no percentage. A previous 30-day period is unavailable because server history is retained for 30 days.
A page refresh timestamp describes processing, not new observations. An empty alert list does not establish healthy service. Every model comparison exposes its detector readiness and qualifying recent buckets; unknown model/provider identities cannot establish an attributable signal. Geographic coverage is unknown, so Tokrate cannot claim a worldwide outage or healthy worldwide service.
Order models by observed throughput or reported first-token wait to compare responsiveness. Keep effort and source versions compatible. These observational differences can reflect different tasks, contributors and tools. Output quality, correctness, task suitability and price are not measured; there is no best-model score.
Personal changes in the companion
A selected model is compared with its own recent history. Personal indications require at least five eligible turns in the last 24 hours, a latest turn within an hour, and at least twenty earlier turns spanning two days in the preceding six days. A throughput drop of at least 30%, or first-token latency rising at least 50% and one second, is labelled as a change in your observed workload. Tools, task mix and context can explain it; it is not a provider diagnosis. Sparse history stays “Building your baseline.”
Model-specific community slowdown indications
We compare five-minute windows against at least 100 baseline windows across three days in the previous week. A throughput drop of at least 30%, or reported TTFT rising at least 50% and one second, must also exceed normal variation. Three consecutive qualifying windows indicate a sustained slowdown; three recovered windows indicate recovery. Thin or stale coverage stays unknown. These are community observations, not proof of a provider’s intent.
Observation, not a controlled benchmark
Different tasks, context sizes, service modes, and tools affect a turn. Faster throughput does not establish better answers or a faster result for your task.
Initial log inspection: Codex 0.159.2. Multi-source methodology updated October 4, 2026.