Skip to content

How fast is each Claude model in Claude Code right now?

Live response speed of Claude models in real Claude Code sessions: tokens per second while the model is answering, for the main session and for subagents. Free, no account.

Output tokens per second while the model is responding, excluding tool runs and your time.

Loading current observations…

What Tokrate measures in Claude Code

  • Response speed, the headline: output tokens per second while the model is responding, including the wait for the first token and excluding tool runs. Available from companion 0.1.14 for the main session and for subagents.
  • Subagent turns, such as Sonnet subagents started by an Opus orchestrator, are measured separately and never merged with main-session turns.
  • Inference provider, from explicit evidence in each API response: Anthropic, Amazon Bedrock or Google Vertex AI. Bedrock and Vertex AI routes are separate comparisons, and Bedrock also records its region group, such as EU or US.
  • Turn speed, from your message to the final answer including tools and waiting, which reflects your workload as much as the model.
  • Not captured: first-token wait (Claude Code does not report it), streaming speed of visible text, answer quality, and the claude.ai website.

Measure your own Claude Code speed

  1. Download the free companion for Mac, Windows or Linux.
  2. Open it. It finds your local Claude Code sessions automatically and starts measuring new turns.
  3. Choose “Yes, let's contribute” to add your numbers to this board, or “Only for local use” to keep everything on your computer.

Common questions

Does Tokrate use my tokens or cost me anything?

No. Tokrate never calls a model and never sends prompts on your behalf. It reads the session logs that Codex, Claude Code and Grok Build already write on your computer and calculates speed from the token counts and timestamps in them, so it adds nothing to your token usage, API bill or subscription limits.

It does not ask for API keys and needs no Tokrate account. The companions and the public board are free, with no ads.

Is my code, prompt or data shared with anyone?

No prompts, responses, code, file paths, session IDs, names or account details ever leave your computer. Sharing is optional: from version 0.1.11 it stays off until you choose “Yes, let's contribute”, and with “Only for local use” no measurements are sent at all (software update checks are a separate switch).

If you do share, each finished turn sends only performance numbers: coding tool, model, inference provider, recorded reasoning effort, client versions, token counts, durations, a five-minute time bucket and a random sample ID. The server stores your continent, never your country or IP address. Contribution data is not sold and is not used for advertising.

What is response speed?

Response speed is output tokens per second while the model is responding: from the request to the finished answer, including the wait for the first token and the model's thinking tokens, excluding tool runs and your own time. Only responses with at least 200 output tokens count. It is Tokrate's headline number because it compares models fairly even when their tasks need different amounts of tool work.

How does Tokrate decide which model is fastest?

The board ranks models by the median response speed across contributors over the selected window, giving every installation equal weight so one heavy user cannot dominate. Models, inference providers, reasoning effort and coding tools stay separate instead of being blended into one number.

All questions