Measured Local LLM Speed, Machine by Machine

Every row is a real run, not an estimate: readers measure one fixed Ollama model with the same prompt on their own Mac or GPU, so the numbers compare across machines. Elsewhere on ModelFit, tokens/sec are engine estimates; this page is where they meet reality.

Measure your machine

One command, about a minute. It downloads the 1GB reference model, times one answer with Ollama, sends the result to this page, then removes the model. It sends only the machine type and the measured speed: no name, no email, no files.

$ npx -y @wecko-ai/modelfit bench --submit --cleanup

Needs Node and Ollama. New results appear here after the next site update.

RUNS COUNTED
1
MACHINES MEASURED
1
FASTEST SO FAR
71.3 tok/s

Leaderboard

#MachineMeasured (median)RunsBandwidth ceilingLast run
1Apple M4 · 16GB71.3 tok/s1122 tok/s2026-10-03

Reference model: qwen2.5:1.5b-instruct-q4_K_M (about 1GB), fixed prompt, generation speed as reported by Ollama. The ceiling column is the memory-bandwidth limit for that model: a real run stays under it. Runs from CLI versions before 1.5.2 are excluded because they recorded prompt-processing speed instead of generation speed. Medians become reliable from about 3 runs per machine.

How is this different from the speeds elsewhere on ModelFit?

The comparator, device pages and guides show estimated tokens/sec computed from memory bandwidth and model size, because no one can measure every model on every machine. This page measures one small model everywhere. It does not tell you how fast a 27B model runs on your Mac, but it shows how close real hardware gets to its bandwidth limit, which is how the estimates get checked. The data is open: JSON export (CC BY 4.0). For coding quality scores, see the SWE-Bench leaderboard.