Local vs Cloud: SWE-Bench Verified Leaderboard
The best open-weight LLM on SWE-Bench Verified is Qwen3.6-27B (77.2%), which runs on a 32GB Mac, 19.8 points behind Claude Opus 5 (97.0%). Every score is raw-confirmed against its primary source, and the table shows when it was last confirmed.
Qwen3.6-27B (77.2%) is the top open-weight model on SWE-Bench Verified that runs on a normal 32GB Mac, 19.8 points behind the cloud leader Claude Opus 5 (97.0%).
Scores are third-party (SWE-Bench Verified), raw-confirmed against each model's primary source. ModelFit runs no benchmarks.
Local vs Cloud: How Close Are We?
SWE-Bench VerifiedReal-world software-engineering benchmark, the headline coding metric. Open-weight models you can run at home keep narrowing the gap with the best closed APIs.
Full Leaderboard
Sources confirmed 2026-07-25| # | Model | SWE-Bench Verified | Type | Runs on | Get it | Source | Verified |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5Cloud API | 97.0% | Cloud | Cloud API | API only | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 2 | GPT-5.6-solCloud API | 96.2% | Cloud | Cloud API | API only | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 3 | Claude Fable 5Cloud API | 95.0% | Cloud | Cloud API | API only | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 4 | Kimi K3Cloud API | 93.4% | Cloud | Cloud API | API only | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 5 | GPT-5.6-lunaCloud API | 93.0% | Cloud | Cloud API | API only | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 6 | Claude Opus 4.8Cloud API | 88.6% | Cloud | Cloud API | API only | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 7 | Grok 4.5Cloud API | 86.6% | Cloud | Cloud API | API only | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 8 | GLM-5.2Cloud API | 82.8% | Cloud | Cloud API | API only | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 9 | GPT-5.5Cloud API | 82.6% | Cloud | Cloud API | API only | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 10 | Claude Opus 4.7Cloud API | 82.0% | Cloud | Cloud API | API only | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 11 | Gemini 3.1 ProCloud API | 78.8% | Cloud | Cloud API | API only | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 12 | Kimi K2.7-Code1T / 32B MoE | 78.2% | Local | Multi-node or 512GB-class rig | — | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 13 | DeepSeek V4 Pro1.6T / 49B MoE | 77.4% | Local | Multi-node rig | — | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 14 | Qwen3.6-27B27B | 77.2% | Local | 32GB Mac | $ qwen3.6:27bregistry-verified | Qwen3.6-27B HuggingFace model card | 2026-07-25 |
| 15 | MiniMax M3428B / 23B MoE | 75.0% | Local | 256GB-class workstation | — | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 16 | Qwen3.5-27B27B | 75.0% | Local | 24GB Mac | $ qwen3.5:27bregistry-verified | Qwen3.5-27B HuggingFace model card | 2026-07-25 |
| 17 | Qwen3.6-35B-A3B35B / 3B MoE | 73.4% | Local | 32GB Mac | $ qwen3.6:35b-a3bregistry-verified | Qwen3.6-35B-A3B HuggingFace model card | 2026-07-25 |
| 18 | NVIDIA Nemotron 3 Ultra550B / 55B MoE | 69.0% | Local | 256GB-class multi-node rig | — | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 19 | Mistral Medium 3.5128B dense | 66.4% | Local | 96GB+ Mac | — | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
| 20 | Devstral 2 (2512)Open-weight 24B | 62.8% | Local | 32GB Mac | — | vals.ai SWE-Bench Verified leaderboard | 2026-07-25 |
Third-party SWE-Bench Verified scores. Each % is re-confirmed against the linked primary source and each Ollama tag re-probed against the registry (HTTP 200) on each refresh. Closed-API parameter counts are undisclosed ("Cloud API"). Tokens-per-second figures elsewhere on ModelFit are estimates, not measured benchmarks.
What Should You Run Today?
MacBook 16GB ⭐
Capable coding on a thin laptop
- Gemma 4 E4B + Qwen3.5-9B
- Qwen3.5-4B for coding (88.8% MMLU-Redux)
MacBook 24GB ⭐
The sweet spot
- Qwen3.6-27B (77.2% SWE-Bench)
- Qwen3.6-35B-A3B for agents
- Gemma 4 26B-A4B for multimodal
Mac Studio 128GB+
Frontier-class local
- DeepSeek V4 Pro (77.4% SWE-Bench)
- MiniMax M3 (75.0%, 256GB+)
- Mistral Medium 3.5 (128B dense)
How this leaderboard stays honest
Every score is SWE-Bench Verified (not the harder SWE-Bench Pro) and links to the primary source it was confirmed against, specifically an independent leaderboard or the model's own card. A scheduled routine re-greps each number from that raw source, re-probes each Ollama tag against the registry (HTTP 200), and only swaps a row when a newly confirmed score beats the incumbent. Closed-API parameter counts are vendor-undisclosed, so cloud rows show "Cloud API". ModelFit runs no benchmarks; tokens-per-second figures elsewhere on the site are estimates.
Want the full analysis?
Detailed comparisons, RAM costs on Apple Silicon, and what runs on your Mac.