By Peter · ModelFit · 2026-07-21

Kimi K3 Hits 93.4% on SWE-Bench: Can You Run It Locally?

SWE-Bench Verified ladder: GPT-5.6 Sol 96.2 and Claude Fable 5 95.0 in the cloud, open-weight Kimi K3 at 93.4, and Qwen3.6-27B at 77.2 as the top score that runs on a 24GB Mac

No, and it will not change on July 27 when the weights drop. Kimi K3 scores 93.40% on SWE-Bench Verified, third overall behind GPT-5.6 Sol at 96.20% and Claude Fable 5 at 95.00% (vals.ai, 2026). But the model is a 2.8-trillion-parameter mixture of experts (Moonshot docs, 2026). At 4-bit that is roughly 1.4TB of weights, about 2.7 times the memory of a maxed 512GB Mac Studio. The best coding score you can hold in your hands today is still Qwen3.6-27B: 77.2% on the same benchmark, on a 24GB Mac.

What is Kimi K3?

Kimi K3 is Moonshot AI's new flagship, released through its API on July 16, 2026. Per Moonshot's developer docs, it packs 2.8 trillion total parameters, activates 16 out of 896 experts per token, and carries a 1M-token context window. The same page commits to open weights: "The full model weights will be released by July 27, 2026."

That makes it the largest open-weight release ever announced, if the drop happens on schedule. For scale, Moonshot's previous flagship Kimi K2 was 1 trillion total parameters with 32 billion activated, under a Modified MIT license (Kimi K2 GitHub, 2025). K3 nearly triples the total parameter count. Moonshot has not yet stated K3's license or its absolute active-parameter count; only the 16-of-896 expert ratio is official.

Demand outran supply fast. Moonshot suspended new API subscriptions within days of launch, citing capacity limits (South China Morning Post, 2026).

How good is Kimi K3 at coding?

On the independent vals.ai SWE-Bench Verified leaderboard, Kimi K3 resolves 93.40% of issues, third overall (vals.ai, 2026). That puts an open-weight model 2.8 points behind the best cloud API. A year ago the gap was measured in double digits. Our local vs cloud leaderboard tracks this same benchmark and updates weekly.

Moonshot's own announcement table, as reported by MarkTechPost (2026), claims 67.5 on DeepSWE, 88.3 on Terminal Bench 2.1, 42.0 on SWE Marathon, and 91.2 on BrowseComp. Those are vendor-run numbers on Moonshot's harness, so treat them as the company's claims rather than independent results.

Can you run Kimi K3 at home? The honest math

Here is the arithmetic, using only Moonshot's disclosed figure of 2.8 trillion parameters:

PrecisionBytes per paramWeights alone
FP162~5.6TB
8-bit1~2.8TB
4-bit0.5~1.4TB

And that is before the KV cache, which grows with context. Every consumer tier misses by a mile:

MachineMax memoryShare of K3 at 4-bit
MacBook Pro M5 Max128GB~9%
Mac Studio M3 Ultra512GB~37%
RTX 509032GB VRAM~2%

Expert streaming from SSD does not rescue this either. We ran the math on disk-streaming a 744B MoE and the result was 0.05 to 1.06 tok/s, which is unusable for real work: Run a 744B Model on 25GB RAM? The Honest Math. K3 is nearly four times that size.

So the July 27 weight release matters for research labs, cloud hosts, and the handful of people with multi-terabyte server rigs. It does not change what your Mac or GPU can run.

What can you actually run instead?

The strongest open-weight coding score that fits consumer hardware is Qwen3.6-27B at 77.2% on SWE-Bench Verified (vals.ai, 2026). It runs on a 24GB Mac:

ollama run qwen3.6:27b

The general ladder, by memory:

Your memoryModel classExample
8GB7-8BLlama 3.1 8B
16GB9-14BQwen3.5 9B, Qwen3 14B
24-32GB20-32BGPT-OSS 20B, Qwen3.6-27B
48-64GB30-35B MoE comfortableQwen3.6 35B-A3B
128GB+70-120B classGPT-OSS 120B

For your exact machine, the ModelFit wizard or the hardware calculator will name the pick, or run npx @wecko-ai/modelfit in a terminal. The sizing logic behind those numbers is in our RAM guide.

Why does Kimi K3 still matter for local AI?

Because the open frontier keeps pulling the local frontier up. Frontier capability has historically reached laptop-class hardware with a lag that verified SWE-Bench data puts at roughly 11 months: How Far Behind Cloud AI Is Your Laptop?. Open releases like K3 are the mechanism: distillations, pruned variants, and smaller models trained with the same recipes tend to follow.

Watch for three things after July 27: the actual license (still unstated), whether a technical report discloses active parameters, and how fast community quants appear. When a K3-derived model lands that fits a 24GB or 48GB machine, it will show up in our live leaderboard and the open dataset.

FAQ

Can I run Kimi K3 on a Mac?

No. Kimi K3 has 2.8 trillion parameters (Moonshot docs, 2026), which is roughly 1.4TB of weights at 4-bit before any context memory. The largest Mac you can buy, a 512GB Mac Studio, holds about a third of that. No quantization level changes the answer for consumer hardware.

When do the Kimi K3 weights come out?

Moonshot's developer docs state the full weights "will be released by July 27, 2026." As of July 21 there is no K3 repository in Moonshot's HuggingFace organization, and the model is API-only. The license has not been announced; Kimi K2 used a Modified MIT license.

What is the best local alternative to Kimi K3 for coding?

Qwen3.6-27B, at 77.2% on SWE-Bench Verified per vals.ai, is the highest verified coding score that runs on a 24GB Mac (ollama run qwen3.6:27b). On 48GB+ machines, Qwen3.6 35B-A3B gives faster generation at similar quality. Both are tracked on the ModelFit benchmark page.

Is Kimi K3 open source?

Not yet, and "open weights" is the more accurate term. The weights are promised by July 27, 2026, but Moonshot has not published the license, the training data, or a technical report. Its predecessor K2 shipped under a Modified MIT license, so something similar is plausible but unconfirmed.

Sources

  • Moonshot AI, Kimi K3 quickstart docs: https://platform.kimi.ai/docs/guide/kimi-k3-quickstart
  • vals.ai SWE-Bench Verified leaderboard: https://www.vals.ai/benchmarks/swebench
  • MarkTechPost, Kimi K3 release report: https://www.marktechpost.com/2026/07/16/moonshot-ai-releases-kimi-k3-a-2-8-trillion-parameter-open-moe-model-with-kimi-delta-attention-and-1m-context/
  • South China Morning Post, subscription suspension: https://www.scmp.com/tech/article/3361172/kimi-k3-developer-suspends-new-subscriptions-amid-compute-constraints
  • Moonshot AI, Kimi K2 repository: https://github.com/moonshotai/kimi-k2
What hardware runs this?

Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.

See how this changes your recommendation
Run the wizard

The weekly local-AI refresh

New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.

By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.

Have questions? Reach out on X/Twitter