Inkling Small

Inkling Small publishes its weights, but at 276B parameters (12B active per token) no consumer machine holds the checkpoint. This page covers the access paths that work and the open models that actually run locally.

PARAMETERS
276B (12B active)
FORMAT
API
RUNS LOCALLY
No
BEST FOR
Agentic coding, Multimodal, Reasoning

Why you can't run Inkling Small locally

At 276B parameters (12B active per token), a Q4-class build of Inkling Small would need roughly 166 GB, a 256GB Mac Studio could hold the weights in principle, but no registry-verified local build exists today. That figure is arithmetic, not opinion: 0.6 GB per billion parameters is the standard Q4 rule we apply across the whole catalog. The weights themselves are public (official source linked below); capacity, not licensing, is the wall.

What works instead: the hosted API. For most workloads, the open models below deliver the same job on hardware that fits under a desk.

Thinking Machines Lab open-weight MoE (Jul 2026, Apache 2.0): 276B total, 12B active, routed to 6 of 256 experts plus 2 shared. Text, image and audio input. SWE-bench Verified 80.2% per its model card. No Ollama tag and mainline llama.cpp does not load it yet; on a Mac the working path is MLX (the mlx-community 4-bit build is about 153 GB, a 192GB-class machine). Unsloth GGUFs are published for when llama.cpp support lands.

Hugging Face model card for Inkling Small

What to run locally instead of Inkling Small

Your machine:

Get told when a better model fits your MacBook Air M5 16 GB

New open-weight models land every week. When one beats today's pick on your exact machine, you get one short email: the model, why it is better, and the command to run it.

You'll also get the weekly ModelFit email. By signing up you agree to our Privacy Policy. Unsubscribe from either email anytime, in one click.

Frequently asked questions

Can I run Inkling Small locally?

Not on hardware you can buy. The weights are public, but Inkling Small is a 276B-parameter model (12B active), and a Q4-class build would need roughly 166 GB, a 256GB Mac Studio could hold the weights in principle, but no registry-verified local build exists today. Until the ecosystem ships a smaller official build, the API or a smaller open model is the practical path.

How do I access Inkling Small?

Through the vendor's hosted API. The official source is linked on this page.

What is the best local alternative to Inkling Small?

Qwen3 235B A22B is the strongest local model we track (235B, from 192 GB machines), with Qwen3.5 122B-A10B Instruct close behind. Both run on a single high-memory Mac or GPU, see their pages for exact hardware.

Cite this page

ModelFit: Inkling Small, specs, memory math and hardware verdicts.
https://modelfit.io/models/inkling-small/ (dataset updated 2026-10-03, CC BY 4.0).