Xiaomi MiMo-V2-Flash
Xiaomi MiMo-V2-Flash is a 309B-parameter model you reach through an API, with 15B parameters active per token — too large to self-host, so this page covers what it does, how to access it, and what to run locally instead.
You don't run Xiaomi MiMo-V2-Flash locally
At 309B parameters (15B active per token), a Q4-class build of Xiaomi MiMo-V2-Flash would need roughly 185 GB — a 256GB Mac Studio could hold the weights in principle, but no registry-verified local build exists today (0.6 GB per billion parameters, our standard Q4 rule). Access is through the hosted API.
Dec 16, 2025 Xiaomi release (MIT, XiaomiMiMo/MiMo-V2-Flash on HF). 309B total / 15B active MoE with hybrid attention and Multi-Token Prediction, 256K context. SWE-Bench Verified 73.4 and AIME 2025 94.1 per Xiaomi's model card. No Ollama build, but community GGUFs exist (bartowski, unsloth), so a ~Q4 build fits a 256GB-class machine via llama.cpp.
Strong local alternatives
Frequently asked questions
Can I run Xiaomi MiMo-V2-Flash locally?
Not realistically. Xiaomi MiMo-V2-Flash is a 309B-parameter model (15B active); a Q4-class build would need roughly 185 GB — a 256GB Mac Studio could hold the weights in principle, but no registry-verified local build exists today. The hosted API or a smaller open model is the practical path.
How do I access Xiaomi MiMo-V2-Flash?
Through the vendor-hosted API. See the official source linked on this page.
What is the best local alternative to Xiaomi MiMo-V2-Flash?
Qwen3 235B A22B is the strongest local model we track (235B). It runs on a single high-memory Mac or GPU; see its page for exact hardware.
Cite this page
ModelFit: Xiaomi MiMo-V2-Flash — specs, memory math and hardware verdicts. https://modelfit.io/models/mimo-v2-flash/ (dataset updated 2026-09-03, CC BY 4.0).