GPT-4o mini
GPT-4o mini is a closed, API-only model of undisclosed size. No weights exist to download, so this page covers what it does, how to reach it, and what to run locally instead.
Why you can't run GPT-4o mini locally
The vendor has not published a parameter count, which means no honest VRAM figure exists — every sizing page that quotes one is guessing. No weights have been released at all (official source linked below), so self-hosting is not a question of hardware budget.
What works instead: the hosted API. For most workloads, the open models below deliver the same job on hardware that fits under a desk.
Cloud/API only compact version of GPT-4o with faster responses. Parameter count undisclosed (cloud API).
What to run locally instead of GPT-4o mini
More OpenAI models
Frequently asked questions
Can I run GPT-4o mini locally?
No — the vendor has not released weights, so there is nothing to run. The hosted API or an open model from the alternatives above is the practical path.
How do I access GPT-4o mini?
Through the vendor's hosted API. The official source is linked on this page.
What is the best local alternative to GPT-4o mini?
Qwen3 235B A22B is the strongest local model we track (235B, from 192 GB machines), with Qwen3.5 122B-A10B Instruct close behind. Both run on a single high-memory Mac or GPU — see their pages for exact hardware.
Cite this page
ModelFit: GPT-4o mini — specs, memory math and hardware verdicts. https://modelfit.io/models/gpt-4o-mini/ (dataset updated 2026-09-03, CC BY 4.0).