By Peter · ModelFit · 2026-09-28

macOS 27 Ships a Free On-Device LLM in Your Terminal: What the fm CLI Does Well and Where the 4,096-Token Limit Bites

An Apple Silicon MacBook on a dark desk with a glowing terminal cursor, lit by cyan-teal rim light, illustrating a preinstalled on-device LLM behind a shell command.

Every Mac running macOS 27 has a language model behind a shell command. It sits at /usr/bin/fm, it answers without an account, an API key or a model download from a hub, and it ships with the operating system rather than with a runtime you install. However, the obvious question is whether that built-in model makes Ollama or MLX unnecessary on Apple Silicon. The short answer: it does not, but it wins one specific job outright, and its 4,096-token ceiling decides which jobs those are.

For this piece we analyzed Apple's own session notes, the framework documentation and both hands-on writeups. We also tested the one claim we could check on our own hardware: on the previous macOS release the command does not exist at all. In other words, every figure below is either Apple's published number or a named third-party measurement, and we measured nothing on the model itself.

Key takeaways
  • The fm CLI is a shell front end to the on-device Foundation Model that also powers Apple Intelligence, with no account and no API key.
  • The whole session budget is shared between your instructions, the prompt and the answer, and it is small.
  • fm serve speaks the OpenAI chat-completions format on loopback, and it silently ignores tool calling.
  • For extraction, tagging and classification it beats installing a runtime. For code, math or reasoning, nothing changes.

What ships in macOS 27

Apple introduced the command in WWDC26 session 334, "Build AI-powered scripts with the fm CLI and Python SDK", and says plainly that it "comes pre-installed with macOS 27" (Apple Developer). Additionally, the same session frames the CLI as a way to test prompts without rebuilding an Xcode project, and to fold a model into shell automation. Specifically, it is aimed at checking a prompt interactively and at wiring a model into a script.

The command surface, verified across a hands-on guide tested on macOS 27.0 build 26A428 (mac.install.guide) and a second hands-on writeup (clews.id.au):

  • fm chat starts an interactive session. Sessions persist in ~/.fm/sessions/ and resume with --continue or --resume <name>.
  • fm respond sends one prompt and streams the answer. Add --no-stream when you redirect output to a file.
  • fm schema object builds a JSON schema, and fm respond --schema constrains the output to it.
  • fm available reports whether the model itself has been downloaded.
  • fm serve starts a local HTTP endpoint speaking the OpenAI chat-completions format.

Nonetheless, three things are worth knowing before you type anything. Importantly, the command requires a machine-wide licence acceptance with sudo fm license, because the notice applies to every user account on the Mac. Until then it does nothing at all, not even help. And there is no fm --version: both hands-on writeups report it returns Error: Unknown option '--version', which leaves the macOS build number from sw_vers as the only version you can pin.

The CLI is macOS 27 only, and macOS 27 runs only on Apple Silicon Macs. In fact, there is no separate download for the command and no way to get it on macOS 26 or earlier. We ran the check on a macOS 26.6.2 machine, and which fm returns nothing at all.

The licence gate, the exit code and the empty variable

The licence wall is the first thing that will break a script, so it deserves its own paragraph. In particular, before acceptance, fm writes its warning to the error stream and exits with status 69. Consequently, a script that captures standard output the usual way gets an empty string and no visible reason. The exit status 69 is reported by two independent sources (mac.install.guide, traictory.com), which is enough to treat it as a real scripting gotcha rather than a one-off.

Meanwhile, the second trap is quieter, and also independently reproduced. Passing the instruction as a positional argument makes fm respond discard anything on standard input:

$ echo 'The secret word is Pineapple.' | fm respond 'What is the secret word?'
The secret word is "apple."

That run exits 0 and answers a prompt with nothing behind it. That said, the fix is --instructions (-i for short), which keeps the piped text. If you write shell pipelines against this tool, put set -e at the top of the script. Then reconcile the output against the input: the mac.install.guide author reports a file-sorting demo that returned eight of ten names and silently dropped two.

The 4,096-token window is Apple's own number

Hands paused over a keyboard beside a face-down printout on a dark desk, evoking a shell script that silently fails behind a licence gate.

Third-party guides report the model at roughly 3 billion parameters. Apple does not publish a parameter count for it, on the framework pages or in the session, so treat 3B as third-party reporting rather than a vendor figure.

The token ceiling is different, however, because Apple states it directly. Its documentation says that "Apple's on-device foundation model has a context window of 4096 tokens per session". In addition, that single budget covers "all prompts, instructions, tool definitions and their input and output, generable type schemas, and all of the model's responses" (Managing the context window). The third-party guides agree on 4,096.

That is a session budget, not an input budget. Importantly, the practical consequence is harsher than the round number suggests. The clews.id.au author measured the wall: a 3,028-token document fit, roughly 3,700 tokens of input worked, 3,950 did not, and the failure message was The session's transcript exceeded the model's context size. That same writeup documents fm count-tokens <file> as the way to check before you try. As a result, one long email plus an instruction can eat most of the window before the model writes a word.

On-device versus Private Cloud Compute

Apple's framework documentation gives the cleanest comparison of what the preinstalled model is and what its larger sibling is not.

CapabilityOn-device modelPrivate Cloud Compute
Preserves privacyYes.Yes.
Works offlineYes.No.
Usage limitsUnlimited.Limit per day.
ReasoningNot supported.Multiple levels.
Context size4K.32K.

Source: Adding server-side intelligence with Private Cloud Compute. Additionally, the same page notes that both models need a Mac that supports Apple Intelligence, and that PCC requires eligibility plus a managed entitlement for developers. Private Cloud Compute is Apple's server-side sibling of the on-device model, while Apple Intelligence is the feature umbrella that owns both.

The routing claim that does not survive contact with a shipping Mac

A compact desktop computer with an ethernet cable looping back into its own chassis, symbolizing a local model endpoint bound to loopback only.

This is the one place where the current coverage contradicts itself, so here is the resolution, source by source.

In contrast, Apple's own session 334 says the CLI can use either the on-device model or the Private Cloud Compute model, and its published code samples include fm respond "..." --model pcc and --image Screenshot.png --model pcc.

Two hands-on reports on macOS 27.0 build 26A428 say the shipped binary rejects it. Both state that the only model the command accepts today is system, and that --model pcc fails with Please provide one of 'system'. mac.install.guide adds, in particular, that PCC "is not in macOS 27.0" and that Apple may add it in a later update.

So which claim should you trust? Apple has documented and demoed PCC as reachable through the CLI, and the 27.0 build in two testers' hands does not expose it. Notably, nothing in either hands-on report describes automatic routing based on request complexity.

An article circulating from iosdev.in describes exactly that automatic routing, complete with fm serve --model unified-3b and --pcc-routing auto. Likewise, it claims the CLI ships with Xcode developer utilities and that fm serve locks about 4 GB of unified memory on the Neural Engine. That piece carries its publisher's own warning at the top: "Speculative Architecture & Preview", which states that its technical details "represent previews and proposals rather than finalized APIs". Its flags do not appear in any first-hand test, one of its claims contradicts Apple's session on where the tool comes from, and it should not be used as a description of the shipped command.

For now, the privacy story is the simple one. Specifically, the on-device model runs locally and the licence terms are about Apple's own terms of use, not about a data path. fm serve binds to loopback only, so nothing off the machine reaches it, and mac.install.guide notes the server does not fetch remote URLs either. The one caveat is a PCC session: PCC is a server model reached over the network. Therefore, if the CLI flag lands in a later update, a request made with it is no longer a request that stays on your desk.

Does it replace Ollama or MLX?

Apple tells developers what the on-device model is for, and the list is narrow: summarization, extraction, classification, rewriting, tagging, and answering a question about a passage you supply. However, the same documentation tells you to avoid code, math and logical reasoning. A hallucination test in both hands-on writeups found the model explaining a Homebrew subcommand that does not exist, calling the current macOS "Sonoma" and then "Monterey", and answering a genuine Java runtime error with "Unable to work with that request".

Importantly, on its advertised strengths it holds up. Both authors report accurate summaries and correct tags, and the structured output path is the strongest part of the tool. Fine-grained structured output goes through guided generation. Guided generation is Apple's term for the structured output path it documents as enforced by constrained decoding rather than by asking nicely.

Speed, as measured by clews.id.au on macOS 27.0, is not a differentiator:

MeasurementTimeMachine and conditions
One-word prompt, cold2.4 smacOS 27.0, model not resident.
Follow-up within seconds0.4 ssame machine, warm.
Same prompt after 6 s idle2.4 sresidency released.
Request after 15-60 s idleabout 2.2 ssame author, read from per-request logs.

We measured nothing on this model ourselves, so none of those timings are modelfit measurements. That said, the useful conclusion from them is that model residency is managed by the system and not by the fm serve process. Consequently, a tight loop is fast and anything with gaps pays roughly two seconds a call.

Therefore, wherever you need a context window you control, a quantization you chose, tool calling, or a model that can write code, the answer is still Ollama or MLX. Our Ollama install guide for macOS covers the single-user path, and the MLX versus Ollama comparison covers which of the two Apple paths fits which workload. In other words, the fm model is not a candidate for either role: you do not pick it, you cannot quantize it, and you cannot extend it.

In contrast, where fm does win is the job that neither runtime makes pleasant: a one-line JSON extractor that already exists on the machine. If you need tags, a classification, or three fields pulled out of a support ticket, fm schema object plus fm respond --schema beats spinning up a 7B model for it, and it costs nothing per token. For example, a shell pipeline that tags incoming mail needs no runtime at all, no weights to download and no memory budget to size.

fm serve: an OpenAI-shaped endpoint with real edges

fm serve is the part most people will try to point existing tooling at. It works, and the edges are documented well enough to plan around them.

  • It listens on 127.0.0.1 by default and reports access loopback-only.
  • The routes are POST /v1/chat/completions, GET /v1/models and GET /health.
  • The model name must be system, or you get a 400.
  • api_key is required by the OpenAI SDK and ignored by the server.
  • Streaming is on by default. An OpenAI SDK call that does not set stream: false receives a stream of chunks as a raw string, and .choices fails.
  • tools are accepted and silently ignored. The request returns 200 and the reply is prose in content, never a tool_calls block. Anything built on function calling will not work here.
  • max_tokens appears to be silently ignored; max_completion_tokens is honoured.
  • The server is not a fast path. Because residency is system-managed, fm serve and a one-shot fm respond perform within about a tenth of a second of each other.

Alternatively, if the CLI and the server are both too narrow, the third path is Apple's Foundation Models SDK for Python: pip install apple_fm_sdk, which requires Python 3.10 or later, Xcode and an Apple Silicon Mac. It mirrors the Swift API and adds what neither the CLI nor the HTTP server offers, tool calling included. It reaches the same on-device model, so the 4,096-token session window follows you there.

Run the server when you need an HTTP surface, when your client expects an OpenAI-compatible API, or when you are calling the model from a language other than shell. However, do not run it expecting the weights to stay hot because a daemon is up.

Can your Mac use it?

Importantly, eligibility is not a memory question here, which is unusual for this site. There is no quantization to fit and no memory figure to clear, because the system owns the model and macOS decides when it is resident. So which Macs actually get it? Every Apple Silicon Mac on the current release, and nothing older.

Machinefm availableWhat it asks forContext you getWhat it cannot do
Any Apple Silicon Mac on macOS 27Yes, preinstalled.7 GB of storage for the Apple Intelligence models.4,096 tokens shared by instructions, prompt and answer.No context beyond 4K, no quant choice, no tool calls over HTTP.
Apple Silicon Mac on macOS 26 or earlierNo.Not applicable.Not applicable.The command does not exist.

Two caveats on that table. The 7 GB figure is Apple's published storage requirement for Apple Intelligence on Mac, cited by several independent guides, and the download is separate from the OS install and happens in the background on Apple's schedule. Apple publishes no model-by-model disk or RAM breakdown, so do not read 7 GB as the size of one model file. Second, the download is not instant and not visible: fm available answers System model available or System model unavailable: modelNotReady, macOS gives no progress bar, and reading /System/Library/AssetsV2/ to watch it returns Operation not permitted.

If you are sizing a machine rather than a runtime, our unified memory cheat sheet is the prerequisite, and the per-machine verdicts have not changed.

Should modelfit list it next to Ollama, MLX and llama.cpp?

We checked, because a preinstalled model is a genuine fourth category on a Mac.

Our model dataset holds 143 entries, each keyed on a downloadable checkpoint with a named quantization, and we built the recommendation logic that reads it. Consequently, the Apple system model has neither a quantization to choose nor a published parameter count, so it cannot become a row in that dataset without inventing a number. It also fails the test our device pages apply: those pages answer "which model fits your memory and your job", and this model does not change what fits. Meanwhile, a 32 GB Mac Mini still gets the same 30B-class recommendation it got last month, because the questions that pick a model, context length, quantization, tool support and coding ability, are exactly the ones the fm model loses on.

So no device page verdict changes. What is worth adding, and what this article argues, is a short note on the pages that recommend a local runtime: for extraction, tagging and classification, the cheapest option on macOS 27 is already installed. Everything else still goes through the stack we already recommend, and if you want the category context for Apple's newer frameworks in this area, our Core AI and Foundation Models breakdown covers it. For a concrete machine verdict alongside all of this, the 32 GB Mac Mini M6 picks are unchanged.

FAQ

Does macOS 27 really ship a language model you can run from the terminal?

Yes, on Apple Silicon. /usr/bin/fm is preinstalled with macOS 27 and drives the on-device Foundation Model that powers Apple Intelligence, with no account and no API key. The command does nothing until someone accepts Apple's machine-wide CLI legal notice with sudo fm license, and the model weights download separately in the background.

What is the fm CLI context limit?

4,096 tokens per session, and it covers your instructions, the prompt, any schema and tool definitions, and the model's answer together. That is Apple's published figure for the on-device foundation model. Additionally, third-party testing found the practical input ceiling closer to 3,700 tokens, with a 3,000-token document plus an instruction fitting comfortably. In other words, size your prompt for a small window rather than for the number on the box.

Can fm serve replace my Ollama or LM Studio endpoint?

It can serve an OpenAI-compatible chat-completions route on loopback, with the model name system and stream: false for non-streaming clients. However, it cannot do tool calling, even though it accepts a tools field and ignores it. For instance, it is a good drop-in endpoint for text transformation and a poor one for anything agentic.

How many parameters does Apple's on-device model have?

Apple publishes no parameter count. Third-party guides put it at roughly 3 billion parameters, and that figure should be read as third-party reporting, not as a vendor specification. What Apple does publish is the 4,096-token context window and the intended task list: summarization, extraction and classification.

Does fm serve route requests to Private Cloud Compute?

Not in macOS 27.0. Specifically, Apple's WWDC26 session documents a --model pcc option and shows it in code samples, and two independent hands-on tests on macOS 27.0 report that the shipped command rejects it and accepts only system. Claims of automatic routing from the local model to PCC above a size threshold come from an article its own publisher labels speculative, and its flags do not appear in any hands-on report.

Every figure in this article comes from a source we link in the text, and we publish no measurement of our own here. How we work: about ModelFit. Corrections welcome: contact.
What hardware runs this?

Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.

See how this changes your recommendation
Run the wizard

The weekly local-AI refresh

New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.

By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.

Have questions? Reach out on X/Twitter