A local runtime choice becomes automatic
Ollama’s v0.40.0 prerelease says architectures supported by MLX will run through the MLX runtime by default on Apple Silicon. Previously, users often had to choose a model variant or runtime deliberately. The change makes the path from pulling a compatible model to running it locally more straightforward.
The release notes also mention support additions for Gemma and Qwen families and MLX versions of several decision models. This is a prerelease, so it should not be read as a stable release announcement or as a guarantee that every model will use MLX.
Why MLX matters on a Mac
MLX is Apple’s machine-learning framework designed for Apple Silicon. A runtime that matches the hardware can improve how local inference uses unified memory and Apple’s compute accelerators, depending on the model and workload. The important user-facing change in Ollama is that a supported architecture can take that route without a separate manual selection.
Local execution can help with offline use, data locality and experimentation without sending prompts to a hosted service. It also puts model size, memory use, thermal limits and generation speed directly in the user’s hands. A default runtime simplifies setup; it does not make a large model fit on every Mac.
A small but useful builder update
For local-AI builders, fewer runtime decisions mean one less packaging issue when testing models. The release can also make app instructions simpler: install Ollama, pull a supported model and let the runtime choose MLX when it can.
The accurate takeaway is narrow and useful: this prerelease changes the default for supported Apple Silicon architectures. Developers should check the release notes and test their specific model before relying on it in a workflow, especially while v0.40.0 remains pre-release.
Source published September 25, 2026; prerelease. Coverage is based on the maker’s announcement and demonstration.
