What changed

Ollama’s v0.40.0 release candidate says supported model architectures on Apple Silicon will run through the MLX runtime by default. The release notes show the simple pull-and-run flow: `ollama pull qwen3.8`, then `ollama run qwen3.8`.

MLX is Apple’s machine-learning framework for Apple Silicon. Routing a supported model to it by default means users do not need to choose a separate runtime for those architectures.

Why local users may care

The improvement is in the path from download to inference: a compatible model can use the hardware-specific runtime automatically, while the same Ollama interface remains in front. That matters to people testing local models who would rather compare models than troubleshoot acceleration settings.

The release notes say additional models will be tested and enabled during the pre-release. That means support is still being expanded, and the behavior applies only to model architectures Ollama marks as supported.

How to try it

Install the v0.40.0 release candidate, pull a supported model, and run it through Ollama. Check the release notes and model support before assuming every architecture is covered.

Because this is a release candidate, it is aimed at early testing. Users who need a stable production setup should stay on their normal version until the project publishes a stable release.

The broader shift

Local model tools are increasingly hiding runtime selection behind a simple command. When the tool detects a compatible accelerator and chooses the matching backend itself, experimenting with models starts to feel more like using an app and less like configuring a stack.

The next useful signal will be which models Ollama enables and how performance compares on the same device; this release announcement gives no benchmark figures.

Explore the original source ↗

Source published 2026-09-25. Coverage is based on the maker’s announcement and demonstration.