The practical change
Ollama 0.34.3 makes each model’s available thinking controls and default visible through GET /api/show. The same information appears in the command-line interface with ollama show. A developer can inspect whether a model accepts a simple on/off setting or multiple effort levels before building a workflow around it.
That sounds minor until two runs behave differently because one model thinks by default and another does not. Reasoning controls can change latency and the amount of work a model performs. Making the controls discoverable allows a user or app to choose intentionally rather than infer behavior from slower responses or unexpected output.
What else shipped
The release notes also add support for Nemotron H vision models on Apple Silicon through MLX. That expands the range of visual models Mac users can try in a local Ollama workflow. Ollama says it fixed a Hugging Face model-pull issue and a macOS behavior where closed app windows reopened on activation.
These are release notes for an existing tool, not a new frontier model. The main value is that the local-AI workflow becomes less mysterious. As open models ship with different reasoning and multimodal features, tools need to expose those capabilities clearly instead of forcing users to guess.
Why it matters
Local AI is often sold as a simple download-and-run experience, but model settings differ sharply. If a team wants predictable latency for a desktop assistant, a CLI or API that reveals defaults can prevent quiet changes after a model switch. For experiments, it also makes comparisons fairer.
The quickest test is to inspect a model with ollama show, record its thinking options, then run the same prompt at each supported setting and compare time and answer quality. The release makes that kind of deliberate evaluation easier, while the actual gain still depends on the specific model and hardware.
Source published 2026-09-19. Coverage is based on the maker’s announcement and demonstration.
