← Back to Model Beat
Models·Aug 4·all news from August 4, 2026

ollama v0.32.6

## What's Changed - Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically - streaming now matches OpenAI's wire format: only on the first chunk, on its own chunk, and usage in a separate chunk with . - Truncated OpenAI responses now report instead of . - now offers for cloud-only models that publish no default tag, instead of failing. - TUI fixes: pipe-delimited prose no longer renders as a table, Enter accepts the highlighted file completion, and scrolling is no longer laggy. - Experimental image generation has been temporarily removed. Continue using 0.32.5 for image generation support - Updated the MLX and llama.cpp engines. **Full Changelog**: https://github.com/ollama/ollama/compare/v0.32.5...v0.32.6-rc0

Covered by 1 source

Related stories

ModelsAlibaba’s Qwen3.8-Max AI Model Claims Benchmark Scores Rivaling AnthropicAug 3 · 82 sourcesModelsResponding to the next frontier of critical cyber capabilitiesAug 7 · 13 sourcesModelsApple Teams Up with Alibaba to Bring Qwen AI to macOS 26.6-Aug 8 · 46 sourcesModelsDeepSeek’s Plan to Raise Prices Have a Whole Industry WatchingAug 6 · 27 sources