← Back to Model Beat
Models·5d ago·all news from July 25, 2026

ollama v0.32.4

## What's Changed - Support Laguna on Apple GPUs via the MLX engine - Quantize draft-model output heads at the requested type when creating speculative-decoding drafts. - Fixed Qwen3 MoE decoding for differently-quantized experts, plus faster packed gate/up projection (~4–9% on M5 Max). **Full Changelog**: https://github.com/ollama/ollama/compare/v0.32.3...v0.32.4

Covered by 1 source · 2 articles

Related stories

ModelsIntroducing Claude Opus 5Jul 24 · 15 sourcesModelsOne tampered ChatGPT link could spawn a rogue AI agent that took orders from an attacker every five minutesJul 22 · 34 sourcesModelsMicrosoft and Mistral strike multi-billion-dollar deal to build AI infrastructure across EuropeJul 21 · 96 sourcesModelsChina’s Moonshot to Release Breakthrough AI Model for DownloadJul 25 · 54 sources