← Back to Model Beat
Models·Jul 6·all news from July 6, 2026

ollama v0.31.2

## What's Changed * Enabled flash attention on older NVIDIA GPUs (compute capability 6.x) * iGPU can now offload vision models with padding to fit available memory * Fixed structured output for thinking models when thinking is disabled * Hardened GGUF model creation * for Claude Code now disables telemetry by default * Fixed loading models on paths with non-UTF-8 characters * Updated the MLX and llama.cpp engines ## New Contributors * @kevinpark1217 made their first contribution in https://github.com/ollama/ollama/pull/16949 **Full Changelog**: https://github.com/ollama/ollama/compare/v0.31.1...v0.31.2

Covered by 1 source

Related stories

ModelsGPT-5.6 is now the preferred model in Microsoft 365 CopilotJul 7 · 24 sourcesModelsIntroducing GPT-LiveJul 8 · 13 sourcesModelsMistral AI Releases Robotics Model to Support Physical AI PushJul 8 · 43 sourcesModelsDeepseek is designing its own AI chipJul 5 · 61 sources