ollama v0.34.1
## What's Changed * MLX safetensors no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization. * Improved MLX memory handling on Apple Silicon * Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g. OCR) * is much faster on large model libraries (3.1 s → 294 ms cold in testing), and model capabilities are now reported consistently. * MLX and llama.cpp updates **Full Changelog**: https://github.com/ollama/ollama/compare/v0.34.0...v0.34.1-rc1
Covered by 1 source
- GGitHub Releases↗1d ago