← Back to Model Beat
Models·16h ago·all news from October 5, 2026

llama.cpp v0.6.0

## Overview llama.cpp v0.6.0 introduces the new extended batch API (with ) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash (GLM5-Next) 320B hybrid model, the Clef decision model (text and vision) and MTP speculative decoding for Qwen4Exp, ships a new server API for decision models (laya, julia-1, lev, openjev, kev, nimble), overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline, adds a Metal tensor-API flash attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan, and updates ggml to v0.26.0. ### Highlights - New extended batch API with , supporting mixed token/embedding batches and per-token "state" embeddings for MTP and deepstack models [#24669](https://github.com/ggml-org/llama.cpp/pull/24669) - New models: GLM-5.3-Flash (GLM5-Next), a 320B text+vision hybrid model [#27773](https://github.com/ggml-org/llama.cpp/pull/27773), and the Clef decision model, fully supported with both text and vision [#29831](https://github.com/ggml-org/llama.cpp/pull/29831) [#29969](https://github.com/ggml-org/llama.cpp/pull/29969) - Qwen4Exp: high-quality support is now available, with…

Covered by 1 source

Related stories

ModelsUS Lead in AI Over China Narrows After DeepSeek Gains, BI SaysOct 4 · 7 sourcesModelsClaude Frontier Academy: $100M to train 10,000 engineersOct 2ModelsAleph Alpha releases Kolibri, an open-weight model that makes the case for European AI sovereigntyOct 4 · 2 sourcesModelsChinese AI models parrot state doctrine or refuse to answer on sensitive topicsOct 2 · 6 sources