← Back to Model Beat
Open Source·Jun 12·all news from June 12, 2026

vllm v0.23.0

# vLLM v0.23.0 Release Notes Please note that Minimax M3 is not yet supported in this version. Please follow [vLLM recipe](https://recipes.vllm.ai/MiniMaxAI/MiniMax-M3) for usage guides for M3. ## Highlights This release features 408 commits from 200 contributors (63 new)! * **DeepSeek-V4 matures across backends**: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass. Its sparse MLA metadata is now decoupled from DeepSeek-V3.2 (#44699), it gained a TRTLLM-gen attention kernel (#43827), EPLB support for the Mega-MoE (#43339), selective prefix-cache retention for sliding-window KV cache (#43447), and an index-share feature for DSA MTP (#44420). The model was also detached from (#43746, #43891), its attention and RoPE paths were refactored (#44569, #44262, #43926), and an XPU attention decode path was added (#42953). * **Model Runner V2 expands to more dense models**: MRv2 is now selected by default for **Llama and Mistral dense models** (#43458) in addition to Qwen3. It gained…

Covered by 1 source

Related stories

Open SourceKPMG fabricated AI case studies in a report designed to sell clients on AI adoptionJun 13 · 6 sourcesOpen SourceWe’re strengthening our presence in Alabama through new investments and community support.Jun 15Open SourceZuckerberg says Meta made 'mistakes' in AI workforce shiftJun 12 · 3 sourcesOpen SourceLLM-ODDR: A Large Language Model Framework for Joint Order Dispatching and Driver RepositioningJun 12