← Back to Model Beat
Open Source·Jul 11·all news from July 11, 2026

vllm v0.25.0

# vLLM v0.25.0 Release Notes ## Highlights This release features 558 commits from 232 contributors (64 new)! * **Model Runner V2 is now the default for all dense models** (#44443). Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953). * **PagedAttention has been removed** (#47361). The legacy attention implementation is deleted now that V1/MRv2 backends are the standard path. * **The Transformers modeling backend is now as fast as native vLLM** (#47187), and gained FP8 MoE support (#46820), CUDA graph + embed scaling fixes (#48010), and migration of GPTBigCode/Starcoder2 (#30966) and RoBERTa (#47452). * **New models**: LLaVA-OneVision-2 (#44785), Unlimited OCR (#46564, #47102), MOSS-Transcribe-Diarize (#47729), openai/privacy-filter (#41026), and Hy3 (#47192). GLM-5 / DeepSeek-V3.2 landed in…

Covered by 1 source

Related stories

Open SourceOpenAI, Meta, SpaceXAI Compete for More Cost-Efficient AI ModelsJul 10 · 10 sourcesOpen SourceXi to Debut at China’s Flagship AI Summit as US Rivalry Heats UpJul 11 · 14 sourcesOpen SourceGerman AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and GermanJul 13 · 2 sourcesOpen SourceDatabricks makes Chinese open-source model GLM 5.2 its default coding engine after it matched Opus at lower costJul 9 · 2 sources