← Back to Model Beat
Open Source·4d ago·all news from July 25, 2026

vllm v0.26.0

# vLLM v0.26.0 Release Notes ## Highlights This release features 411 commits from 212 contributors (61 new)! * **New Inkling model family** with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990). * **DeepSeek-V4 performance push** across vendors: a specialized routing kernel (2.94% E2E TPOT, #48660), (1.5–2x kernel, #47463), and redundant repeat/copy removal (1.8% E2E TPOT, #48137), plus ROCm two-stage compressor for HCA prefill (#47718), sparse decode/prefill optimizations (#48519, #48788, #46275), and DSpark speculative decoding on AMD (#47419) and XPU (#47677). * **fp32 for generation models via ** (#48390), extended to the LoRA path (#48525) and given a ROCm fast path (#48688), improving accuracy for generation heads. * **Flexible attention backends**: the attention backend can now be selected per KV-cache group (#48012), and sliding-window support is now an explicit backend…

Covered by 1 source

Related stories

Open SourceOpenAI admits its autonomous AI models also compromised credentials on other platforms during security evalJul 27 · 24 sourcesOpen SourceNew reports reveal the extent of OpenAI's loss of control during the autonomous hack on Hugging FaceJul 24 · 7 sourcesOpen SourceScientific computing in the age of agentic AIJul 28 · 2 sourcesOpen SourceGoogle just had its first negative cash flow quarter due to massive AI spendingJul 22 · 4 sources