← Back to Model Beat
Open Source·Aug 10·all news from August 10, 2026

vllm v0.27.0

# vLLM v0.27.0 Release Notes ## Highlights This release features 561 commits from 242 contributors (64 new)! * **Kimi K3 support** with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656). * **More new models**: Qwen3.5 text-only dense and MoE models (#50210) with EVS video token pruning (#48912), K-EXAONE-2.0-750B-A37B (#50524), VaultGemma via the Transformers modeling backend (#49803), and jina-embeddings-v5-text-nano (#50688). * **PyTorch 2.13.0 upgrade** along with torchvision 0.28.0 and Triton 3.7.1 (#48155) — this is a breaking environment change; XPU (#48677) and CPU (#50412) followed to torch 2.13 as well. * **FlashAttention 4 integration deepens on SM100**: FP8 KV cache support (#42569) and headdim-256 support (#42669), backed by a new JIT warmup…

Covered by 1 source

Related stories

Open SourceChina's Zhipu Says Open-Source GLM-5.3 Outperforms Anthropic's Restricted Model in Vulnerability DetectionAug 14 · 7 sourcesOpen SourceMotif 3: Technical ReportAug 11Open SourceOh Lord, AI Reporters Are Actually Breaking Big NewsAug 12 · 3 sourcesOpen SourceAuditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reportingAug 14