← Back to Model Beat
Models·Apr 9·all news from April 9, 2026

transformers v5.5.2 — Patch release: v5.5.2

Small patch dedicated to optimizing gemma4, fixing inference with due to k/v states sharing between layers, as well as conversion mappings for some models that would inconsistently serialize their weight names. It contains the following PRs: - Add MoE to Gemma4 TP plan (#45219) by @sywangyi and @Cyrilvallez - [gemma4] Dissociate kv states sharing from the Cache (#45312) by @Cyrilvallez - [gemma4] Remove all shared weights, and silently skip them during loading (#45336) by @Cyrilvallez - Fix conversion mappings for vlms (#45340) by @Cyrilvallez

Covered by 1 source

Related stories

ModelsDeepSeek V4 Expected to Launch in Late April with Massive Parameter Scale - Gizchina.comApr 11ModelsAlibaba's Qwen AI Leads Global Open-Source Downloads in 2026 | Market Analysis - News and Statistics - IndexBoxApr 11ModelsNew DeepSeek model to test China’s AI ambitions - Taipei TimesApr 11ModelsWaiting for DeepSeek: New model to test China’s AI ambitions - Hong Kong Free Press HKFPApr 11