← Back to Model Beat
Research·Jul 28·all news from July 28, 2026

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Researchers have introduced AngelSpec, a framework designed to improve speculative decoding efficiency for large language models. The method dynamically adapts drafting structures to different real-world workloads, addressing the inconsistency in performance often found when using standard autoregressive multi-token prediction.

Covered by 2 sources · 4 articles

  • MMarkTechPost↗Michal SutterJul 30
  • AarXiv CS.AI↗Hong Liu, Rui Cen, Junhan Shi, Guangshuo Qin, Jiebin Zhang, Tianyu Liu, Runzhi Fan, Guoliang Zhao, Ruobing Xie, Kai Zhang, Song Liu, Guanghua Yu, Jianchen ZhuJul 29
  • AarXiv CS.AI↗Dengke HanJul 28
  • AarXiv CS.AI↗Zheng Wang, Zhifan Ye, Qi Cheng, Yonggan Fu, Ziyan Wang, Feng Zhu, Haozhe Zhao, Jan Kautz, Pavlo Molchanov, Humphrey Shi, Minjia ZhangJul 28

Related stories

ResearchAdvancing responsible AI across EuropeJul 29 · 24 sourcesResearchAI Investment Boom Faces Reality Check From Markets and RegulatorsJul 29 · 6 sourcesResearchAccelerating scientific discovery with ChatGPT for Academic ResearchersJul 29ResearchMoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action TokenizationJul 28 · 9 sources