← Back to Model Beat
Research·Jul 28·all news from July 28, 2026

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Researchers have introduced AngelSpec, a framework designed to improve speculative decoding efficiency for large language models. The method dynamically adapts drafting structures to different real-world workloads, addressing the inconsistency in performance often found when using standard autoregressive multi-token prediction.

Covered by 2 sources · 4 articles

  • MMarkTechPostMichal SutterJul 30
  • AarXiv CS.AIDengke HanJul 28
  • AarXiv CS.AIZheng Wang, Zhifan Ye, Qi Cheng, Yonggan Fu, Ziyan Wang, Feng Zhu, Haozhe Zhao, Jan Kautz, Pavlo Molchanov, Humphrey Shi, Minjia ZhangJul 28
  • AarXiv CS.AIHong Liu, Rui Cen, Junhan Shi, Guangshuo Qin, Jiebin Zhang, Tianyu Liu, Runzhi Fan, Guoliang Zhao, Ruobing Xie, Kai Zhang, Song Liu, Guanghua Yu, Jianchen ZhuJul 29

Related stories

ResearchAdvancing responsible AI across EuropeJul 29 · 24 sourcesResearchAI Investment Boom Faces Reality Check From Markets and RegulatorsJul 29 · 6 sourcesResearchAccelerating scientific discovery with ChatGPT for Academic ResearchersJul 29ResearchMoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action TokenizationJul 28 · 9 sources