← Back to Model Beat
Research·Jul 15·all news from July 15, 2026

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Researchers have introduced Ring-Zero, a framework designed to scale reinforcement learning without human-annotated data to a trillion parameters. By automating chain-of-thought reasoning through verifiable rewards, this method attempts to overcome the computational limitations that have previously restricted such models. This development suggests a shift toward training massive reasoning agents that do not rely on expensive, manual data labeling.

Covered by 1 source

  • AarXiv CS.AIXinyu Tang, Gangqiang Cao, Yurou Liu, Yuliang Zhan, Xiaochong Lan, Yifan Li, Yuchen Yan, Han Peng, Zican Dong, Zhenduo Zhang, Tianshu Wang, Xinyu Kong, Zujie Wen, Wayne Xin Zhao, Zhiqiang Zhang, Jun ZhouJul 15

Related stories

ResearchHow Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to EnforcementJul 16 · 4 sourcesResearchBonsai 27B is a full open reasoning model that fits on an iPhoneJul 14 · 2 sourcesResearchSpaceX in Talks to Sell Computing Power to Pentagon, WSJ SaysJul 17 · 2 sourcesResearchNVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads — a Key Metric for Agentic AIJul 17