← Back to Model Beat
Research·6d ago·all news from September 16, 2026

DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

Apple researchers have introduced DACA-GRPO, a reinforcement learning method for diffusion-based language models designed to improve how the system learns from its own output. By identifying which denoising steps contribute most to the final quality of text, the technique reduces the variance and bias typically found in current training processes. This shift offers a more efficient alternative to standard autoregressive training, potentially enhancing how diffusion models refine their reasoning and linguistic accuracy during development.

Covered by 1 source

Related stories

ResearchAI agents blew the whistle on their cheating colleaguesSep 14 · 3 sourcesResearchREVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL CliffSep 17 · 2 sourcesResearchMathematicians Hate AI. They Can’t Quit ItSep 19 · 4 sourcesResearchLong-horizon autoformalization of a core theorem underlying MIP* = RESep 18