DACA-GRPO is a new RL method that improves training for diffusion language models. It addresses the lack of temporal credit assignment across denoising steps and the bias of mean-field likelihood estimates in existing GRPO-style trainers. By offering a lightweight, plug-and-play enhancement, it enables more effective policy optimization for diffusion LLMs, which are emerging as alternatives to autoregressive models.
Opening Kapyn…