kapynResearch

DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

DACA-GRPO is a new RL method that improves training for diffusion language models. It addresses the lack of temporal credit assignment across denoising steps and the bias of mean-field likelihood estimates in existing GRPO-style trainers. By offering a lightweight, plug-and-play enhancement, it enables more effective policy optimization for diffusion LLMs, which are emerging as alternatives to autoregressive models.

Apple ML Research·Sep 16, 2026

Opening Kapyn…