GRPO Beyond English is a large-scale study evaluating Group Relative Policy Optimization across multilingual settings. Researchers investigate how native-language reasoning rewards impact pretrained language models compared to traditional English-centric training. The findings reveal that training models to reason in their native language achieves performance comparable to English reasoning, offering key insights for global LLM alignment.
Opening Kapyn…