
GSPO: Stable MoE RL Training Without Routing Replay
Qwen's GSPO shifts optimization from token-level to sequence-level, eliminating Routing Replay overhead and enabling stable large-scale RLHF training for Qwen3 models.
via Qwen

Qwen's GSPO shifts optimization from token-level to sequence-level, eliminating Routing Replay overhead and enabling stable large-scale RLHF training for Qwen3 models.