Geometric-Mean Policy Optimization

Group Relative Policy Optimization (GRPO) has significantly enhanced the reasoning capability of large language models by optimizing the arithmetic mean of token-level rewards. Unfortunately, GRPO is observed to suffer from unstable policy updates when facing tokens with outlier importance-weighted rewards, which manifest as extreme importance sampling ratios during training. In this study, we propose Geometric-Mean Policy Optimization (GMPO), with the aim to improve the stability of GRPO through suppressing token reward outliers. Instead of optimizing the arithmetic mean, GMPO maximizes the geometric mean of token-level rewards, which is inherently less sensitive to outliers and maintains a more stable range of importance sampling ratio. GMPO is plug-and-play-simply replacing GRPO's arithmetic mean with the geometric mean of token-level rewards, as the latter is inherently less sensitive to outliers. GMPO is theoretically plausible-analysis reveals that both GMPO and GRPO are weighted forms of the policy gradient while the former enjoys more stable weights, which consequently benefits policy optimization and performance. Experiments on multiple mathematical reasoning benchmarks show that GMPO-7B improves the average Pass@1 of GRPO by up to 4.1%, outperforming many state-of-the-art approaches. Code is available at https://github.com/callsys/GMPO.

SEED-GRPO: SemanticEntropy Enhanced GRPO…SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy OptimizationRePO: Replay-EnhancedPolicy OptimizationRePO: Replay-Enhanced Policy OptimizationGPG: A Simple and StrongReinforcement Learning…GPG: A Simple and Strong Reinforcement Learning Baseline for Model ReasoningUnderstandingR1-Zero-Like Training…Understanding R1-Zero-Like Training: A Critical PerspectiveCPPO: Accelerating theTraining of Group…CPPO: Accelerating the Training of Group Relative Policy Optimization-Based Reasoning ModelsGRPO-LEAD: ADifficulty-Aware…GRPO-LEAD: A Difficulty-Aware Reinforcement Learning Approach for Concise Mathematical Reasoning in Language ModelsRight Question isAlready Half the Answer…Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning IncentivizationBNPO: Beta NormalizationPolicy OptimizationBNPO: Beta Normalization Policy OptimizationDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleOpen-Reasoner-Zero: AnOpen Source Approach to…Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelReasoning withExploration: An Entropy…Reasoning with Exploration: An Entropy PerspectiveNot All Rollouts areUseful: Down-Sampling…Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement LearningDemystifyingReinforcement Learning…Demystifying Reinforcement Learning in Agentic ReasoningRiskPO: Risk-basedPolicy Optimization via…RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-TrainingRandom Policy Valuationis Enough for LLM…Random Policy Valuation is Enough for LLM Reasoning with Verifiable RewardsStabilizing MoEReinforcement Learning…Stabilizing MoE Reinforcement Learning by Aligning Training and Inference RoutersSimpleTIR: End-to-EndReinforcement Learning…SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated ReasoningESPO: Entropy ImportanceSampling Policy…ESPO: Entropy Importance Sampling Policy OptimizationOne Ring to Rule ThemAll: Unifying…One Ring to Rule Them All: Unifying Group-Based RL via Dynamic Power-Mean GeometryHolder PolicyOptimisationHolder Policy OptimisationOnline Causal KalmanFiltering for Stable an…Online Causal Kalman Filtering for Stable and Effective Policy OptimizationHarder Is Better:Boosting Mathematical…Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question ReformulationScaling ReasoningEfficiently via Relaxed…Scaling Reasoning Efficiently via Relaxed On-Policy DistillationDeep Dense Explorationfor LLM Reinforcement…Deep Dense Exploration for LLM Reinforcement Learning via Pivot-Driven ResamplingGeometric-Mean PolicyOptimizationGeometric-Mean Policy OptimizationEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.