Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reinforcement learning with verifiable rewards (RLVR) has emerged as the leading approach for enhancing reasoning capabilities in large language models. However, it faces a fundamental compute and memory asymmetry: rollout generation is embarrassingly parallel and memory-light, whereas policy updates are communication-heavy and memory-intensive. To address this, we introduce PODS (Policy Optimization with Down-Sampling), which decouples rollout generation from policy updates by training only on a strategically selected subset of rollouts, maintaining learning quality while dramatically reducing update costs. We propose a principled subset selection criterion, max-variance down-sampling, that maximizes reward diversity, and provide an efficient $O(n\log n)$ implementation. Empirically, Group Relative Policy Optimization (GRPO) with PODS achieves the peak test accuracy of vanilla GRPO at least $\mathbf{1.7\times}$ faster across the different reasoning benchmarks and hardware configurations we tested.

DeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsAct Only When It Pays:Efficient Reinforcement…Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective RolloutsDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleSRPO: A Cross-DomainImplementation of…SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLMWhat's Behind PPO'sCollapse in Long-CoT?…What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the SecretA Minimalist Approach toLLM Reasoning: from…A Minimalist Approach to LLM Reasoning: from Rejection Sampling to ReinforceVAPO: Efficient andReliable Reinforcement…VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning TasksWhat Makes a RewardModel a Good Teacher? A…What Makes a Reward Model a Good Teacher? An Optimization PerspectiveDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningREINFORCE++: A Simpleand Efficient Approach…REINFORCE++: A Simple and Efficient Approach for Aligning Large Language ModelsVinePPO: Unlocking RLPotential For LLM…VinePPO: Unlocking RL Potential For LLM Reasoning Through Refined Credit AssignmentProcess Reinforcementthrough Implicit RewardsProcess Reinforcement through Implicit RewardsAct Only When It Pays:Efficient Reinforcement…Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective RolloutsSEED-GRPO: SemanticEntropy Enhanced GRPO…SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy OptimizationRePO: Replay-EnhancedPolicy OptimizationRePO: Replay-Enhanced Policy OptimizationSPEED-RL: FasterTraining of Reasoning…SPEED-RL: Faster Training of Reasoning Models via Online Curriculum LearningPrompt CurriculumLearning for Efficient…Prompt Curriculum Learning for Efficient LLM Post-TrainingThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningGeometric-Mean PolicyOptimizationGeometric-Mean Policy OptimizationReinforce-Ada: AnAdaptive Sampling…Reinforce-Ada: An Adaptive Sampling Framework for Reinforce-Style LLM TrainingBeyond Correctness:Harmonizing Process and…Beyond Correctness: Harmonizing Process and Outcome Rewards through RL TrainingTrain at Moving Edge:Online-Verified Prompt…Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning ModelSmall GeneralizablePrompt Predictive Model…Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning ModelsDemystifying GroupRelative Policy…Demystifying Group Relative Policy Optimization: Its Policy Gradient is a U-StatisticNot All Rollouts areUseful: Down-Sampling…Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement LearningEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.