Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Parallel thinking has emerged as a novel approach for enhancing the reasoning capabilities of large language models (LLMs) by exploring multiple reasoning paths concurrently. However, activating such capabilities through training remains challenging, as existing methods predominantly rely on supervised fine-tuning (SFT) over synthetic data, which encourages teacher-forced imitation rather than exploration and generalization. Different from them, we propose \textbf{Parallel-R1}, the first reinforcement learning (RL) framework that enables parallel thinking behaviors for complex real-world reasoning tasks. Our framework employs a progressive curriculum that explicitly addresses the cold-start problem in training parallel thinking with RL. We first use SFT on prompt-generated trajectories from easier tasks to instill the parallel thinking ability, then transition to RL to explore and generalize this skill on harder problems. Experiments on various math benchmarks, including MATH, AMC23, and AIME, show that Parallel-R1 successfully instills parallel thinking, leading to 8.4% accuracy improvements over the sequential thinking model trained directly on challenging tasks with RL. Further analysis reveals a clear shift in the model's thinking behavior: at an early stage, it uses parallel thinking as an exploration strategy, while in a later stage, it uses the same capability for multi-perspective verification. Most significantly, we validate parallel thinking as a \textbf{mid-training exploration scaffold}, where this temporary exploratory phase unlocks a higher performance ceiling after RL, yielding a 42.9% improvement over the baseline on AIME25. Our model, data, and code will be open-source at https://github.com/zhengkid/Parallel-R1.

Multiverse: YourLanguage Models Secretl…Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge GenerationHogwild! Inference:Parallel LLM Generation…Hogwild! Inference: Parallel LLM Generation via Concurrent AttentionLearning AdaptiveParallel Reasoning with…Learning Adaptive Parallel Reasoning with Language ModelsLearning to Reason viaMixture-of-Thought for…Learning to Reason via Mixture-of-Thought for Logical ReasoningASPD: Unlocking AdaptiveSerial-Parallel Decodin…ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMsGroup Think: MultipleConcurrent Reasoning…Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level GranularitySelf-RewardingVision-Language Model…Self-Rewarding Vision-Language Model via Reasoning DecompositionBeyond the 80/20 Rule:High-Entropy Minority…Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningR-Zero: Self-EvolvingReasoning LLM from Zero…R-Zero: Self-Evolving Reasoning LLM from Zero DataLearning to Keep aPromise: Scaling…Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous DecodingDeep Think withConfidenceDeep Think with ConfidenceR1-RE: Cross-DomainRelationship Extraction…R1-RE: Cross-Domain Relationship Extraction with RLVREvolving Language Modelswithout Labels: Majorit…Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes VariationHogwild! Inference:Parallel LLM Generation…Hogwild! Inference: Parallel LLM Generation via Concurrent AttentionUniRel-R1: RL-tuned LLMReasoning for Knowledge…UniRel-R1: RL-tuned LLM Reasoning for Knowledge Graph Relational Question AnsweringStable and EfficientSingle-Rollout RL for…Stable and Efficient Single-Rollout RL for Multimodal ReasoningVOGUE: GuidingExploration with Visual…VOGUE: Guiding Exploration with Visual Uncertainty Improves Multimodal ReasoningThe Era of AgenticOrganization: Learning…The Era of Agentic Organization: Learning to Organize with Language ModelsOne Token to FoolLLM-as-a-JudgeOne Token to Fool LLM-as-a-JudgeRethinking ThinkingTokens: LLMs as…Rethinking Thinking Tokens: LLMs as Improvement OperatorsRelayLLM: EfficientReasoning via…RelayLLM: Efficient Reasoning via Collaborative DecodingSave the Good Prefix:Precise Error…Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM ReasoningToo Correct to Learn:Reinforcement Learning…Too Correct to Learn: Reinforcement Learning on Saturated Reasoning DataParallel-Probe: TowardsEfficient Parallel…Parallel-Probe: Towards Efficient Parallel Thinking via 2D ProbingParallel-R1: TowardsParallel Thinking via…Parallel-R1: Towards Parallel Thinking via Reinforcement LearningEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.