Agent Learning via Early Experience

A long-term goal of language agents is to learn and improve through their own experience, ultimately outperforming humans in complex, real-world tasks. However, training agents from experience data with reinforcement learning remains difficult in many environments, which either lack verifiable rewards (e.g., websites) or require inefficient long-horizon rollouts (e.g., multi-turn tool use). As a result, most current agents rely on supervised fine-tuning on expert data, which is challenging to scale and generalizes poorly. This limitation stems from the nature of expert demonstrations: they capture only a narrow range of scenarios, and expose the agent to limited environment diversity. We address this limitation with a middle-ground paradigm we call early experience: interaction data generated by the agent's own actions, where the resulting future states serve as supervision without reward signals. Within this paradigm, we study two strategies of using such data: (1) implicit world modeling, which uses collected states to ground the policy in environment dynamics; and (2) self-reflection, where the agent learns from its suboptimal actions to improve reasoning and decision-making. Evaluation across eight diverse environments and multiple model families shows that our approaches consistently improve effectiveness and out-of-domain generalization, highlighting the value of early experience. Moreover, in environments with verifiable rewards, our results provide promising signals that early experience offers a strong foundation for subsequent reinforcement learning, making it a practical bridge between imitation learning and fully experience-driven agents.

Playing Atari with DeepReinforcement LearningPlaying Atari with Deep Reinforcement LearningLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsIs Your LLM Secretly aWorld Model of the…Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web AgentsDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsWebAgent-R1: TrainingWeb Agents via…WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement LearningRAGEN: UnderstandingSelf-Evolution in LLM…RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement LearningAgentOccam: A Simple YetStrong Baseline for…AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web AgentsToolRL: Reward is AllTool Learning NeedsToolRL: Reward is All Tool Learning NeedsSearch-R1: Training LLMsto Reason and Leverage…Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement LearningInternalizing WorldModels via Self-Play…Internalizing World Models via Self-Play Finetuning for Agentic RLLearn-by-interact: AData-Centric Framework…Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic EnvironmentsDyna-Think: SynergizingReasoning, Acting, and…Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI AgentsScaling Agent Learningvia Experience SynthesisScaling Agent Learning via Experience SynthesisFrom Word to World: CanLarge Language Models b…From Word to World: Can Large Language Models be Implicit Text-based World Models?Memory in the Age of AIAgentsMemory in the Age of AI AgentsAgentic Learner withGrow-and-Refine…Agentic Learner with Grow-and-Refine Multimodal Semantic MemoryOn the Interplay ofPre-Training…On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsReinforcement WorldModel Learning for…Reinforcement World Model Learning for LLM-based AgentsExpanding theCapabilities of…Expanding the Capabilities of Reinforcement Learning via Text FeedbackReinforcement Learningvia Self-DistillationReinforcement Learning via Self-DistillationQwen3-Coder-NextTechnical ReportQwen3-Coder-Next Technical ReportImagine-then-Plan: AgentLearning from Adaptive…Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World ModelsDr. MAS: StableReinforcement Learning…Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM SystemsAligning Agentic WorldModels via Knowledgeabl…Aligning Agentic World Models via Knowledgeable Experience LearningAgent Learning via EarlyExperienceAgent Learning via Early ExperienceEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.