Voyager: An Open-Ended Embodied Agent with Large Language Models

We introduce Voyager, the first LLM-powered embodied lifelong learning agent in Minecraft that continuously explores the world, acquires diverse skills, and makes novel discoveries without human intervention. Voyager consists of three key components: 1) an automatic curriculum that maximizes exploration, 2) an ever-growing skill library of executable code for storing and retrieving complex behaviors, and 3) a new iterative prompting mechanism that incorporates environment feedback, execution errors, and self-verification for program improvement. Voyager interacts with GPT-4 via blackbox queries, which bypasses the need for model parameter fine-tuning. The skills developed by Voyager are temporally extended, interpretable, and compositional, which compounds the agent's abilities rapidly and alleviates catastrophic forgetting. Empirically, Voyager shows strong in-context lifelong learning capability and exhibits exceptional proficiency in playing Minecraft. It obtains 3.3x more unique items, travels 2.3x longer distances, and unlocks key tech tree milestones up to 15.3x faster than prior SOTA. Voyager is able to utilize the learned skill library in a new Minecraft world to solve novel tasks from scratch, while other techniques struggle to generalize. We open-source our full codebase and prompts at https://voyager.minedojo.org/.

Do As I Can, Not As ISay: Grounding Language…Do As I Can, Not As I Say: Grounding Language in Robotic AffordancesMineDojo: BuildingOpen-Ended Embodied…MineDojo: Building Open-Ended Embodied Agents with Internet-Scale KnowledgeInner Monologue:Embodied Reasoning…Inner Monologue: Embodied Reasoning through Planning with Language ModelsVideo PreTraining (VPT):Learning to Act by…Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosDescribe, Explain, Planand Select: Interactive…Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task AgentsReflexion: an autonomousagent with dynamic…Reflexion: an autonomous agent with dynamic memory and self-reflectionCode as Policies:Language Model Programs…Code as Policies: Language Model Programs for Embodied ControlReAct: SynergizingReasoning and Acting in…ReAct: Synergizing Reasoning and Acting in Language ModelsPaLM-E: An EmbodiedMultimodal Language…PaLM-E: An Embodied Multimodal Language ModelPlan4MC: SkillReinforcement Learning…Plan4MC: Skill Reinforcement Learning and Planning for Open-World Minecraft TasksGenerative Agents:Interactive Simulacra o…Generative Agents: Interactive Simulacra of Human BehaviorSPRING: GPT-4Out-performs RL…SPRING: GPT-4 Out-performs RL Algorithms by Studying Papers and ReasoningVoxPoser: Composable 3DValue Maps for Robotic…VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language ModelsPlan4MC: SkillReinforcement Learning…Plan4MC: Skill Reinforcement Learning and Planning for Open-World Minecraft TasksCross-EpisodicCurriculum for…Cross-Episodic Curriculum for Transformer AgentsLLM-Assist: EnhancingClosed-Loop Planning…LLM-Assist: Enhancing Closed-Loop Planning with Language-Based ReasoningEureka: Human-LevelReward Design via Codin…Eureka: Human-Level Reward Design via Coding Large Language ModelsSee and Think: EmbodiedAgent in Virtual…See and Think: Embodied Agent in Virtual EnvironmentLanguage Agents withReinforcement Learning…Language Agents with Reinforcement Learning for Strategic Play in the Werewolf GameV-IRL: Grounding VirtualIntelligence in Real…V-IRL: Grounding Virtual Intelligence in Real LifeAutoRT: EmbodiedFoundation Models for…AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic AgentsMOKA: Open-VocabularyRobotic Manipulation…MOKA: Open-Vocabulary Robotic Manipulation through Mark-Based Visual PromptingDo LLM Agents HaveRegret? A Case Study in…Do LLM Agents Have Regret? A Case Study in Online Learning and GamesSPINE: Online SemanticPlanning for Missions…SPINE: Online Semantic Planning for Missions with Incomplete Natural Language Specifications in Unstructured EnvironmentsVoyager: An Open-EndedEmbodied Agent with…Voyager: An Open-Ended Embodied Agent with Large Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。