WebArena: A Realistic Web Environment for Building Autonomous Agents

Openstreetmap docker files required to self-host the WebArena benchmark, as described here:https://webarena.dev/https://arxiv.org/abs/2307.13854https://github.com/web-arena-x/webarena/tree/main/environment_docker Copyright to openstreetmaphttps://www.openstreetmap.org/copyright

SQuAD: 100, 000+Questions for Machine…SQuAD: 100, 000+ Questions for Machine Comprehension of TextHotpotQA: A Dataset forDiverse, Explainable…HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question AnsweringKnow What You Don'tKnow: Unanswerable…Know What You Don't Know: Unanswerable Questions for SQuADMapping Natural LanguageInstructions to Mobile…Mapping Natural Language Instructions to Mobile UI Action SequencesWebGPT: Browser-assistedquestion-answering with…WebGPT: Browser-assisted question-answering with human feedbackWebShop: TowardsScalable Real-World Web…WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsReAct: SynergizingReasoning and Acting in…ReAct: Synergizing Reasoning and Acting in Language ModelsTree of Thoughts:Deliberate Problem…Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsLanguage Models canSolve Computer TasksLanguage Models can Solve Computer TasksFrom Pixels to UIActions: Learning to…From Pixels to UI Actions: Learning to Follow Instructions via Graphical User InterfacesA Real-World WebAgentwith Planning, Long…A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisVoyager: An Open-EndedEmbodied Agent with…Voyager: An Open-Ended Embodied Agent with Large Language ModelsWorkArena: How CapableAre Web Agents at…WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?AgentDojo: A DynamicEnvironment to Evaluate…AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM AgentsA Trembling House ofCards? Mapping…A Trembling House of Cards? Mapping Adversarial Attacks against Language AgentsFeedback Loops WithLanguage Models Drive…Feedback Loops With Language Models Drive In-Context Reward HackingTowards General ComputerControl: A Multimodal…Towards General Computer Control: A Multimodal Agent for Red Dead Redemption II as a Case StudySafeArena: Evaluatingthe Safety of Autonomou…SafeArena: Evaluating the Safety of Autonomous Web AgentsThe rise and potentialof large language model…The rise and potential of large language model based agents: a surveyShieldAgent: ShieldingAgents via Verifiable…ShieldAgent: Shielding Agents via Verifiable Safety Policy ReasoningRE-Bench: EvaluatingFrontier AI R&D…RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human ExpertsA Survey on (M)LLM-BasedGUI AgentsA Survey on (M)LLM-Based GUI AgentsFrom System 1 to System2: A Survey of Reasonin…From System 1 to System 2: A Survey of Reasoning Large Language ModelsDeep Research: ASystematic SurveyDeep Research: A Systematic SurveyWebArena: A RealisticWeb Environment for…WebArena: A Realistic Web Environment for Building Autonomous Agents過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。