Authors: Shuyan Zhou , Frank F. Xu , Hao Zhu , Xuhui Zhou , Robert Lo , Abishek Sridhar , Xianyi Cheng , Tianyue Ou , Yonatan Bisk , Daniel Fried , Uri Alon , Graham Neubig - International Conference on Learning Representations, ICLR 2024 cited by 1,673
Openstreetmap docker files required to self-host the WebArena benchmark, as described here:https://webarena.dev/https://arxiv.org/abs/2307.13854https://github.com/web-arena-x/webarena/tree/main/environment_docker Copyright to openstreetmaphttps://www.openstreetmap.org/copyright
✨ Checking sign-in… PDF Cited by View BibTeX Hide BibTeX View BibTeX Cite
SQuAD: 100, 000+ Questions for Machine… SQuAD: 100, 000+ Questions for Machine Comprehension of Text HotpotQA: A Dataset for Diverse, Explainable… HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering Know What You Don't Know: Unanswerable… Know What You Don't Know: Unanswerable Questions for SQuAD Mapping Natural Language Instructions to Mobile… Mapping Natural Language Instructions to Mobile UI Action Sequences WebGPT: Browser-assisted question-answering with… WebGPT: Browser-assisted question-answering with human feedback WebShop: Towards Scalable Real-World Web… WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents ReAct: Synergizing Reasoning and Acting in… ReAct: Synergizing Reasoning and Acting in Language Models Tree of Thoughts: Deliberate Problem… Tree of Thoughts: Deliberate Problem Solving with Large Language Models Language Models can Solve Computer Tasks Language Models can Solve Computer Tasks From Pixels to UI Actions: Learning to… From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces A Real-World WebAgent with Planning, Long… A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis Voyager: An Open-Ended Embodied Agent with… Voyager: An Open-Ended Embodied Agent with Large Language Models WorkArena: How Capable Are Web Agents at… WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? AgentDojo: A Dynamic Environment to Evaluate… AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents A Trembling House of Cards? Mapping… A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents Feedback Loops With Language Models Drive… Feedback Loops With Language Models Drive In-Context Reward Hacking Towards General Computer Control: A Multimodal… Towards General Computer Control: A Multimodal Agent for Red Dead Redemption II as a Case Study SafeArena: Evaluating the Safety of Autonomou… SafeArena: Evaluating the Safety of Autonomous Web Agents The rise and potential of large language model… The rise and potential of large language model based agents: a survey ShieldAgent: Shielding Agents via Verifiable… ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning RE-Bench: Evaluating Frontier AI R&D… RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human Experts A Survey on (M)LLM-Based GUI Agents A Survey on (M)LLM-Based GUI Agents From System 1 to System 2: A Survey of Reasonin… From System 1 to System 2: A Survey of Reasoning Large Language Models Deep Research: A Systematic Survey Deep Research: A Systematic Survey WebArena: A Realistic Web Environment for… WebArena: A Realistic Web Environment for Building Autonomous Agents Earlier references Focus paper Citing papers Older Newer Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.