Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

We present a new question set, text corpus, and baselines assembled to encourage AI research in advanced question answering. Together, these constitute the AI2 Reasoning Challenge (ARC), which requires far more powerful knowledge and reasoning than previous challenges such as SQuAD or SNLI. The ARC question set is partitioned into a Challenge Set and an Easy Set, where the Challenge Set contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurence algorithm. The dataset contains only natural, grade-school science questions (authored for human tests), and is the largest public-domain set of this kind (7,787 questions). We test several baselines on the Challenge Set, including leading neural models from the SQuAD and SNLI tasks, and find that none are able to significantly outperform a random baseline, reflecting the difficult nature of this task. We are also releasing the ARC Corpus, a corpus of 14M science sentences relevant to the task, and implementations of the three neural baseline models tested. Can your model perform better? We pose ARC as a challenge to the community.

MCTest: A ChallengeDataset for the…MCTest: A Challenge Dataset for the Open-Domain Machine Comprehension of TextCombining Retrieval,Statistics, and…Combining Retrieval, Statistics, and Inference to Answer Elementary Science QuestionsMy Computer Is an HonorStudent - but How…My Computer Is an Honor Student - but How Intelligent Is It? Standardized Tests as a Measure of AIQuestion Answering viaInteger Programming ove…Question Answering via Integer Programming over Semi-Structured KnowledgeTowards AI-CompleteQuestion Answering: A…Towards AI-Complete Question Answering: A Set of Prerequisite Toy TasksA Decomposable AttentionModel for Natural…A Decomposable Attention Model for Natural Language InferenceCrowdsourcing MultipleChoice Science QuestionsCrowdsourcing Multiple Choice Science QuestionsTriviaQA: A Large ScaleDistantly Supervised…TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading ComprehensionNewsQA: A MachineComprehension DatasetNewsQA: A Machine Comprehension DatasetAre You Smarter Than aSixth Grader? Textbook…Are You Smarter Than a Sixth Grader? Textbook Question Answering for Multimodal Machine ComprehensionSciTaiL: A TextualEntailment Dataset from…SciTaiL: A Textual Entailment Dataset from Science Question AnsweringConstructing Datasetsfor Multi-hop Reading…Constructing Datasets for Multi-hop Reading Comprehension Across DocumentsKG^2: Learning to ReasonScience Exam Questions…KG^2: Learning to Reason Science Exam Questions with Contextual Knowledge Graph EmbeddingsWhat Disease does thisPatient Have? A…What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical ExamsEnglish Machine ReadingComprehension Datasets…English Machine Reading Comprehension Datasets: A SurveyGLaM: Efficient Scalingof Language Models with…GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsRWKV: Reinventing RNNsfor the Transformer EraRWKV: Reinventing RNNs for the Transformer EraEQ-Bench: An EmotionalIntelligence Benchmark…EQ-Bench: An Emotional Intelligence Benchmark for Large Language ModelsMAP-Neo: Highly Capableand Transparent…MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model SeriesNemotron-4 15B TechnicalReportNemotron-4 15B Technical ReportSeerAttention: LearningIntrinsic Sparse…SeerAttention: Learning Intrinsic Sparse Attention in Your LLMsSALAD-Bench: AHierarchical and…SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language ModelsSearch-R1: Training LLMsto Reason and Leverage…Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement LearningMMAU: A MassiveMulti-Task Audio…MMAU: A Massive Multi-Task Audio Understanding and Reasoning BenchmarkThink you have SolvedQuestion Answering? Try…Think you have Solved Question Answering? Try ARC, the AI2 Reasoning ChallengeEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.