Instruction-Following Evaluation for Large Language Models

One core capability of Large Language Models (LLMs) is to follow natural language instructions. However, the evaluation of such abilities is not standardized: Human evaluations are expensive, slow, and not objectively reproducible, while LLM-based auto-evaluation is potentially biased or limited by the ability of the evaluator LLM. To overcome these issues, we introduce Instruction-Following Eval (IFEval) for large language models. IFEval is a straightforward and easy-to-reproduce evaluation benchmark. It focuses on a set of "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times". We identified 25 types of those verifiable instructions and constructed around 500 prompts, with each prompt containing one or more verifiable instructions. We show evaluation results of two widely available LLMs on the market. Our code and data can be found at https://github.com/google-research/google-research/tree/master/instruction_following_eval

Evaluating LargeLanguage Models Trained…Evaluating Large Language Models Trained on CodeScalingInstruction-Finetuned…Scaling Instruction-Finetuned Language ModelsLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsGPT-4 Technical ReportGPT-4 Technical ReportPaLM 2 Technical ReportPaLM 2 Technical ReportInstruction Tuning withGPT-4Instruction Tuning with GPT-4G-Eval: NLG Evaluationusing GPT-4 with Better…G-Eval: NLG Evaluation using GPT-4 with Better Human AlignmentPaLM: Scaling LanguageModeling with PathwaysPaLM: Scaling Language Modeling with PathwaysLarge Language Modelsare not Fair EvaluatorsLarge Language Models are not Fair EvaluatorsLarge Language Modelsare Not Yet Human-Level…Large Language Models are Not Yet Human-Level Evaluators for Abstractive SummarizationGPTScore: Evaluate asYou DesireGPTScore: Evaluate as You DesireA Survey on Evaluationof Large Language ModelsA Survey on Evaluation of Large Language ModelsFOFO: A Benchmark toEvaluate LLMs'…FOFO: A Benchmark to Evaluate LLMs' Format-Following CapabilityDeepSeek LLM: ScalingOpen-Source Language…DeepSeek LLM: Scaling Open-Source Language Models with LongtermismMulti-IF: BenchmarkingLLMs on Multi-Turn and…Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions FollowingApple IntelligenceFoundation Language…Apple Intelligence Foundation Language ModelsEvaluation andmitigation of the…Evaluation and mitigation of the limitations of large language models in clinical decision-makingCodecLM: AligningLanguage Models with…CodecLM: Aligning Language Models with Tailored Synthetic DataParallel Scaling Law forLanguage ModelsParallel Scaling Law for Language ModelsHow to Evaluate RewardModels for RLHFHow to Evaluate Reward Models for RLHFThe Price of Format:Diversity Collapse in…The Price of Format: Diversity Collapse in LLMsEvaluating Judges asEvaluators: The JETTS…Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling EvaluatorsTraining on the TestTask Confounds…Training on the Test Task Confounds Evaluation and EmergenceMind the Gap! ChoiceIndependence in Using…Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different LanguagesInstruction-FollowingEvaluation for Large…Instruction-Following Evaluation for Large Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。