Gemini: A Family of Highly Capable Multimodal Models

This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging from complex reasoning tasks to on-device memory-constrained use-cases. Evaluation on a broad range of benchmarks shows that our most-capable Gemini Ultra model advances the state of the art in 30 of 32 of these benchmarks - notably being the first model to achieve human-expert performance on the well-studied exam benchmark MMLU, and improving the state of the art in every one of the 20 multimodal benchmarks we examined. We believe that the new capabilities of the Gemini family in cross-modal reasoning and language understanding will enable a wide variety of use cases. We discuss our approach toward post-training and deploying Gemini models responsibly to users through services including Gemini, Gemini Advanced, Google AI Studio, and Cloud Vertex AI.

Training Verifiers toSolve Math Word ProblemsTraining Verifiers to Solve Math Word ProblemsMeasuring MathematicalProblem Solving With th…Measuring Mathematical Problem Solving With the MATH DatasetChain of ThoughtPrompting Elicits…Chain of Thought Prompting Elicits Reasoning in Large Language ModelsImproving alignment ofdialogue agents via…Improving alignment of dialogue agents via targeted human judgementsTraining a Helpful andHarmless Assistant with…Training a Helpful and Harmless Assistant with Reinforcement Learning from Human FeedbackConstitutional AI:Harmlessness from AI…Constitutional AI: Harmlessness from AI FeedbackLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsSelf-ConsistencyImproves Chain of…Self-Consistency Improves Chain of Thought Reasoning in Language ModelsChallenging BIG-BenchTasks and Whether…Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve ThemMathVista: EvaluatingMath Reasoning in Visua…MathVista: Evaluating Math Reasoning in Visual Contexts with GPT-4V, Bard, and Other Large Multimodal ModelsWizardLM: EmpoweringLarge Pre-Trained…WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsMathVista: EvaluatingMath Reasoning in Visua…MathVista: Evaluating Math Reasoning in Visual Contexts with GPT-4V, Bard, and Other Large Multimodal ModelsCollectiveConstitutional AI…Collective Constitutional AI: Aligning a Language Model with Public InputSALAD-Bench: AHierarchical and…SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language ModelsUnlock the Correlationbetween Supervised…Unlock the Correlation between Supervised Fine-Tuning and Reinforcement Learning in Training Code Large Language ModelsTowards UnifiedAlignment Between…Towards Unified Alignment Between Agents, Humans, and EnvironmentIntactKV: ImprovingLarge Language Model…IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens IntactCan We Trust LargeLanguage Models…Can We Trust Large Language Models Generated Code? A Framework for In-Context Learning, Security Patterns, and Code Evaluations Across Diverse LLMsBoldly Going Where NoBenchmark Has Gone…Boldly Going Where No Benchmark Has Gone Before: Exposing Bias and Shortcomings in Code Generation EvaluationMono-InternVL: Pushingthe Boundaries of…Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-trainingLLaVA-PruMerge: AdaptiveToken Reduction for…LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsVarying Shades of Wrong:Aligning LLMs with Wron…Varying Shades of Wrong: Aligning LLMs with Wrong Answers OnlyPrism: EfficientTest-Time Scaling via…Prism: Efficient Test-Time Scaling via Hierarchical Search and Self-Verification for Discrete Diffusion Language ModelsGemini: A Family ofHighly Capable…Gemini: A Family of Highly Capable Multimodal ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.