Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal understanding and it is now able to process up to 3 hours of video content. Its unique combination of long context, multimodal and reasoning capabilities can be combined to unlock new agentic workflows. Gemini 2.5 Flash provides excellent reasoning abilities at a fraction of the compute and latency requirements and Gemini 2.0 Flash and Flash-Lite provide high performance at low latency and cost. Taken together, the Gemini 2.X model generation spans the full Pareto frontier of model capability vs cost, allowing users to explore the boundaries of what is possible with complex agentic problem solving.

Distilling the Knowledgein a Neural NetworkDistilling the Knowledge in a Neural NetworkPreventing VerbatimMemorization in Languag…Preventing Verbatim Memorization in Language Models Gives a False Sense of PrivacySwitch Transformers:Scaling to Trillion…Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient SparsityConstitutional AI:Harmlessness from AI…Constitutional AI: Harmlessness from AI FeedbackMADLAD-400: AMultilingual And…MADLAD-400: A Multilingual And Document-Level Large Audited DatasetPaLM: Scaling LanguageModeling with PathwaysPaLM: Scaling Language Modeling with PathwaysPaLM 2 Technical ReportPaLM 2 Technical ReportQuantifying MemorizationAcross Neural Language…Quantifying Memorization Across Neural Language ModelsGemini 1.5: Unlockingmultimodal understandin…Gemini 1.5: Unlocking multimodal understanding across millions of tokens of contextThe Llama 3 Herd ofModelsThe Llama 3 Herd of ModelsGemma: Open Models Basedon Gemini Research and…Gemma: Open Models Based on Gemini Research and TechnologyHumanity's Last ExamHumanity's Last ExamAgenTracer: Who IsInducing Failure in the…AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?Self-RewardingVision-Language Model…Self-Rewarding Vision-Language Model via Reasoning DecompositionSAM Audio: SegmentAnything in AudioSAM Audio: Segment Anything in AudioOptimus-3: TowardsGeneralist Multimodal…Optimus-3: Towards Generalist Multimodal Minecraft Agents with Scalable Task Expertsdots.ocr: MultilingualDocument Layout Parsing…dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language ModelCorrective DiffusionLanguage ModelsCorrective Diffusion Language ModelsScienceBoard: EvaluatingMultimodal Autonomous…ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific WorkflowsHEART:Emotionally-driven…HEART: Emotionally-driven test-time scaling of Language ModelsLLM-FE: AutomatedFeature Engineering for…LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary OptimizersOn-PolicySelf-Distillation for…On-Policy Self-Distillation for Reasoning CompressionDIFFA: Large LanguageDiffusion Models Can…DIFFA: Large Language Diffusion Models Can Listen and UnderstandA.X K1 Technical ReportA.X K1 Technical ReportGemini 2.5: Pushing theFrontier with Advanced…Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.