Qwen3 Technical Report

In this work, we present Qwen3, the latest version of the Qwen model family. Qwen3 comprises a series of large language models (LLMs) designed to advance performance, efficiency, and multilingual capabilities. The Qwen3 series includes models of both dense and Mixture-of-Expert (MoE) architectures, with parameter scales ranging from 0.6 to 235 billion. A key innovation in Qwen3 is the integration of thinking mode (for complex, multi-step reasoning) and non-thinking mode (for rapid, context-driven responses) into a unified framework. This eliminates the need to switch between different models--such as chat-optimized models (e.g., GPT-4o) and dedicated reasoning models (e.g., QwQ-32B)--and enables dynamic mode switching based on user queries or chat templates. Meanwhile, Qwen3 introduces a thinking budget mechanism, allowing users to allocate computational resources adaptively during inference, thereby balancing latency and performance based on task complexity. Moreover, by leveraging the knowledge from the flagship models, we significantly reduce the computational resources required to build smaller-scale models, while ensuring their highly competitive performance. Empirical evaluations demonstrate that Qwen3 achieves state-of-the-art results across diverse benchmarks, including tasks in code generation, mathematical reasoning, agent tasks, etc., competitive against larger MoE models and proprietary models. Compared to its predecessor Qwen2.5, Qwen3 expands multilingual support from 29 to 119 languages and dialects, enhancing global accessibility through improved cross-lingual understanding and generation capabilities. To facilitate reproducibility and community-driven research and development, all Qwen3 models are publicly accessible under Apache 2.0.

Qwen2.5 Technical ReportQwen2.5 Technical ReportDeepSeek-V3 TechnicalReportDeepSeek-V3 Technical ReportPhi-4 Technical ReportPhi-4 Technical ReportCRUXEval: A Benchmarkfor Code Reasoning…CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionDOGE: Domain Reweightingwith Generalization…DOGE: Domain Reweighting with Generalization EstimationMulti-IF: BenchmarkingLLMs on Multi-Turn and…Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions FollowingGemma 3 Technical ReportGemma 3 Technical ReportDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningQwen2.5-VL TechnicalReportQwen2.5-VL Technical ReportSuperGPQA: Scaling LLMEvaluation across 285…SuperGPQA: Scaling LLM Evaluation across 285 Graduate DisciplinesRegMix: Data Mixture asRegression for Language…RegMix: Data Mixture as Regression for Language Model Pre-trainingThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningCUDA-L1: Improving CUDAOptimization via…CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement LearningEvolving Language Modelswithout Labels: Majorit…Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes VariationTritonRL: Training LLMsto Think and Code Trito…TritonRL: Training LLMs to Think and Code Triton Without CheatingHow to Train a Leader:Hierarchical Reasoning…How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMsTrajectory Balance withAsynchrony: Decoupling…Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-TrainingRepurposing SyntheticData for Fine-grained…Repurposing Synthetic Data for Fine-grained Search Agent SupervisionLook Back to ReasonForward: Revisitable…Look Back to Reason Forward: Revisitable Memory for Long-Context LLM AgentsP1-VL: Bridging VisualPerception and…P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics OlympiadsOn-PolicySelf-Distillation for…On-Policy Self-Distillation for Reasoning CompressionLearning Rate Scalingacross LoRA Ranks and…Learning Rate Scaling across LoRA Ranks and Transfer to Full FinetuningNovBench: EvaluatingLarge Language Models o…NovBench: Evaluating Large Language Models on Academic Paper Novelty AssessmentQwen3 Technical ReportQwen3 Technical ReportEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.