A Survey of Large Language Models

Abstract The rapid evolution of large language models (LLMs) has driven a transformative shift in artificial intelligence (AI), reshaping both research paradigms and practical applications. Distinguished from their predecessors by unprecedented scale and advanced capabilities, LLMs necessitate new frameworks for understanding their development, behavior, and societal impact. This survey systematically reviews recent advancements in LLM techniques across four key dimensions: (1) pre-training methodologies, which establish core model capabilities through large-scale self-supervised training, architectural innovations, and data curation strategies; (2) post-training techniques, including supervised fine-tuning and reinforcement learning, which adapt foundational models to downstream tasks and enhance their alignment and safety; (3) utilization strategies, such as in-context learning, prompt engineering, and agentic reasoning, that optimize real-world deployment and enable effective interaction with external environments; and (4) evaluation methods, encompassing benchmarks for key ability dimensions such as core language capabilities, reasoning, and safety, which support comprehensive and reliable assessment of model performance. Additionally, we identify critical research issues, including those concerning theoretical foundations, efficient scaling, alignment, and agentic capability, and highlight the open challenges they present. By synthesizing state-of-the-art insights and emerging trends, this survey aims to provide a systematic and comprehensive framework for understanding the trajectory, current limitations, and future directions of LLM progress.

GPTQ: AccuratePost-Training…GPTQ: Accurate Post-Training Quantization for Generative Pre-trained TransformersLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsSiren's Song in the AIOcean: A Survey on…Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language ModelsRetrieval-AugmentedGeneration for Large…Retrieval-Augmented Generation for Large Language Models: A SurveyPrinciple-DrivenSelf-Alignment of…Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionBaichuan 2: OpenLarge-scale Language…Baichuan 2: Open Large-scale Language ModelsDoReMi: Optimizing DataMixtures Speeds Up…DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsSheared LLaMA:Accelerating Language…Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningQwen2 Technical ReportQwen2 Technical ReportOWQ: Outlier-AwareWeight Quantization for…OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language ModelsEfficient Large LanguageModels: A SurveyEfficient Large Language Models: A SurveyLarge Language ModelAlignment: A SurveyLarge Language Model Alignment: A SurveyData Race DetectionUsing Large Language…Data Race Detection Using Large Language ModelsLarge Language Models inEducation: Vision and…Large Language Models in Education: Vision and OpportunitiesReEvo: Large LanguageModels as…ReEvo: Large Language Models as Hyper-Heuristics with Reflective EvolutionLLM-TOPLA: Efficient LLMEnsemble by Maximising…LLM-TOPLA: Efficient LLM Ensemble by Maximising DiversityInternet of IntelligentThings: A convergence o…Internet of Intelligent Things: A convergence of embedded systems, edge computing and machine learningSearch-R1: Training LLMsto Reason and Leverage…Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement LearningCODESIM: Multi-AgentCode Generation and…CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and DebuggingLightPROF: A LightweightReasoning Framework for…LightPROF: A Lightweight Reasoning Framework for Large Language Model on Knowledge GraphToward large reasoningmodels: A survey of…Toward large reasoning models: A survey of reinforced reasoning with large language modelsLEGO-GraphRAG:Modularizing Graph-base…LEGO-GraphRAG: Modularizing Graph-based Retrieval-Augmented Generation for Design Space ExplorationdLLM-Cache: AcceleratingDiffusion Large Languag…dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive CachingA Survey of LargeLanguage ModelsA Survey of Large Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.