DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

As Large Language Models (LLMs) become increasingly integrated into secure software development workflows, a critical question remains unanswered: can these models not only detect insecure code but also reliably classify vulnerabilities according to standardized taxonomies? In this work, we conduct a systematic evaluation of three state-of-the-art LLMs - Llama3, Codestral, and Deepseek R1 - using a carefully filtered subset of the Big-Vul dataset annotated with eight representative Common Weakness Enumeration categories. Adopting a closed-world classification setup, we assess each model’s performance in both identifying the presence of vulnerabilities and mapping them to the correct CWE label. Our findings reveal a sharp contrast between high detection rates and markedly poor classification accuracy, with frequent overgeneralization and misclassification. Moreover, we analyze model-specific biases and common failure modes, shedding light on the limitations of current LLMs in performing fine-grained security reasoning.These insights are especially relevant in educational contexts, where LLMs are being adopted as learning aids despite their limitations. A nuanced understanding of their behaviour is essential to prevent the propagation of misconceptions among students. Our results expose key challenges that must be addressed before LLMs can be reliably deployed in security-sensitive environments.

Scaling Laws for NeuralLanguage ModelsScaling Laws for Neural Language ModelsTraining Verifiers toSolve Math Word ProblemsTraining Verifiers to Solve Math Word ProblemsTraining a Helpful andHarmless Assistant with…Training a Helpful and Harmless Assistant with Reinforcement Learning from Human FeedbackGPT-4 Technical ReportGPT-4 Technical ReportLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeek-V3 TechnicalReportDeepSeek-V3 Technical ReportLarge Language Monkeys:Scaling Inference…Large Language Monkeys: Scaling Inference Compute with Repeated SamplingDiagnosingNon-Intermittent…Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper)AlphaZero-LikeTree-Search can Guide…AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and TrainingTraining Language Modelsto Self-Correct via…Training Language Models to Self-Correct via Reinforcement LearningLiveCodeBench: Holisticand Contamination Free…LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeEvolving Alignment viaAsymmetric Self-PlayEvolving Alignment via Asymmetric Self-PlaySafeChain: Safety ofLanguage Models with…SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning CapabilitiesThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningToward large reasoningmodels: A survey of…Toward large reasoning models: A survey of reinforced reasoning with large language modelsCUDA-L1: Improving CUDAOptimization via…CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement LearningSelf-RewardingVision-Language Model…Self-Rewarding Vision-Language Model via Reasoning DecompositionSycophancy underPressure: Evaluating an…Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QATrajectory Balance withAsynchrony: Decoupling…Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-TrainingRank-R1: EnhancingReasoning in LLM-based…Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement LearningSaffron-1: Towards anInference Scaling…Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety AssuranceRethinking the TrustRegion in LLM…Rethinking the Trust Region in LLM Reinforcement LearningReasoning over SemanticIDs Enhances Generative…Reasoning over Semantic IDs Enhances Generative RecommendationDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.