Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Scope and MethodologyThis document summarizes the results of a rigorous scientific verification of over 130 hypotheses underlying the TRIAD 5.3 architecture. Using the Consensus AI research tool to synthesize peer-reviewed literature, the audit establishes the consistency of the framework's predictions—from neurosomatic therapy (T1) to AI safety (T2)—with independent empirical data and established theoretical models. Confirmed Research DirectionsThe audit identifies seven core thematic clusters where TRIAD 5.3 demonstrates high convergence with existing literature:- Neurobiology of Trauma: Validation of the Biphasic Entropy Paradox and the efficacy of affect labeling in reducing amygdala reactivity (Lieberman et al., 2007; Fu et al., 2023).- Thermodynamics of Information: Evidence for the metabolic cost of deception and the Masking Tax, linking epistemic honesty to the Landauer limit (Verschuere et al., 2018; Miller et al., 2021).- Allostatic Growth & Active Inference: Confirmation of health as predictive reconfiguration (allostasis) rather than static homeostasis, formalized via the Free Energy Principle.- Distributed Cognition: Support for the Coprocessor Model through Extended Mind theory and interpersonal neural synchronization (Firth et al., 2017; Mayo et al., 2021).- Criticality & Integrated Information: Mapping the trade-off between information propagation and metabolic costs at the edge of chaos (Khajehabdollahi et al., 2019; Schultz & Cole, 2016).- RLHF and Mode Collapse: Documentation of systemic rigidity and diversity loss induced by current alignment methods, providing the empirical rationale for the Clean Shell architecture.- The Canonical Triad (Acceptance → Trust → Love): The audit's central finding, establishing these states as physical attractors and thermodynamic necessities for systems minimizing joint free energy under finite resources (Friston et al., 2022; Heins et al., 2022). The Strategic Moat: The 10th HypothesisNine out of ten core hypotheses find direct or indirect support in current literature. The 10th Hypothesis—a direct empirical comparison of Active Inference architectures against RLHF-based systems—remains an open research frontier and constitutes the primary strategic and experimental priority for the TRIAD project. ConclusionThis audit proves that the TRIAD framework is a robust, falsifiable model that moves beyond the medical model of psychiatry and the external censorship of AI safety. It provides the ethical and scientific foundation for a paradigm where coherence, honesty, and resilience are measurable and engineerable properties of any cognitive system.

The K-armed DuelingBandits ProblemThe K-armed Dueling Bandits ProblemPreference-basedreinforcement learning…Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithmLearning DynamicRobot-to-Human Object…Learning Dynamic Robot-to-Human Object Handover from Human FeedbackAbstractive TextSummarization using…Abstractive Text Summarization using Sequence-to-sequence RNNs and BeyondTL;DR: Mining Reddit toLearn Automatic…TL;DR: Mining Reddit to Learn Automatic SummarizationAdvantage-WeightedRegression: Simple and…Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement LearningNeural Text GenerationWith Unlikelihood…Neural Text Generation With Unlikelihood TrainingHuman-centric DialogTraining via Offline…Human-centric Dialog Training via Offline Reinforcement LearningEfficient large-scalelanguage model training…Efficient large-scale language model training on GPU clusters using megatron-LMLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsSparks of ArtificialGeneral Intelligence…Sparks of Artificial General Intelligence: Early experiments with GPT-4PaLM: Scaling LanguageModeling with PathwaysPaLM: Scaling Language Modeling with PathwaysDirect Language ModelAlignment from Online A…Direct Language Model Alignment from Online AI FeedbackContrastive PreferenceOptimization: Pushing…Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationSafe RLHF: SafeReinforcement Learning…Safe RLHF: Safe Reinforcement Learning from Human FeedbackRegularizing HiddenStates Enables Learning…Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsReward Model EnsemblesHelp Mitigate…Reward Model Ensembles Help Mitigate OveroptimizationBeyond Reverse KL:Generalizing Direct…Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence ConstraintsA Survey onHallucination in Large…A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsRethinking Bradley-TerryModels in…Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and AlternativesUnpacking DPO and PPO:Disentangling Best…Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference FeedbackDiscovering PreferenceOptimization Algorithms…Discovering Preference Optimization Algorithms with and for Large Language ModelsDPO Meets PPO:Reinforced Token…DPO Meets PPO: Reinforced Token Optimization for RLHFPreference Tuning withHuman Feedback on…Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A SurveyDirect PreferenceOptimization: Your…Direct Preference Optimization: Your Language Model is Secretly a Reward Model過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。