Meta-Harness: End-to-End Optimization of Model Harnesses

The performance of large language model (LLM) systems depends not only on model weights, but also on their harness: the code that determines what information to store, retrieve, and present to the model. Yet harnesses are still designed largely by hand, and existing text optimizers are poorly matched to this setting because they compress feedback too aggressively. We introduce Meta-Harness, an outer-loop system that searches over harness code for LLM applications. It uses an agentic proposer that accesses the source code, scores, and execution traces of all prior candidates through a filesystem. On online text classification, Meta-Harness improves over a state-of-the-art context management system by 7.7 points while using 4x fewer context tokens. On retrieval-augmented math reasoning, a single discovered harness improves accuracy on 200 IMO-level problems by 4.7 points on average across five held-out models. On agentic coding, discovered harnesses surpass the best hand-engineered baselines on TerminalBench-2. Together, these results show that richer access to prior experience can enable automated harness engineering.

Character-levelConvolutional Networks…Character-level Convolutional Networks for Text ClassificationDSPy: CompilingDeclarative Language…DSPy: Compiling Declarative Language Model Calls into Self-Improving PipelinesTextGrad: Automatic"Differentiation" via…TextGrad: Automatic "Differentiation" via TextMemEvolve:Meta-Evolution of Agent…MemEvolve: Meta-Evolution of Agent Memory SystemsRecursive LanguageModelsRecursive Language ModelsFeedback Descent:Open-Ended Text…Feedback Descent: Open-Ended Text Optimization via Pairwise ComparisonAgentic ContextEngineering: Evolving…Agentic Context Engineering: Evolving Contexts for Self-Improving Language ModelsAlphaEvolve: A codingagent for scientific an…AlphaEvolve: A coding agent for scientific and algorithmic discoveryGEPA: Reflective PromptEvolution Can Outperfor…GEPA: Reflective Prompt Evolution Can Outperform Reinforcement LearningMeta Context Engineeringvia Agentic Skill…Meta Context Engineering via Agentic Skill EvolutionTerminal-Bench:Benchmarking Agents on…Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesAdaEvolve: Adaptive LLMDriven Zeroth-Order…AdaEvolve: Adaptive LLM Driven Zeroth-Order OptimizationAgentic HarnessEngineering…Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent HarnessesAdapting the Interface,Not the Model: Runtime…Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM AgentsContinual Harness:Online Adaptation for…Continual Harness: Online Adaptation for Self-Improving Foundation AgentsMUSE: A Unified AgenticHarness for MLLMsMUSE: A Unified Agentic Harness for MLLMsOpenClaw-RL: Train AnyAgent Simply by TalkingOpenClaw-RL: Train Any Agent Simply by TalkingShepherd: A RuntimeSubstrate Empowering…Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution TraceVeRO: An EvaluationHarness for Agents to…VeRO: An Evaluation Harness for Agents to Optimize AgentsHarnesses forInference-Time Alignmen…Harnesses for Inference-Time Alignment over Execution TrajectoriesDemoEvolve: OvercomingSparse Feedback in…DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with DemonstrationsHarness Updating Is NotHarness Benefit…Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM AgentsHarness as an Asset:Enforcing Determinism…Harness as an Asset: Enforcing Determinism via the Convergent AI Agent Framework (CAAF)FlashEvolve:Accelerating Agent…FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage OrchestrationMeta-Harness: End-to-EndOptimization of Model…Meta-Harness: End-to-End Optimization of Model Harnesses過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。