Learning to Discover at Test Time

How can we use AI to discover a new state of the art for a scientific problem? Prior work in test-time scaling, such as AlphaEvolve, performs search by prompting a frozen LLM. We perform reinforcement learning at test time, so the LLM can continue to train, but now with experience specific to the test problem. This form of continual learning is quite special, because its goal is to produce one great solution rather than many good ones on average, and to solve this very problem rather than generalize to other problems. Therefore, our learning objective and search subroutine are designed to prioritize the most promising solutions. We call this method Test-Time Training to Discover (TTT-Discover). Following prior work, we focus on problems with continuous rewards. We report results for every problem we attempted, across mathematics, GPU kernel engineering, algorithm design, and biology. TTT-Discover sets the new state of the art in almost all of them: (i) Erdős' minimum overlap problem and an autocorrelation inequality; (ii) a GPUMode kernel competition (up to $2\times$ faster than prior art); (iii) past AtCoder algorithm competitions; and (iv) denoising problem in single-cell analysis. Our solutions are reviewed by experts or the organizers. All our results are achieved with an open model, OpenAI gpt-oss-120b, and can be reproduced with our publicly available code, in contrast to previous best results that required closed frontier models. Our test-time training runs are performed using Tinker, an API by Thinking Machines, with a cost of only a few hundred dollars per problem.

Efficient Estimation ofWord Representations in…Efficient Estimation of Word Representations in Vector SpaceAdam: A Method forStochastic OptimizationAdam: A Method for Stochastic OptimizationBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingDiagnosingNon-Intermittent…Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper)ThetaEvolve: Test-timeLearning on Open…ThetaEvolve: Test-time Learning on Open ProblemsEnd-to-End Test-TimeTraining for Long…End-to-End Test-Time Training for Long ContextShinkaEvolve: TowardsOpen-Ended And…ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program EvolutionMathematical explorationand discovery at scaleMathematical exploration and discovery at scaleTTRL: Test-TimeReinforcement LearningTTRL: Test-Time Reinforcement LearningAlphaEvolve: A codingagent for scientific an…AlphaEvolve: A coding agent for scientific and algorithmic discoveryQwen3 Technical ReportQwen3 Technical ReportAlgorithm Discovery WithLLMs: Evolutionary…Algorithm Discovery With LLMs: Evolutionary Search Meets Reinforcement LearningEvaluation-drivenScaling for Scientific…Evaluation-driven Scaling for Scientific DiscoveryAdaEvolve: Adaptive LLMDriven Zeroth-Order…AdaEvolve: Adaptive LLM Driven Zeroth-Order OptimizationEvoX: Meta-Evolution forAutomated DiscoveryEvoX: Meta-Evolution for Automated DiscoveryMLS-Bench: A Holisticand Rigorous Assessment…MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AICORAL: TowardsAutonomous Multi-Agent…CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended DiscoveryPACEvolve++: ImprovingTest-time Learning for…PACEvolve++: Improving Test-time Learning for Evolutionary Search AgentsReinforcement Learningvia Self-DistillationReinforcement Learning via Self-DistillationK-Search: LLM KernelGeneration via…K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World ModelWhat Do EvolutionaryCoding Agents Evolve?What Do Evolutionary Coding Agents Evolve?DeltaEvolve:Accelerating Scientific…DeltaEvolve: Accelerating Scientific Discovery through Momentum-Driven EvolutionMeta-Harness: End-to-EndOptimization of Model…Meta-Harness: End-to-End Optimization of Model HarnessesTool Verification forTest-Time Reinforcement…Tool Verification for Test-Time Reinforcement LearningLearning to Discover atTest TimeLearning to Discover at Test Time過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。