Program Synthesis with Large Language Models

This paper explores the limits of the current generation of large language models for program synthesis in general purpose programming languages. We evaluate a collection of such models (with between 244M and 137B parameters) on two new benchmarks, MBPP and MathQA-Python, in both the few-shot and fine-tuning regimes. Our benchmarks are designed to measure the ability of these models to synthesize short Python programs from natural language descriptions. The Mostly Basic Programming Problems (MBPP) dataset contains 974 programming tasks, designed to be solvable by entry-level programmers. The MathQA-Python dataset, a Python version of the MathQA benchmark, contains 23914 problems that evaluate the ability of the models to synthesize code from more complex text. On both datasets, we find that synthesis performance scales log-linearly with model size. Our largest models, even without finetuning on a code dataset, can synthesize solutions to 59.6 percent of the problems from MBPP using few-shot learning with a well-designed prompt. Fine-tuning on a held-out portion of the dataset improves performance by about 10 percentage points across most model sizes. On the MathQA-Python dataset, the largest fine-tuned model achieves 83.8 percent accuracy. Going further, we study the model's ability to engage in dialog about code, incorporating human feedback to improve its solutions. We find that natural language feedback from a human halves the error rate compared to the model's initial prediction. Additionally, we conduct an error analysis to shed light on where these models fall short and what types of programs are most difficult to generate. Finally, we explore the semantic grounding of these models by fine-tuning them to predict the results of program execution. We find that even our best models are generally unable to predict the output of a program given a specific input.

openalex_id:w2963935794openalex_id:w2963935794Probabilistic model forcode with decision treesProbabilistic model for code with decision treesRobustFill: NeuralProgram Learning under…RobustFill: Neural Program Learning under Noisy I/ODeepBugs: A LearningApproach to Name-based…DeepBugs: A Learning Approach to Name-based Bug DetectionCodeSearchNet Challenge:Evaluating the State of…CodeSearchNet Challenge: Evaluating the State of Semantic Code SearchLanguage Models areFew-Shot LearnersLanguage Models are Few-Shot LearnersLearning to RepresentPrograms with Property…Learning to Represent Programs with Property SignaturesExploring the Limits ofTransfer Learning with…Exploring the Limits of Transfer Learning with a Unified Text-to-Text TransformerGlobal Relational Modelsof Source CodeGlobal Relational Models of Source CodeEvaluating LargeLanguage Models Trained…Evaluating Large Language Models Trained on CodeMeasuring CodingChallenge Competence…Measuring Coding Challenge Competence With APPSProject CodeNet: ALarge-Scale AI for Code…Project CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding TasksUnsolved Problems in MLSafetyUnsolved Problems in ML SafetyOn Distribution Shift inLearning-based Bug…On Distribution Shift in Learning-based Bug DetectorsCodeGen: An Open LargeLanguage Model for Code…CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisHow Important Are GoodMethod Names in Neural…How Important Are Good Method Names in Neural Code Generation? A Model Robustness PerspectiveEureka: Human-LevelReward Design via Codin…Eureka: Human-Level Reward Design via Coding Large Language ModelsMAP-Neo: Highly Capableand Transparent…MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model SeriesBoldly Going Where NoBenchmark Has Gone…Boldly Going Where No Benchmark Has Gone Before: Exposing Bias and Shortcomings in Code Generation EvaluationDELLA-Merging: ReducingInterference in Model…DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based SamplingParallel Scaling Law forLanguage ModelsParallel Scaling Law for Language ModelsCODESIM: Multi-AgentCode Generation and…CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and DebuggingMcEval: MassivelyMultilingual Code…McEval: Massively Multilingual Code EvaluationAgentConductor: TopologyEvolution for…AgentConductor: Topology Evolution for Multi-Agent Competition-Level Code GenerationProgram Synthesis withLarge Language ModelsProgram Synthesis with Large Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.