Language Modeling Is Compression

It has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine learning community has focused on training increasingly large and powerful self-supervised (language) models. Since these large language models exhibit impressive predictive capabilities, they are well-positioned to be strong compressors. In this work, we advocate for viewing the prediction problem through the lens of compression and evaluate the compression capabilities of large (foundation) models. We show that large language models are powerful general-purpose predictors and that the compression viewpoint provides novel insights into scaling laws, tokenization, and in-context learning. For example, Chinchilla 70B, while trained primarily on text, compresses ImageNet patches to 43.4% and LibriSpeech samples to 16.4% of their raw size, beating domain-specific compressors like PNG (58.5%) or FLAC (30.3%), respectively. Finally, we show that the prediction-compression equivalence allows us to use any compressor (like gzip) to build a conditional generative model.

Asymmetric numeralsystemsAsymmetric numeral systemsA Survey of ModelCompression and…A Survey of Model Compression and Acceleration for Deep Neural NetworksScaling Laws for NeuralLanguage ModelsScaling Laws for Neural Language ModelsOn the Opportunities andRisks of Foundation…On the Opportunities and Risks of Foundation ModelsScaling Language Models:Methods, Analysis &…Scaling Language Models: Methods, Analysis & Insights from Training GopherTraining Compute-OptimalLarge Language ModelsTraining Compute-Optimal Large Language ModelsLLMZip: Lossless TextCompression using Large…LLMZip: Lossless Text Compression using Large Language ModelsLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsSparks of ArtificialGeneral Intelligence…Sparks of Artificial General Intelligence: Early experiments with GPT-4Large Language Models asGeneral Pattern MachinesLarge Language Models as General Pattern MachinesIn-context Autoencoderfor Context Compression…In-context Autoencoder for Context Compression in a Large Language ModelLarge Language ModelsAre Zero-Shot Time…Large Language Models Are Zero-Shot Time Series ForecastersIn-context Autoencoderfor Context Compression…In-context Autoencoder for Context Compression in a Large Language ModelTraining Compute-OptimalProtein Language ModelsTraining Compute-Optimal Protein Language ModelsDiff-eRank: A NovelRank-Based Metric for…Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language ModelsA Survey onHallucination in Large…A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsHallucination ofMultimodal Large…Hallucination of Multimodal Large Language Models: A SurveyA Comprehensive Study ofKnowledge Editing for…A Comprehensive Study of Knowledge Editing for Large Language ModelsFine-Tuned LanguageModels Generate Stable…Fine-Tuned Language Models Generate Stable Inorganic Materials as TextScaling Synthetic DataCreation with…Scaling Synthetic Data Creation with 1,000,000,000 PersonasThe Information of LargeLanguage Model GeometryThe Information of Large Language Model GeometryLLMs may DominateInformation Access…LLMs may Dominate Information Access: Neural Retrievers are Biased Towards LLM-Generated TextsIn-Context LearningStrategies Emerge…In-Context Learning Strategies Emerge RationallyLanguage Modeling IsCompressionLanguage Modeling Is Compression過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。