The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of Transformers

Recently, many datasets have been proposed to test the systematic generalization ability of neural networks. The companion baseline Transformers, typically trained with default hyper-parameters from standard tasks, are shown to fail dramatically. Here we demonstrate that by revisiting model configurations as basic as scaling of embeddings, early stopping, relative positional embedding, and Universal Transformer variants, we can drastically improve the performance of Transformers on systematic generalization. We report improvements on five popular datasets: SCAN, CFQ, PCFG, COGS, and Mathematics dataset. Our models improve accuracy from 50% to 85% on the PCFG productivity split, and from 35% to 81% on COGS. On SCAN, relative positional embedding largely mitigates the EOS decision problem (Newman et al., 2020), yielding 100% accuracy on the length split with a cutoff at 26. Importantly, performance differences between these models are typically invisible on the IID data split. This calls for proper generalization validation sets for developing neural networks that generalize systematically. We publicly release the code to reproduce our results.

Attention Is All YouNeedAttention Is All You NeedMemorize or generalize?Searching for a…Memorize or generalize? Searching for a compositional RNN in a haystackCompositionalGeneralization for…Compositional Generalization for Primitive SubstitutionsCompositionalgeneralization in a dee…Compositional generalization in a deep seq2seq model by separating syntax and semanticsUniversal TransformersUniversal TransformersPermutation EquivariantModels for Compositiona…Permutation Equivariant Models for Compositional Generalization in LanguageCompositionalGeneralization in…Compositional Generalization in Semantic Parsing: Pre-training vs. Specialized ArchitecturesLearning CompositionalRules via Neural Progra…Learning Compositional Rules via Neural Program SynthesisHierarchical PosetDecoding for…Hierarchical Poset Decoding for Compositional Generalization in LanguageUnlocking CompositionalGeneralization in…Unlocking Compositional Generalization in Pre-trained Models Using Intermediate RepresentationsCompositionalGeneralization and…Compositional Generalization and Natural Language Variation: Can a Semantic Parsing Approach Handle Both?Making TransformersSolve Compositional…Making Transformers Solve Compositional TasksSystematicGeneralization with Edg…Systematic Generalization with Edge TransformersLAGr: Labeling AlignedGraphs for Improving…LAGr: Labeling Aligned Graphs for Improving Systematic Generalization in Semantic ParsingMaking TransformersSolve Compositional…Making Transformers Solve Compositional TasksImproving CompositionalGeneralization with…Improving Compositional Generalization with Latent Structure and Data AugmentationLAGr: Label AlignedGraphs for Better…LAGr: Label Aligned Graphs for Better Systematic Generalization in Semantic ParsingEvaluating the Impact ofModel Scale for…Evaluating the Impact of Model Scale for Compositional Generalization in Semantic ParsingCompositionalgeneralization with a…Compositional generalization with a broad-coverage semantic parserUnobserved LocalStructures Make…Unobserved Local Structures Make Compositional Generalization HardRandomized PositionalEncodings Boost Length…Randomized Positional Encodings Boost Length Generalization of TransformersSparse UniversalTransformerSparse Universal TransformerReCOGS: How IncidentalDetails of a Logical…ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic InterpretationConsistencyRegularization Training…Consistency Regularization Training for Compositional GeneralizationThe Devil is in theDetail: Simple Tricks…The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of Transformers過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。