Neural Machine Translation by Jointly Learning to Align and Translate

Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neural network that can be jointly tuned to maximize the translation performance. The models proposed recently for neural machine translation often belong to a family of encoder-decoders and consists of an encoder that encodes a source sentence into a fixed-length vector from which a decoder generates a translation. In this paper, we conjecture that the use of a fixed-length vector is a bottleneck in improving the performance of this basic encoder-decoder architecture, and propose to extend this by allowing a model to automatically (soft-)search for parts of a source sentence that are relevant to predicting a target word, without having to form these parts as a hard segment explicitly. With this new approach, we achieve a translation performance comparable to the existing state-of-the-art phrase-based system on the task of English-to-French translation. Furthermore, qualitative analysis reveals that the (soft-)alignments found by the model agree well with our intuition.

Bidirectional recurrentneural networksBidirectional recurrent neural networks10.1162/15324430332253322310.1162/153244303322533223Statistical Phrase-BasedTranslationStatistical Phrase-Based TranslationStatistical MachineTranslationStatistical Machine TranslationADADELTA: An AdaptiveLearning Rate MethodADADELTA: An Adaptive Learning Rate MethodSequence Transductionwith Recurrent Neural…Sequence Transduction with Recurrent Neural NetworksContinuous SpaceTranslation Models for…Continuous Space Translation Models for Phrase-Based Statistical Machine TranslationHybrid speechrecognition with Deep…Hybrid speech recognition with Deep Bidirectional LSTMLOL: An Investigationinto Cybernetic Humor…LOL: An Investigation into Cybernetic Humor, or: Can Machines Laugh?Recurrent ContinuousTranslation ModelsRecurrent Continuous Translation ModelsOn the Properties ofNeural Machine…On the Properties of Neural Machine Translation: Encoder-Decoder ApproachesFast and Robust NeuralNetwork Joint Models fo…Fast and Robust Neural Network Joint Models for Statistical Machine TranslationJoint CTC-Attentionbased End-to-End Speech…Joint CTC-Attention based End-to-End Speech Recognition using Multi-task LearningDeep Learning approachfor sentiment analysis…Deep Learning approach for sentiment analysis of short textsGenerating High-Qualityand Informative…Generating High-Quality and Informative Conversation Responses with Sequence-to-Sequence ModelsHigh-Order AttentionModels for Visual…High-Order Attention Models for Visual Question AnsweringAdvancing ConnectionistTemporal Classification…Advancing Connectionist Temporal Classification With Attention ModelingRecurrent NeuralNetworks for Time Serie…Recurrent Neural Networks for Time Series Forecasting: Current Status and Future DirectionsSynchronous Transformersfor End-to-End Speech…Synchronous Transformers for End-to-End Speech RecognitionActivate or Not:Learning Customized…Activate or Not: Learning Customized ActivationEnvironmental soundclassification using…Environmental sound classification using temporal-frequency attention based convolutional neural networkCodeGen: An Open LargeLanguage Model for Code…CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisExploring the Responsesof Large Language Model…Exploring the Responses of Large Language Models to Beginner Programmers' Help RequestsLeave No Context Behind:Efficient Infinite…Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attentionNeural MachineTranslation by Jointly…Neural Machine Translation by Jointly Learning to Align and Translate過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。