Diverse Beam Search for Improved Description of Complex Scenes

A single image captures the appearance and position of multiple entities in a scene as well as their complex interactions. As a consequence, natural language grounded in visual contexts tends to be diverse---with utterances differing as focus shifts to specific objects, interactions, or levels of detail. Recently, neural sequence models such as RNNs and LSTMs have been employed to produce visually-grounded language. Beam Search, the standard work-horse for decoding sequences from these models, is an approximate inference algorithm that decodes the top-B sequences in a greedy left-to-right fashion. In practice, the resulting sequences are often minor rewordings of a common utterance, failing to capture the multimodal nature of source images. To address this shortcoming, we propose Diverse Beam Search (DBS), a diversity promoting alternative to BS for approximate inference. DBS produces sequences that are significantly different from each other by incorporating diversity constraints within groups of candidate sequences during decoding; moreover, it achieves this with minimal computational or memory overhead. We demonstrate that our method improves both diversity and quality of decoded sequences over existing techniques on two visually-grounded language generation tasks---image captioning and visual question generation---particularly on complex scenes containing diverse visual content. We also show similar improvements at language-only machine translation tasks, highlighting the generality of our approach.

Bleu: a Method forAutomatic Evaluation of…Bleu: a Method for Automatic Evaluation of Machine TranslationMicrosoft COCO: CommonObjects in ContextMicrosoft COCO: Common Objects in ContextDeep Visual-SemanticAlignments for…Deep Visual-Semantic Alignments for Generating Image DescriptionsCIDEr: Consensus-basedImage Description…CIDEr: Consensus-based Image Description EvaluationSequence to Sequence -Video to TextSequence to Sequence - Video to TextShow and Tell: A NeuralImage Caption GeneratorShow and Tell: A Neural Image Caption GeneratorMutual Information andDiverse Decoding Improv…Mutual Information and Diverse Decoding Improve Neural Machine TranslationDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionSPICE: SemanticPropositional Image…SPICE: Semantic Propositional Image Caption EvaluationA Diversity-PromotingObjective Function for…A Diversity-Promoting Objective Function for Neural Conversation ModelsSequence-to-SequenceLearning as Beam-Search…Sequence-to-Sequence Learning as Beam-Search OptimizationVisual DialogVisual Dialogopenalex_id:w3102877762openalex_id:w3102877762Show, Control and Tell:A Framework for…Show, Control and Tell: A Framework for Generating Controllable and Grounded CaptionsDiverse and ControllableImage Captioning with…Diverse and Controllable Image Captioning with Part-of-Speech GuidanceDiverse Image Captioningwith Context-Object…Diverse Image Captioning with Context-Object Split Latent SpacesComprehensive ImageCaptioning via Scene…Comprehensive Image Captioning via Scene Graph DecompositionTowards Unique andInformative Captioning…Towards Unique and Informative Captioning of ImagesThe Curious Case ofNeural Text DegenerationThe Curious Case of Neural Text DegenerationChallenges in BuildingIntelligent Open-domain…Challenges in Building Intelligent Open-domain Dialog SystemsAttacking ImageCaptioning Towards…Attacking Image Captioning Towards Accuracy-Preserving Target Words RemovalConZIC: ControllableZero-shot Image…ConZIC: Controllable Zero-shot Image Captioning by Sampling-Based PolishingFully-attentiveiterative networks for…Fully-attentive iterative networks for region-based controllable image and video captioningRetracted: Deep LearningApproaches for Image…Retracted: Deep Learning Approaches for Image Captioning: Opportunities, Challenges and Future PotentialDiverse Beam Search forImproved Description of…Diverse Beam Search for Improved Description of Complex Scenes過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。