DMRFNet: Deep Multimodal Reasoning and Fusion for Visual Question Answering and explanation generation | Litlas