LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation

  • 2021-10-19 16:43:54
  • Mohammad Abuzar Shaikh, Zhanghexuan Ji, Dana Moukheiber, Yan Shen, Sargur Srihari, Mingchen Gao
  • 0

Abstract

Pre-training visual and textual representations from large-scale image-textpairs is becoming a standard approach for many downstream vision-languagetasks. The transformer-based models learn inter and intra-modal attentionthrough a list of self-supervised learning tasks. This paper proposes LAViTeR,a novel architecture for visual and textual representation learning. The mainmodule, Visual Textual Alignment (VTA) will be assisted by two auxiliary tasks,GAN-based image synthesis and Image Captioning. We also propose a newevaluation metric measuring the similarity between the learnt visual andtextual embedding. The experimental results on two public datasets, CUB andMS-COCO, demonstrate superior visual and textual representation alignment inthe joint feature embedding space

 

Quick Read (beta)

loading the full paper ...