Connecting Vision and Language with Localized Narratives

Abstract

This paper proposes Localized Narratives, a new form of multimodal imageannotations connecting vision and language. We ask annotators to describe animage with their voice while simultaneously hovering their mouse over theregion they are describing. Since the voice and the mouse pointer aresynchronized, we can localize every single word in the description. This densevisual grounding takes the form of a mouse trace segment per word and is uniqueto our data. We annotate 628k images with Localized Narratives: the whole COCOdataset and 504k images of the Open Images dataset, which we make publiclyavailable. We provide an extensive analysis of these annotations showing theyare diverse, accurate, and efficient to produce. We also demonstrate theirutility on the application of controlled image captioning.

Quick Read (beta)

loading the full paper ...