ISAAQ -- Mastering Textbook Questions with Pre-trained Transformers and Bottom-Up and Top-Down Attention

Abstract

Textbook Question Answering is a complex task in the intersection of MachineComprehension and Visual Question Answering that requires reasoning withmultimodal information from text and diagrams. For the first time, this papertaps on the potential of transformer language models and bottom-up and top-downattention to tackle the language and visual understanding challenges this taskentails. Rather than training a language-visual transformer from scratch werely on pre-trained transformers, fine-tuning and ensembling. We add bottom-upand top-down attention to identify regions of interest corresponding to diagramconstituents and their relationships, improving the selection of relevantvisual information for each question and answer options. Our system ISAAQreports unprecedented success in all TQA question types, with accuracies of81.36%, 71.11% and 55.12% on true/false, text-only and diagram multiple choicequestions. ISAAQ also demonstrates its broad applicability, obtainingstate-of-the-art results in other demanding datasets.

Quick Read (beta)

loading the full paper ...