DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Abstract

This work presents DocPedia, a novel large multimodal model (LMM) forversatile OCR-free document understanding, capable of parsing images up to2,560$\times$2,560 resolution. Unlike existing work either struggle withhigh-resolution documents or give up the large language model thus vision orlanguage ability constrained, our DocPedia directly processes visual input inthe frequency domain rather than the pixel space. The unique characteristicenables DocPedia to capture a greater amount of visual and textual informationusing a limited number of visual tokens. To consistently enhance bothperception and comprehension abilities of our model, we develop a dual-stagetraining strategy and enrich instructions/annotations of all training taskscovering multiple document types. Extensive quantitative and qualitativeexperiments conducted on various publicly available benchmarks confirm themutual benefits of jointly learning perception and comprehension tasks. Theresults provide further evidence of the effectiveness and superior performanceof our DocPedia over other methods.

Quick Read (beta)

loading the full paper ...