Publication | Closed Access
DocFormer: End-to-End Transformer for Document Understanding
267
Citations
53
References
2021
Year
EngineeringMachine LearningVisual Document UnderstandingSemantic WebText MiningNatural Language ProcessingMultimodal LlmVisual GroundingSpatial EmbeddingsDocument EngineeringComputational LinguisticsDocument UnderstandingVisual Question AnsweringLanguage StudiesMachine TranslationMulti-modal TransformerVision Language ModelComputer VisionStructured DocumentLinguisticsDocument Processing
Visual Document Understanding is a challenging problem that seeks to interpret documents in varied formats and layouts. The authors present DocFormer, a multi‑modal transformer architecture for Visual Document Understanding. DocFormer is an unsupervised, multi‑modal transformer that fuses text, vision, and spatial features through a novel self‑attention layer, shares spatial embeddings across modalities, and is evaluated on four datasets. DocFormer achieves state‑of‑the‑art performance on all four datasets, sometimes outperforming models four times its size in terms of parameters.
We present DocFormer - a multi-modal transformer based architecture for the task of Visual Document Understanding (VDU). VDU is a challenging problem which aims to understand documents in their varied formats (forms, receipts etc.) and layouts. In addition, DocFormer is pre-trained in an unsupervised fashion using carefully designed tasks which encourage multi-modal interaction. DocFormer uses text, vision and spatial features and combines them using a novel multi-modal self-attention layer. DocFormer also shares learned spatial embeddings across modalities which makes it easy for the model to correlate text to visual tokens and vice versa. DocFormer is evaluated on 4 different datasets each with strong baselines. DocFormer achieves state-of-the-art results on all of them, sometimes beating models 4x its size (in no. of parameters).
| Year | Citations | |
|---|---|---|
Page 1
Page 1