paper-with-me

홈 › Papers

Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning

2020-09-30 · Subhojeet Pramanik, Shashank Mujumdar, Hima Patel

Recent approaches in literature have exploited the multi-modal information in documents (text, layout, image) to serve specific downstream document tasks. However, they are limited by their - (i) inability to learn cross-modal representations across text, layout and image dimensions for documents and (ii) inability to process multi-page documents. Pre-training techniques have been shown in Natural Language Processing (NLP) domain to learn generic textual representations from large unlabelled datasets, applicable to various downstream NLP tasks. In this paper, we propose a multi-task learning-based framework that utilizes a combination of self-supervised and supervised pre-training tasks to learn a generic document representation applicable to various downstream document tasks. Specifically, we introduce Document Topic Modelling and Document Shuffle Prediction as novel pre-training tasks to learn rich image representations along with the text and layout representations for documents. We utilize the Longformer network architecture as the backbone to encode the multi-modal information from multi-page documents in an end-to-end fashion. We showcase the applicability of our pre-training framework on a variety of different real-world document tasks such as document classification, document information extraction, and document retrieval. We evaluate our framework on different standard document datasets and conduct exhaustive experiments to compare performance against various ablations of our framework and state-of-the-art baselines.

📄 PDF Abstract BibTeX arXiv:2009.14457

Code (0)

등록된 구현이 없습니다.

Tasks

Document ClassificationMulti-Task LearningRepresentation LearningRetrieval

Similar Papers 제목 키워드 기반

FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning

2026-05-10 · Xing Han, Shravan Chaudhari, Tanvi Ranade, Rama Chellappa 외 arxiv

Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: (1) multi-task pretraining, tasks are co-available at design time where related tasks could borro…

Multi-Task LearningContinual Learning

M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-training

2020-06-04 · CVPR 2021 1 · Minheng Ni, Haoyang Huang, Lin Su, Edward Cui 외

We present M3P, a Multitask Multilingual Multimodal Pre-trained model that combines multilingual pre-training and multimodal pre-training into a unified framework via multitask pre-training. Our goal is to learn universa…

Image CaptioningImage RetrievalMachine TranslationMultimodal Machine Translation+3

Multi-modal Multi-task Pre-training for Improved Point Cloud Understanding

2025-07-23 · Liwen Liu, Weidong Yang, Lipeng Ma, Ben Fei arxiv

Recent advances in multi-modal pre-training methods have shown promising effectiveness in learning 3D representations by aligning multi-modal features between 3D shapes and their corresponding 2D counterparts. However, e…

Contrastive Learning

MultiMAE: Multi-modal Multi-task Masked Autoencoders

2022-04-04 · Roman Bachmann, David Mizrahi, Andrei Atanov, Amir Zamir

We propose a pre-training strategy called Multi-modal Multi-task Masked Autoencoders (MultiMAE). It differs from standard Masked Autoencoding in two key aspects: I) it can optionally accept additional modalities of infor…

Depth Estimationimage-classificationImage ClassificationSemantic Segmentation

Unimodal Face Classification with Multimodal Training

2021-12-08 · Wenbin Teng, Chongyang Bai

Face recognition is a crucial task in various multimedia applications such as security check, credential access and motion sensing games. However, the task is challenging when an input face is noisy (e.g. poor-condition …

ClassificationFace Recognition