paper-with-me

Papers

FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction

2023-05-04 · Chen-Yu Lee, Chun-Liang Li, Hao Zhang, Timothy Dozat, Vincent Perot, Guolong Su, Xiang Zhang, Kihyuk Sohn, Nikolai Glushnev, Renshen Wang, Joshua Ainslie, Shangbang Long, Siyang Qin, Yasuhisa Fujii, Nan Hua, Tomas Pfister

The recent advent of self-supervised pre-training techniques has led to a surge in the use of multimodal learning in form document understanding. However, existing approaches that extend the mask language modeling to other modalities require careful multi-task tuning, complex reconstruction target designs, or additional pre-training data. In FormNetV2, we introduce a centralized multimodal graph contrastive learning strategy to unify self-supervised pre-training for all modalities in one loss. The graph contrastive objective maximizes the agreement of multimodal representations, providing a natural interplay for all modalities without special customization. In addition, we extract image features within the bounding box that joins a pair of tokens connected by a graph edge, capturing more targeted visual cues without loading a sophisticated and separately pre-trained image embedder. FormNetV2 establishes new state-of-the-art performance on FUNSD, CORD, SROIE and Payment benchmarks with a more compact model size.

📄 PDF Abstract BibTeX arXiv:2305.02549

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learningdocument understandingFormLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval

2025-11-02 · Ahmed Masry, Megh Thakkar, Patrice Bechard, Sathwik Tejaswi Madhusudhan 외 arxiv

Retrieval-augmented generation has proven practical when models require specialized knowledge or access to the latest data. However, existing methods for multimodal document retrieval often replicate techniques developed…

Representation LearningContrastive Learning

Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents

2025-10-21 · Yiqi Lin, Alex Jinpeng Wang, Linjie Li, Zhengyuan Yang 외 arxiv

Contrastive vision-language models such as CLIP have demonstrated strong performance across a wide range of multimodal tasks by learning from aligned image-text pairs. However, their ability to handle complex, real-world…

Representation LearningCross-Modal RetrievalContrastive Learning

Linking Representations with Multimodal Contrastive Learning

2023-04-07 · Abhishek Arora, Xinmei Yang, Shao-Yu Jheng, Melissa Dell

Many applications require linking individuals, firms, or locations across datasets. Most widely used methods, especially in social science, do not employ deep learning, with record linkage commonly approached using strin…

Contrastive LearningOptical Character RecognitionOptical Character Recognition (OCR)

CMDR: Contextual Multimodal Document Retrieval

2026-07-07 · Ryota Tanaka, Taku Hasegawa, Kyosuke Nishida arxiv

Multimodal document retrieval aims to retrieve relevant pages while preserving both textual and visual content from the original document. However, existing benchmarks primarily evaluate simple lexical or semantic matchi…

Contrastive Learning

Graph Contrastive Topic Model

2023-07-05 · Zheheng Luo, Lei Liu, Qianqian Xie, Sophia Ananiadou

Existing NTMs with contrastive learning suffer from the sample bias problem owing to the word frequency-based sampling strategy, which may result in false negative samples with similar semantics to the prototypes. In thi…

Contrastive LearningmodelRepresentation Learning