paper-with-me

Papers

Understanding Long Documents with Different Position-Aware Attentions

2022-08-17 · Hai Pham, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang

Despite several successes in document understanding, the practical task for long document understanding is largely under-explored due to several challenges in computation and how to efficiently absorb long multimodal input. Most current transformer-based approaches only deal with short documents and employ solely textual information for attention due to its prohibitive computation and memory limit. To address those issues in long document understanding, we explore different approaches in handling 1D and new 2D position-aware attention with essentially shortened context. Experimental results show that our proposed models have advantages for this task based on various evaluation metrics. Furthermore, our model makes changes only to the attention and thus can be easily adapted to any transformer-based architecture.

📄 PDF Abstract BibTeX arXiv:2208.08201

Code (0)

등록된 구현이 없습니다.

Tasks

document understandingPosition

Similar Papers 제목 키워드 기반

Position Masking for Improved Layout-Aware Document Understanding

2021-09-01 · Anik Saha, Catherine Finegan-Dollak, Ashish Verma

Natural language processing for document scans and PDFs has the potential to enormously improve the efficiency of business processes. Layout-aware word embeddings such as LayoutLM have shown promise for classification of…

document understandingPositionWord Embeddings

Probing Position-Aware Attention Mechanism in Long Document Understanding

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Long document understanding is a challenging problem in natural language understanding. Most current transformer-based models only employ textual information for attention calculation due to high computation limit. To ad…

document understandingNatural Language UnderstandingPosition

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

2026-07-11 · Abhigya Verma, Khyati Mahajan, Amit Kumar Saha, Shruthan Radhakrishna 외 arxiv

Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as leng…

Supervised Contrastive Learning for Interpretable Long-Form Document Matching

2021-08-20 · Akshita Jha, Vineeth Rakesh, Jaideep Chandrashekar, Adithya Samavedhi 외

Recent advancements in deep learning techniques have transformed the area of semantic text matching. However, most state-of-the-art models are designed to operate with short documents such as tweets, user reviews, commen…

ArticlesContrastive LearningFormSemantic Text Matching+1

PharmaShip: An Entity-Centric, Reading-Order-Supervised Benchmark for Chinese Pharmaceutical Shipping Documents

2025-11-29 · Tingwei Xie, Tianyi Zhou, Yonghong Song arxiv

We present PharmaShip, a real-world Chinese dataset of scanned pharmaceutical shipping documents designed to stress-test pre-trained text-layout models under noisy OCR and heterogeneous templates. PharmaShip covers three…

Relation Extraction