paper-with-me

Papers

MATrIX -- Modality-Aware Transformer for Information eXtraction

2022-05-17 · Thomas Delteil, Edouard Belval, Lei Chen, Luis Goncalves, Vijay Mahadevan

We present MATrIX - a Modality-Aware Transformer for Information eXtraction in the Visual Document Understanding (VDU) domain. VDU covers information extraction from visually rich documents such as forms, invoices, receipts, tables, graphs, presentations, or advertisements. In these, text semantics and visual information supplement each other to provide a global understanding of the document. MATrIX is pre-trained in an unsupervised way with specifically designed tasks that require the use of multi-modal information (spatial, visual, or textual). We consider the spatial and text modalities all at once in a single token set. To make the attention more flexible, we use a learned modality-aware relative bias in the attention mechanism to modulate the attention between the tokens of different modalities. We evaluate MATrIX on 3 different datasets each with strong baselines.

📄 PDF Abstract BibTeX arXiv:2205.08094

Code (0)

등록된 구현이 없습니다.

Tasks

document understanding

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Residual Connection 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…

Similar Papers 제목 키워드 기반

Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object Detection

2023-08-17 · Runmin Cong, Hongyu Liu, Chen Zhang, Wei zhang 외

By integrating complementary information from RGB image and depth map, the ability of salient object detection (SOD) for complex and challenging scenes can be improved. In recent years, the important role of Convolutiona…

object-detectionObject DetectionRGB-D Salient Object DetectionSalient Object Detection

NestedFormer: Nested Modality-Aware Transformer for Brain Tumor Segmentation

2022-08-31 · Zhaohu Xing, Lequan Yu, Liang Wan, Tong Han 외

Multi-modal MR imaging is routinely used in clinical practice to diagnose and investigate brain tumors by providing rich complementary information. Previous multi-modal MRI segmentation methods usually perform modal fusi…

Brain Tumor SegmentationDecoderMRI segmentationSegmentation+1

Cross-Modality Fusion Transformer for Multispectral Object Detection

2021-10-30 · Fang Qingyun, Han Dapeng, Wang Zhaokui

Multispectral image pairs can provide the combined information, making object detection applications more reliable and robust in the open world. To fully exploit the different modalities, we present a simple yet effectiv…

Multispectral Object DetectionObjectobject-detectionObject Detection+1

Modality-aware Transformer for Financial Time series Forecasting

2023-10-02 · Hajar Emami, Xuan-Hong Dang, Yousaf Shah, Petros Zerfos

Time series forecasting presents a significant challenge, particularly when its accuracy relies on external data sources rather than solely on historical values. This issue is prevalent in the financial sector, where the…

Feature ImportanceTime SeriesTime Series Forecasting

Multi-Dimension-Embedding-Aware Modality Fusion Transformer for Psychiatric Disorder Clasification

2023-10-04 · Guoxin Wang, Xuyang Cao, Shan An, Fengmei Fan 외

Deep learning approaches, together with neuroimaging techniques, play an important role in psychiatric disorders classification. Previous studies on psychiatric disorders diagnosis mainly focus on using functional connec…

Functional ConnectivityTime Series