paper-with-me

홈 › Papers

MoCA: Incorporating Multi-stage Domain Pretraining and Cross-guided Multimodal Attention for Textbook Question Answering

2021-12-06 · Fangzhi Xu, Qika Lin, Jun Liu, Lingling Zhang, Tianzhe Zhao, Qi Chai, Yudai Pan

Textbook Question Answering (TQA) is a complex multimodal task to infer answers given large context descriptions and abundant diagrams. Compared with Visual Question Answering (VQA), TQA contains a large number of uncommon terminologies and various diagram inputs. It brings new challenges to the representation capability of language model for domain-specific spans. And it also pushes the multimodal fusion to a more complex level. To tackle the above issues, we propose a novel model named MoCA, which incorporates multi-stage domain pretraining and multimodal cross attention for the TQA task. Firstly, we introduce a multi-stage domain pretraining module to conduct unsupervised post-pretraining with the span mask strategy and supervised pre-finetune. Especially for domain post-pretraining, we propose a heuristic generation algorithm to employ the terminology corpus. Secondly, to fully consider the rich inputs of context and diagrams, we propose cross-guided multimodal attention to update the features of text, question diagram and instructional diagram based on a progressive strategy. Further, a dual gating mechanism is adopted to improve the model ensemble. The experimental results show the superiority of our model, which outperforms the state-of-the-art methods by 2.21% and 2.43% for validation and test split respectively.

📄 PDF Abstract BibTeX arXiv:2112.02839

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModellingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

MMVP: A Multimodal MoCap Dataset with Vision and Pressure Sensors

2024-03-26 · CVPR 2024 1 · He Zhang, Shenghao Ren, Haolei Yuan, Jianhui Zhao 외

Foot contact is an important cue for human motion capture, understanding, and generation. Existing datasets tend to annotate dense foot contact using visual matching with thresholding or incorporating pressure signals. H…

Translation

BundleMoCap: Efficient, Robust and Smooth Motion Capture from Sparse Multiview Videos

2023-11-21 · Georgios Albanis, Nikolaos Zioulis, Kostas Kolomvatsos

Capturing smooth motions from videos using markerless techniques typically involves complex processes such as temporal constraints, multiple stages with data-driven regression and optimization, and bundle solving over te…

3D human pose and shape estimation3D Human Pose EstimationMarkerless Motion Capture

m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder

2026-05-19 · Yaoxiang Wang, Simiao Zuo, Qingguo Hu, Yucheng Ding 외 arxiv

Embedding models are pivotal in industrial information retrieval systems like search and advertising. However, existing pretrained models often exhibit fixed architectures and embedding dimensionalities, posing significa…

Continual PretrainingInformation Retrieval

Align Your Query: Representation Alignment for Multimodality Medical Object Detection

2025-10-03 · Ara Seo, Bryan Sangwoo Kim, Hyungjin Chung, Jong Chul Ye arxiv

Medical object detection suffers when a single detector is trained on mixed medical modalities (e.g., CXR, CT, MRI) due to heterogeneous statistics and disjoint representation spaces. To address this challenge, we turn t…

Medical Object Detection

MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos

2025-12-11 · Kehong Gong, Zhengyu Wen, Weixia He, Mingxi Xu 외 arxiv

Motion capture now underpins content creation far beyond digital humans, yet most existing pipelines remain species- or template-specific. We formalize this gap as Category-Agnostic Motion Capture (CAMoCap): given a mono…