paper-with-me

홈 › Papers

Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning

2024-07-22 · Xiangyan Qu, Jing Yu, Keke Gai, Jiamin Zhuang, Yuanmin Tang, Gang Xiong, Gaopeng Gou, Qi Wu

Recent work shows that documents from encyclopedias serve as helpful auxiliary information for zero-shot learning. Existing methods align the entire semantics of a document with corresponding images to transfer knowledge. However, they disregard that semantic information is not equivalent between them, resulting in a suboptimal alignment. In this work, we propose a novel network to extract multi-view semantic concepts from documents and images and align the matching rather than entire concepts. Specifically, we propose a semantic decomposition module to generate multi-view semantic embeddings from visual and textual sides, providing the basic concepts for partial alignment. To alleviate the issue of information redundancy among embeddings, we propose the local-to-semantic variance loss to capture distinct local details and multiple semantic diversity loss to enforce orthogonality among embeddings. Subsequently, two losses are introduced to partially align visual-semantic embedding pairs according to their semantic relevance at the view and word-to-patch levels. Consequently, we consistently outperform state-of-the-art methods under two document sources in three standard benchmarks for document-based zero-shot learning. Qualitatively, we show that our model learns the interpretable partial association.

📄 PDF Abstract BibTeX arXiv:2407.15613

Code (1)

morningstarovo/emdepart 공식 구현 pytorch

Tasks

DiversityZero-Shot Learning

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval

2025-07-28 · Yili Li, Gang Xiong, Gaopeng Gou, Xiangyan Qu 외 arxiv

Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrated by image-text pretrained models such …

Text to Video Retrieval

Bilingual Document Alignment with Latent Semantic Indexing

2017-07-29 · WS 2016 8 · Ulrich Germann

We apply cross-lingual Latent Semantic Indexing to the Bilingual Document Alignment Task at WMT16. Reduced-rank singular value decomposition of a bilingual term-document matrix derived from known English/French page pair…

Which Information Matters? Dissecting Human-written Multi-document Summaries with Partial Information Decomposition

2024-05-23 · Laura Mascarell, Yan L'Homme, Majed El Helou

Understanding the nature of high-quality summaries is crucial to further improve the performance of multi-document summarization. We propose an approach to characterize human-written summaries using partial information d…

Document SummarizationMulti-Document Summarization

Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection

2026-03-25 · Sa Zhu, Wanqian Zhang, Lin Wang, Xiaohua Chen 외 arxiv

Open-Vocabulary Temporal Action Detection (OV-TAD) aims to classify and localize action segments in untrimmed videos for unseen categories. Previous methods rely solely on global alignment between label-level semantics a…

Action Detection

Decomposing and Recomposing Event Structure

2021-03-18 · William Gantt, Lelia Glass, Aaron Steven White

We present an event structure classification empirically derived from inferential properties annotated on sentence- and document-level Universal Decompositional Semantics (UDS) graphs. We induce this classification joint…

ClassificationSentence