paper-with-me

Papers

META: Metadata-Empowered Weak Supervision for Text Classification

2020-11-01 · EMNLP 2020 11 · Dheeraj Mekala, Xinyang Zhang, Jingbo Shang

Recent advances in weakly supervised learning enable training high-quality text classifiers by only providing a few user-provided seed words. Existing methods mainly use text data alone to generate pseudo-labels despite the fact that metadata information (e.g., author and timestamp) is widely available across various domains. Strong label indicators exist in the metadata and it has been long overlooked mainly due to the following challenges: (1) metadata is multi-typed, requiring systematic modeling of different types and their combinations, (2) metadata is noisy, some metadata entities (e.g., authors, venues) are more compelling label indicators than others. In this paper, we propose a novel framework, META, which goes beyond the existing paradigm and leverages metadata as an additional source of weak supervision. Specifically, we organize the text data and metadata together into a text-rich network and adopt network motifs to capture appropriate combinations of metadata. Based on seed words, we rank and filter motif instances to distill highly label-indicative ones as {``}seed motifs{''}, which provide additional weak supervision. Following a bootstrapping manner, we train the classifier and expand the seed words and seed motifs iteratively. Extensive experiments and case studies on real-world datasets demonstrate superior performance and significant advantages of leveraging metadata as weak supervision.

📄 PDF Abstract BibTeX

Code (1)

dheeraj7596/META 공식 구현 tf

Tasks

ClassificationGeneral Classificationtext-classificationText ClassificationWeakly-supervised Learning

Similar Papers 제목 키워드 기반

Hierarchical Metadata-Aware Document Categorization under Weak Supervision

2020-10-26 · Yu Zhang, Xiusi Chen, Yu Meng, Jiawei Han

Categorizing documents into a given label hierarchy is intuitively appealing due to the ubiquity of hierarchical topic structures in massive text corpora. Although related studies have achieved satisfying performance in …

Data AugmentationDocument ClassificationRepresentation Learning

Learning from Acquisition: Metadata-driven Multimodal Pre-training for Cardiac MRI

2026-06-27 · Xueyi Fu, Liwei Hu, Zi Wang, Guang Yang arxiv

Cardiac magnetic resonance imaging (CMR) routinely records structured acquisition metadata, yet most CMR foundation models rely primarily on image-only pre-training and leave this naturally available source of weak seman…

Representation Learning

MotifClass: Weakly Supervised Text Classification with Higher-order Metadata Information

2021-11-07 · Yu Zhang, Shweta Garg, Yu Meng, Xiusi Chen 외

We study the problem of weakly supervised text classification, which aims to classify text documents into a set of pre-defined categories with category surface names only and without any annotated training document provi…

text-classificationText Classification

Invisible Shortcuts: Why Vision Encoders Know Your Camera

2026-08-05 · Vladan Stojnić, Ryan Ramos, Giorgos Kordopatis-Zilos, Noa Garcia 외 hf

Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correlations. We identify a different source o…

How Universal is Genre in Universal Dependencies?

2021-12-09 · ACL (TLT, SyntaxFest) 2021 12 · Max Müller-Eberstein, Rob van der Goot, Barbara Plank

This work provides the first in-depth analysis of genre in Universal Dependencies (UD). In contrast to prior work on genre identification which uses small sets of well-defined labels in mono-/bilingual setups, UD contain…

Specificity