paper-with-me

홈 › Papers

A Two-Stage Masked LM Method for Term Set Expansion

2020-05-03 · ACL 2020 6 · Guy Kushilevitz, Shaul Markovitch, Yoav Goldberg

We tackle the task of Term Set Expansion (TSE): given a small seed set of example terms from a semantic class, finding more members of that class. The task is of great practical utility, and also of theoretical utility as it requires generalization from few examples. Previous approaches to the TSE task can be characterized as either distributional or pattern-based. We harness the power of neural masked language models (MLM) and propose a novel TSE algorithm, which combines the pattern-based and distributional approaches. Due to the small size of the seed set, fine-tuning methods are not effective, calling for more creative use of the MLM. The gist of the idea is to use the MLM to first mine for informative patterns with respect to the seed set, and then to obtain more members of the seed class by generalizing these patterns. Our method outperforms state-of-the-art TSE algorithms. Implementation is available at: https://github.com/ guykush/TermSetExpansion-MPB/

📄 PDF Abstract BibTeX arXiv:2005.01063

Code (1)

guykush/TermSetExpansion-MPB 공식 구현

Tasks

Vocal Bursts Valence Prediction

Similar Papers 제목 키워드 기반

$ρ$-$\texttt{EOS}$: Training-free Bidirectional Variable-Length Control for Masked Diffusion LLMs

2026-01-30 · Jingyi Yang, Yuxian Jiang, Jing Shao arxiv

Beyond parallel generation and global context modeling, current masked diffusion large language models (masked dLLMs, i.e., LLaDA) suffer from a fundamental limitation: they require a predefined, fixed generation length,…

Computational Efficiency

Dynin-Omni: Omnimodal Unified Large Diffusion Language Model

2026-03-09 · Jaeik Kim, Woojin Kim, Jihwan Hong, Yejoon Lee 외 arxiv

We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlik…

Cross-Modal RetrievalSpeech RecognitionImage Generation

Multi-Signal Reconstruction Using Masked Autoencoder From EEG During Polysomnography

2023-11-14 · Young-Seok Kweon, Gi-Hwan Shin, Heon-Gyu Kwak, Ha-Na Jo 외

Polysomnography (PSG) is an indispensable diagnostic tool in sleep medicine, essential for identifying various sleep disorders. By capturing physiological signals, including EEG, EOG, EMG, and cardiorespiratory metrics, …

DiagnosticEEG

VoidPadding: Let [VOID] Handle Padding in Masked Diffusion Language Models so that [EOS] Can Focus on Semantic Termination

2026-06-16 · Chunyu Liu, Zhengyang Fan, Kaisen Yang, Alex Lamb arxiv

MDLMs generate text by denoising a preallocated masked response canvas, making response-length modeling central to instruction tuning. Existing MDLMs often inherit the autoregressive convention of using repeated \texttt{…

Mathematical ReasoningCode Generation

Beyond two-stage models for lung carcinogenesis in the Mayak workers: Implications for Plutonium risk

2015-04-27

Mechanistic multi-stage models are used to analyze lung-cancer mortality after Plutonium exposure in the Mayak-workers cohort, with follow-up until 2008. Besides the established two-stage model with clonal expansion, mod…