paper-with-me

홈 › Papers

PRISM: PRIor from corpus Statistics for topic Modeling

2026-03-31 · Tal Ishon, Yoav Goldberg, Uri Shaham arxiv

Topic modeling seeks to uncover latent semantic structure in text, with LDA providing a foundational probabilistic framework. While recent methods often incorporate external knowledge (e.g., pre-trained embeddings), such reliance limits applicability in emerging or underexplored domains. We introduce \textbf{PRISM}, a corpus-intrinsic method that derives a Dirichlet parameter from word co-occurrence statistics to initialize LDA without altering its generative process. Experiments on text and single cell RNA-seq data show that PRISM improves topic coherence and interpretability, rivaling models that rely on external knowledge. These results underscore the value of corpus-driven initialization for topic modeling in resource-constrained settings. Code is available at: https://github.com/shaham-lab/PRISM.

📄 PDF Abstract BibTeX arXiv:2603.29406

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PRISM: LLM-Guided Semantic Clustering for High-Precision Topics

2026-04-03 · Connor Douglas, Utkucan Balci, Joseph Aylett-Bullock arxiv

In this paper, we propose Precision-Informed Semantic Modeling (PRISM), a structured topic modeling framework combining the benefits of rich representations captured by LLMs with the low cost and interpretability of late…

Topic Models

Anchor-Free Correlated Topic Modeling: Identifiability and Algorithm

2016-11-15 · NeurIPS 2016 12 · Kejun Huang, Xiao Fu, Nicholas D. Sidiropoulos

In topic modeling, many algorithms that guarantee identifiability of the topics have been developed under the premise that there exist anchor words -- i.e., words that only appear (with positive probability) in one topic…

Clustering

Conditional Language Learning with Context

2024-06-04 · Xiao Zhang, Miao Li, Ji Wu

Language models can learn sophisticated language understanding skills from fitting raw text. They also unselectively learn useless corpus statistics and biases, especially during finetuning on domain-specific corpora. In…

Causal Language ModelingLanguage ModelingLanguage ModellingLifelong learning

Adaptive Mixed Component LDA for Low Resource Topic Modeling

2021-04-01 · EACL 2021 2 · Suzanna Sia, Kevin Duh

Probabilistic topic models in low data resource scenarios are faced with less reliable estimates due to sparsity of discrete word co-occurrence counts, and do not have the luxury of retraining word or topic embeddings us…

Topic Models

Specific Aspects of Intellectual Property Management in the Knowledge-Based Economy

2025-01-14 · Aurel Mihail Titu, Alina Bianca Pop, Camelia Oprean-Stan, Sebastian Emanuel Stan

This paper addresses the issue of intellectual property management in the knowledge-based economy. The starting point in carrying out the study is the presentation of some concepts regarding in the first phase, the intel…

Management