paper-with-me

Papers

Refining Dimensions for Improving Clustering-based Cross-lingual Topic Models

2024-12-17 · Chia-Hsuan Chang, Tien-Yuan Huang, Yi-Hang Tsai, Chia-Ming Chang, San-Yih Hwang

Recent works in clustering-based topic models perform well in monolingual topic identification by introducing a pipeline to cluster the contextualized representations. However, the pipeline is suboptimal in identifying topics across languages due to the presence of language-dependent dimensions (LDDs) generated by multilingual language models. To address this issue, we introduce a novel, SVD-based dimension refinement component into the pipeline of the clustering-based topic model. This component effectively neutralizes the negative impact of LDDs, enabling the model to accurately identify topics across languages. Our experiments on three datasets demonstrate that the updated pipeline with the dimension refinement component generally outperforms other state-of-the-art cross-lingual topic models.

📄 PDF Abstract BibTeX arXiv:2412.12433

Code (1)

Text-Analytics-and-Retrieval/Clustering-based-Cross-Lingual-Topic-Model 공식 구현

Tasks

ClusteringTopic Models

Similar Papers 제목 키워드 기반

Research on Multilingual News Clustering Based on Cross-Language Word Embeddings

2023-05-30 · Lin Wu, Rui Li, Wong-Hing Lam

Classifying the same event reported by different countries is of significant importance for public opinion control and intelligence gathering. Due to the diverse types of news, relying solely on transla-tors would be cos…

ClusteringKnowledge DistillationSentenceTranslation+1

German Text Embedding Clustering Benchmark

2024-01-05 · Silvan Wehrli, Bert Arnrich, Christopher Irrgang

This work introduces a benchmark assessing the performance of clustering German text embeddings in different domains. This benchmark is driven by the increasing use of clustering neural text embeddings in tasks that requ…

BenchmarkingClustering

CLTC: A Chinese-English Cross-lingual Topic Corpus

2012-05-01 · LREC 2012 5 · Yunqing Xia, Guoyu Tang, Peng Jin, Xia Yang

Cross-lingual topic detection within text is a feasible solution to resolving the language barrier in accessing the information. This paper presents a Chinese-English cross-lingual topic corpus (CLTC), in which 90,000 Ch…

ArticlesClusteringText Clustering

Hierarchical Level-Wise News Article Clustering via Multilingual Matryoshka Embeddings

2025-05-30 · Hans W. A. Hanley, Zakir Durumeric

Contextual large language model embeddings are increasingly utilized for topic modeling and clustering. However, current methods often scale poorly, rely on opaque similarity metrics, and struggle in multilingual setting…

ArticlesClusteringLanguage ModelingLanguage Modelling+1

Batch Clustering for Multilingual News Streaming

2020-04-17 · Mathis Linger, Mhamed Hajaiej

Nowadays, digital news articles are widely available, published by various editors and often written in different languages. This large volume of diverse and unorganized information makes human reading very difficult or …

ArticlesClustering