Refining Dimensions for Improving Clustering-based Cross-lingual Topic Models
Recent works in clustering-based topic models perform well in monolingual topic identification by introducing a pipeline to cluster the contextualized representations. However, the pipeline is suboptimal in identifying topics across languages due to the presence of language-dependent dimensions (LDDs) generated by multilingual language models. To address this issue, we introduce a novel, SVD-based dimension refinement component into the pipeline of the clustering-based topic model. This component effectively neutralizes the negative impact of LDDs, enabling the model to accurately identify topics across languages. Our experiments on three datasets demonstrate that the updated pipeline with the dimension refinement component generally outperforms other state-of-the-art cross-lingual topic models.
Code (1)
Tasks
ClusteringTopic ModelsSimilar Papers 제목 키워드 기반
Research on Multilingual News Clustering Based on Cross-Language Word Embeddings
Classifying the same event reported by different countries is of significant importance for public opinion control and intelligence gathering. Due to the diverse types of news, relying solely on transla-tors would be cos…
ClusteringKnowledge DistillationSentenceTranslation+1German Text Embedding Clustering Benchmark
This work introduces a benchmark assessing the performance of clustering German text embeddings in different domains. This benchmark is driven by the increasing use of clustering neural text embeddings in tasks that requ…
BenchmarkingClusteringCLTC: A Chinese-English Cross-lingual Topic Corpus
Cross-lingual topic detection within text is a feasible solution to resolving the language barrier in accessing the information. This paper presents a Chinese-English cross-lingual topic corpus (CLTC), in which 90,000 Ch…
ArticlesClusteringText ClusteringHierarchical Level-Wise News Article Clustering via Multilingual Matryoshka Embeddings
Contextual large language model embeddings are increasingly utilized for topic modeling and clustering. However, current methods often scale poorly, rely on opaque similarity metrics, and struggle in multilingual setting…
ArticlesClusteringLanguage ModelingLanguage Modelling+1Batch Clustering for Multilingual News Streaming
Nowadays, digital news articles are widely available, published by various editors and often written in different languages. This large volume of diverse and unorganized information makes human reading very difficult or …
ArticlesClustering