paper-with-me

홈 › Papers

scLLM-DSC: LLM-Knowledge Enhanced Cross-Modal Deep Structural Clustering for Single-Cell RNA Sequencing

2026-06-11 · Ping Xu, Pengjiang Li, Tian Du, Zaitian Wang, Jiawei Gu, Zhiyuan Ning, Ziyue Qiao, Pengfei Wang, Yuanchun Zhou arxiv

Clustering is fundamental to scRNA-seq analysis, serving as a cornerstone for identifying cell populations and resolving tissue heterogeneity. However, existing methods focus on mining numerical statistical patterns, suffering from semantic agnosticism by neglecting the intrinsic biological functions encoded by genes. While Large Language Models (LLMs) offer promising semantic capabilities, their direct adaptation to cell clustering is hindered by the structural mismatch between generative pre-training objectives and discriminative downstream tasks. To bridge this gap, we propose scLLM-DSC, a novel LLM-Knowledge Enhanced Cross-Modal Deep Structural Clustering framework. Diverging from data-driven paradigms, scLLM-DSC establishes a semantically-grounded representation by synergizing two views: a Knowledge-Driven Semantic View derived from NCBI gene priors and contextualized Cell2Sentence embeddings, and a Structure-Aware Topological View extracted via a graph-guided encoder. Crucially, we introduce a cross-modal contrastive alignment mechanism to enforce consistency between biological semantics and transcriptomic features within a unified latent space. Extensive benchmarks demonstrate that scLLM-DSC significantly outperforms eleven state-of-the-art baselines in clustering accuracy.

📄 PDF Abstract BibTeX arXiv:2606.13007

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Language-Enhanced Representation Learning for Single-Cell Transcriptomics

2025-03-12 · Yaorui Shi, Jiaqi Yang, Changhao Nai, Sihang Li 외

Single-cell RNA sequencing (scRNA-seq) offers detailed insights into cellular heterogeneity. Recent advancements leverage single-cell large language models (scLLMs) for effective representation learning. These models foc…

Language ModelingLanguage ModellingRepresentation Learning

CLOP: Video-and-Language Pre-Training with Knowledge Regularizations

2022-11-07 · Guohao Li, Hu Yang, Feng He, Zhifan Feng 외

Video-and-language pre-training has shown promising results for learning generalizable representations. Most existing approaches usually model video and text in an implicit manner, without considering explicit structural…

Contrastive LearningRetrievalVideo Retrieval

Knowledge-Enhanced Hierarchical Information Correlation Learning for Multi-Modal Rumor Detection

2023-06-28 · Jiawei Liu, Jingyi Xie, Fanrui Zhang, Qiang Zhang 외

The explosive growth of rumors with text and images on social media platforms has drawn great attention. Existing studies have made significant contributions to cross-modal information interaction and fusion, but they fa…

CMV-Fuse: Cross Modal-View Fusion of AMR, Syntax, and Knowledge Representations for Aspect Based Sentiment Analysis

2025-12-07 · Smitha Muthya Sudheendra, Mani Deep Cherukuri, Jaideep Srivastava arxiv

Natural language understanding inherently depends on integrating multiple complementary perspectives spanning from surface syntax to deep semantics and world knowledge. However, current Aspect-Based Sentiment Analysis (A…

Natural Language UnderstandingComputational EfficiencyConstituency ParsingContrastive Learning

Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model

2022-07-16 · Xiaolin Chen, Xuemeng Song, Liqiang Jing, Shuo Li 외

Text response generation for multimodal task-oriented dialog systems, which aims to generate the proper text response given the multimodal context, is an essential yet challenging task. Although existing efforts have ach…

DecoderLanguage ModelingLanguage ModellingResponse Generation