paper-with-me

홈 › Papers

The Critical Role of Aspects in Measuring Document Similarity

2026-01-06 · Eftekhar Hossain, Tarnika Hazra, Ahatesham Bhuiyan, Santu Karmaker arxiv

We introduce ASPECTSIM, a simple and interpretable framework that requires conditioning document similarity on an explicitly specified aspect, which is different from the traditional holistic approach in measuring document similarity. Experimenting with a newly constructed benchmark of 26K aspect-document pairs, we found that ASPECTSIM, when implemented with direct GPT-4o prompting, achieves substantially higher human-machine agreement ($\approx$80% higher) than the same for holistic similarity without explicit aspects. These findings underscore the importance of explicitly accounting for aspects when measuring document similarity and highlight the need to revise standard practice. Next, we conducted a large-scale meta-evaluation using 16 smaller open-source LLMs and 9 embedding models with a focus on making ASPECTSIM accessible and reproducible. While directly prompting LLMs to produce ASPECTSIM scores turned out be ineffective (20-30% human-machine agreement), a simple two-stage refinement improved their agreement by $\approx$140%. Nevertheless, agreement remains well below that of GPT-4o-based models, indicating that smaller open-source LLMs still lag behind large proprietary models in capturing aspect-conditioned similarity.

📄 PDF Abstract BibTeX arXiv:2601.03435

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unsupervised Concept Representation Learning for Length-Varying Text Similarity

2021-06-01 · NAACL 2021 4 · Xuchao Zhang, Bo Zong, Wei Cheng, Jingchao Ni 외

Measuring document similarity plays an important role in natural language processing tasks. Most existing document similarity approaches suffer from the information gap caused by context and vocabulary mismatches when co…

Representation Learningtext similarity

Methods for Computing Legal Document Similarity: A Comparative Study

2020-04-26 · Paheli Bhattacharya, Kripabandhu Ghosh, Arindam Pal, Saptarshi Ghosh

Computing similarity between two legal documents is an important and challenging task in the domain of Legal Information Retrieval. Finding similar legal documents has many applications in downstream tasks, including pri…

ArticlesInformation RetrievalRetrieval

We've had this conversation before: A Novel Approach to Measuring Dialog Similarity

2021-10-12 · Ofer Lavi, Ella Rabinovich, Segev Shlomov, David Boaz 외

Dialog is a core building block of human natural language interactions. It contains multi-party utterances used to convey information from one party to another in a dynamic and evolving manner. The ability to compare dia…

We’ve had this conversation before: A Novel Approach to Measuring Dialog Similarity

2021-11-01 · EMNLP 2021 11 · Ofer Lavi, Ella Rabinovich, Segev Shlomov, David Boaz 외

Dialog is a core building block of human natural language interactions. It contains multi-party utterances used to convey information from one party to another in a dynamic and evolving manner. The ability to compare dia…

Multi-Vector Models with Textual Guidance for Fine-Grained Scientific Document Similarity

2021-12-17 · ACL ARR December 2022 12 · Anonymous

We present a new scientific document similarity model based on matching fine-grained aspects of texts. To train our model, we exploit a naturally-occurring source of supervision: sentences in the full-text of papers that…

Sentence