paper-with-me

홈 › Papers

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models

2025-01-28 · Zeping Min, Xinshang Wang

We introduce a novel index, the Distribution of Cosine Similarity (DOCS), for quantitatively assessing the similarity between weight matrices in Large Language Models (LLMs), aiming to facilitate the analysis of their complex architectures. Leveraging DOCS, our analysis uncovers intriguing patterns in the latest open-source LLMs: adjacent layers frequently exhibit high weight similarity and tend to form clusters, suggesting depth-wise functional specialization. Additionally, we prove that DOCS is theoretically effective in quantifying similarity for orthogonal matrices, a crucial aspect given the prevalence of orthogonal initializations in LLMs. This research contributes to a deeper understanding of LLM architecture and behavior, offering tools with potential implications for developing more efficient and interpretable models.

📄 PDF Abstract BibTeX arXiv:2501.16650

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DocScanner: Robust Document Image Rectification with Progressive Learning

2021-10-28 · Hao Feng, Wengang Zhou, Jiajun Deng, Qi Tian 외

Compared with flatbed scanners, portable smartphones provide more convenience for physical document digitization. However, such digitized documents are often distorted due to uncontrolled physical deformations, camera po…

Optical Character Recognition (OCR)

Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models

2024-10-17 · Jiatao Li, Xinyu Hu, Xunjian Yin, Xiaojun Wan

The integration of documents generated by LLMs themselves (Self-Docs) alongside retrieved documents has emerged as a promising strategy for retrieval-augmented generation systems. However, previous research primarily foc…

Language ModellingLarge Language ModelQuestion AnsweringRAG+2

Contrastive Similarity Learning for Market Forecasting: The ContraSim Framework

2025-02-22 · Nicholas Vinden, Raeid Saqur, Zining Zhu, Frank Rudzicz

We introduce the Contrastive Similarity Space Embedding Algorithm (ContraSim), a novel framework for uncovering the global semantic relationships between daily financial headlines and market movements. ContraSim operates…

Contrastive Learning

A cohomology-based Gromov-Hausdorff metric approach for quantifying molecular similarity

2024-11-21 · JunJie Wee, Xue Gong, Wilderich Tuschmann, Kelin Xia

We introduce, for the first time, a cohomology-based Gromov-Hausdorff ultrametric method to analyze 1-dimensional and higher-dimensional (co)homology groups, focusing on loops, voids, and higher-dimensional cavity struct…

Clustering

Towards Quantifying the Distance between Opinions

2020-01-27 · Saket Gurukar, Deepak Ajwani, Sourav Dutta, Juho Lauri 외

Increasingly, critical decisions in public policy, governance, and business strategy rely on a deeper understanding of the needs and opinions of constituent members (e.g. citizens, shareholders). While it has become easi…

Navigatetext similarity