paper-with-me

Papers

Do LLM Embedding Spaces Recover Expert Structure?

2026-06-22 · Yixuan Zhu, Zhenke Duan, Fanghen Li arxiv

Pretrained text embeddings are increasingly used as representational maps, yet high category separability does not imply that their geometry recovers expert-defined structure. We study this problem in mental-health-related language, where symptom relations provide an external reference and online communities introduce strong domain, affective, stylistic, and discourse confounds. Using 28 Reddit communities, we compare pretrained and supervised fine-tuned Qwen3 embedding spaces at two scales (0.6B and 4B). We construct category prototypes, evaluate their representational dissimilarity matrices against an expert symptom matrix with representational similarity analysis, and complement this global test with prototype-based typicality and multi-baseline confound controls. Pretrained embeddings show measurable alignment with expert structure within the mental-health subset; fine-tuning strengthens this alignment most at the finest category level; and larger scale improves both zero-shot alignment and supervision-induced gains. Residual alignment remains substantial after controlling for VAD, LIWC, lexical style, and topic-distribution structure. These results suggest that LLM embeddings can recover expert-relevant category geometry, but this recovery is level-dependent and should be tested against explicit confounds rather than inferred from classification alone.

📄 PDF Abstract BibTeX arXiv:2606.23394

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Emblaze: Illuminating Machine Learning Representations through Interactive Comparison of Embedding Spaces

2022-02-05 · Venkatesh Sivaraman, Yiwei Wu, Adam Perer

Modern machine learning techniques commonly rely on complex, high-dimensional embedding representations to capture underlying structure in the data and improve performance. In order to characterize model flaws and choose…

BIG-bench Machine Learning

Assessing Multimodal Chronic Wound Embeddings with Expert Triplet Agreement

2026-03-31 · Fabian Kabus, Julia Hindel, Jelena Bratulić, Meropi Karakioulaki 외 arxiv

Recessive dystrophic epidermolysis bullosa (RDEB) is a rare genetic skin disorder for which clinicians greatly benefit from finding similar cases using images and clinical text. However, off-the-shelf foundation models d…

Representation Learning

Word Embeddings as Metric Recovery in Semantic Spaces

2016-01-01 · TACL 2016 1 · Tatsunori B. Hashimoto, David Alvarez-Melis, Tommi S. Jaakkola

Continuous word representations have been remarkably useful across NLP tasks but remain poorly understood. We ground word embeddings in semantic spaces studied in the cognitive-psychometric literature, taking these space…

Named Entity Recognition (NER)Semantic SimilaritySemantic Textual SimilarityWord Embeddings

LORE: Jointly Learning the Intrinsic Dimensionality and Relative Similarity Structure From Ordinal Data

2026-02-04 · Vivek Anand, Alec Helbling, Mark A. Davenport, Gordon J. Berman 외 arxiv

Learning the intrinsic dimensionality of subjective perceptual spaces such as taste, smell, or aesthetics from ordinal data is a challenging problem. We introduce LORE (Low Rank Ordinal Embedding), a scalable framework t…

Probing BERT in Hyperbolic Spaces

2021-04-08 · ICLR 2021 1 · Boli Chen, Yao Fu, Guangwei Xu, Pengjun Xie 외

Recently, a variety of probing tasks are proposed to discover linguistic properties learned in contextualized word embeddings. Many of these works implicitly assume these embeddings lay in certain metric spaces, typicall…

Word Embeddings