paper-with-me

홈 › Papers

Idiosyncrasies in Large Language Models

2025-02-17 · MingJie Sun, Yida Yin, Zhiqiu Xu, J. Zico Kolter, Zhuang Liu

In this work, we unveil and study idiosyncrasies in Large Language Models (LLMs) -- unique patterns in their outputs that can be used to distinguish the models. To do so, we consider a simple classification task: given a particular text output, the objective is to predict the source LLM that generates the text. We evaluate this synthetic task across various groups of LLMs and find that simply fine-tuning existing text embedding models on LLM-generated texts yields excellent classification accuracy. Notably, we achieve 97.1% accuracy on held-out validation data in the five-way classification problem involving ChatGPT, Claude, Grok, Gemini, and DeepSeek. Our further investigation reveals that these idiosyncrasies are rooted in word-level distributions. These patterns persist even when the texts are rewritten, translated, or summarized by an external LLM, suggesting that they are also encoded in the semantic content. Additionally, we leverage LLM as judges to generate detailed, open-ended descriptions of each model's idiosyncrasies. Finally, we discuss the broader implications of our findings, particularly for training on synthetic data and inferring model similarity. Code is available at https://github.com/locuslab/llm-idiosyncrasies.

📄 PDF Abstract BibTeX arXiv:2502.12150

Code (1)

locuslab/llm-idiosyncrasies 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Learnings from Federated Learning in the Real world

2022-02-08 · Christophe Dupuy, Tanya G. Roosta, Leo Long, Clement Chung 외

Federated Learning (FL) applied to real world data may suffer from several idiosyncrasies. One such idiosyncrasy is the data distribution across devices. Data across devices could be distributed such that there are some …

Federated LearningNatural Language Understanding

Idiosyncrasies and challenges of data driven learning in electronic trading

2018-11-30

We outline the idiosyncrasies of neural information processing and machine learning in quantitative finance. We also present some of the approaches we take towards solving the fundamental challenges we face.

BIG-bench Machine Learning

Asymmetric Idiosyncrasies in Multimodal Models

2026-02-26 · Muzi Tao, Chufan Shi, Huijuan Wang, Shengbang Tong 외 arxiv

In this work, we study idiosyncrasies in the caption models and their downstream impact on text-to-image models. We design a systematic analysis: given either a generated caption or the corresponding image, we train neur…

Text Classification

Not an Interlingua, But Close: Comparison of English AMRs to Chinese and Czech

2014-05-01 · LREC 2014 5 · Nianwen Xue, Ond{\v{r}}ej Bojar, Jan Haji{\v{c}}, Martha Palmer 외

Abstract Meaning Representations (AMRs) are rooted, directional and labeled graphs that abstract away from morpho-syntactic idiosyncrasies such as word category (verbs and nouns), word order, and function words (determin…

Machine TranslationSemantic ParsingSemantic Role LabelingTranslation

Cultural Communication Idiosyncrasies in Human-Computer Interaction

2016-09-01 · WS 2016 9 · Juliana Miehle, Koichiro Yoshino, Louisa Pragst, Stefan Ultes 외
Spoken Dialogue Systems