paper-with-me

홈 › Papers

Quantifying Association Capabilities of Large Language Models and Its Implications on Privacy Leakage

2023-05-22 · Hanyin Shao, Jie Huang, Shen Zheng, Kevin Chen-Chuan Chang

The advancement of large language models (LLMs) brings notable improvements across various applications, while simultaneously raising concerns about potential private data exposure. One notable capability of LLMs is their ability to form associations between different pieces of information, but this raises concerns when it comes to personally identifiable information (PII). This paper delves into the association capabilities of language models, aiming to uncover the factors that influence their proficiency in associating information. Our study reveals that as models scale up, their capacity to associate entities/information intensifies, particularly when target pairs demonstrate shorter co-occurrence distances or higher co-occurrence frequencies. However, there is a distinct performance gap when associating commonsense knowledge versus PII, with the latter showing lower accuracy. Despite the proportion of accurately predicted PII being relatively small, LLMs still demonstrate the capability to predict specific instances of email addresses and phone numbers when provided with appropriate prompts. These findings underscore the potential risk to PII confidentiality posed by the evolving capabilities of LLMs, especially as they continue to expand in scale and power.

📄 PDF Abstract BibTeX arXiv:2305.12707

Code (1)

hanyins/lm_association_quantification 공식 구현

Similar Papers 제목 키워드 기반

Sycophancy in Large Language Models: Causes and Mitigations

2024-11-22 · Lars Malmqvist

Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing tasks. However, their tendency to exhibit sycophantic behavior - excessively agreeing with or flat…

Hallucination

WinoGrande: An Adversarial Winograd Schema Challenge at Scale

2019-07-24 · Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin Choi

The Winograd Schema Challenge (WSC) (Levesque, Davis, and Morgenstern 2011), a benchmark for commonsense reasoning, is a set of 273 expert-crafted pronoun resolution problems originally designed to be unsolvable for stat…

Common Sense ReasoningCoreference ResolutionQuestion AnsweringTransfer Learning+1

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models

2025-01-28 · Zeping Min, Xinshang Wang

We introduce a novel index, the Distribution of Cosine Similarity (DOCS), for quantitatively assessing the similarity between weight matrices in Large Language Models (LLMs), aiming to facilitate the analysis of their co…

She Elicits Requirements and He Tests: Software Engineering Gender Bias in Large Language Models

2023-03-17 · Christoph Treude, Hideaki Hata

Implicit gender bias in software development is a well-documented issue, such as the association of technical roles with men. To address this bias, it is important to understand it in more detail. This study uses data mi…

Quantifying and Analyzing Entity-level Memorization in Large Language Models

2023-08-30 · Zhenhong Zhou, Jiuyang Xiang, Chaomeng Chen, Sen Su

Large language models (LLMs) have been proven capable of memorizing their training data, which can be extracted through specifically designed prompts. As the scale of datasets continues to grow, privacy risks arising fro…

Language ModelingLanguage ModellingMemorizationProbing Language Models