paper-with-me

홈 › Papers

Interrogating LLM design under a fair learning doctrine

2025-02-22 · Johnny Tian-Zheng Wei, Maggie Wang, Ameya Godbole, Jonathan H. Choi, Robin Jia

The current discourse on large language models (LLMs) and copyright largely takes a "behavioral" perspective, focusing on model outputs and evaluating whether they are substantially similar to training data. However, substantial similarity is difficult to define algorithmically and a narrow focus on model outputs is insufficient to address all copyright risks. In this interdisciplinary work, we take a complementary "structural" perspective and shift our focus to how LLMs are trained. We operationalize a notion of "fair learning" by measuring whether any training decision substantially affected the model's memorization. As a case study, we deconstruct Pythia, an open-source LLM, and demonstrate the use of causal and correlational analyses to make factual determinations about Pythia's training decisions. By proposing a legal standard for fair learning and connecting memorization analyses to this standard, we identify how judges may advance the goals of copyright law through adjudication. Finally, we discuss how a fair learning standard might evolve to enhance its clarity by becoming more rule-like and incorporating external technical guidelines.

📄 PDF Abstract BibTeX arXiv:2502.16290

Code (0)

등록된 구현이 없습니다.

Tasks

Memorization

Methods 이 논문이 사용한 방법론

Pythia Pythia is a suite of decoder-only autoregressive language models all trained on public data seen in the exact same order and ranging in size from 70M to 12B parameters. The…
Focus 설명 없음

Similar Papers 제목 키워드 기반

Towards Substantive Conceptions of Algorithmic Fairness: Normative Guidance from Equal Opportunity Doctrines

2022-07-06 · Falaah Arif Khan, Eleni Manis, Julia Stoyanovich

In this work we use Equal Oppportunity (EO) doctrines from political philosophy to make explicit the normative judgements embedded in different conceptions of algorithmic fairness. We contrast formal EO approaches that n…

FairnessPhilosophy

LUCID: Exposing Algorithmic Bias through Inverse Design

2022-08-26 · Carmen Mazijn, Carina Prunkl, Andres Algaba, Jan Danckaert 외

AI systems can create, propagate, support, and automate bias in decision-making processes. To mitigate biased decisions, we both need to understand the origin of the bias and define what it means for an algorithm to make…

Decision MakingFairness

Towards Fair Deep Clustering With Multi-State Protected Variables

2019-01-29 · Bokun Wang, Ian Davidson

Fair clustering under the disparate impact doctrine requires that population of each protected group should be approximately equal in every cluster. Previous work investigated a difficult-to-scale pre-processing step for…

AttributeClusteringDeep ClusteringFairness

Fair Clustering Through Fairlets

2018-02-15 · NeurIPS 2017 12 · Flavio Chierichetti, Ravi Kumar, Silvio Lattanzi, Sergei Vassilvitskii

We study the question of fair clustering under the {\em disparate impact} doctrine, where each protected class must have approximately equal representation in every cluster. We formulate the fair clustering problem under…

Clustering

Fairness as Equality of Opportunity: Normative Guidance from Political Philosophy

2021-06-15 · Falaah Arif Khan, Eleni Manis, Julia Stoyanovich

Recent interest in codifying fairness in Automated Decision Systems (ADS) has resulted in a wide range of formulations of what it means for an algorithmic system to be fair. Most of these propositions are inspired by, bu…

EthicsFairnessPhilosophy