paper-with-me

홈 › Papers

The Impact of Copyrighted Material on Large Language Models: A Norwegian Perspective

2024-12-12 · Javier de la Rosa, Vladislav Mikhailov, Lemei Zhang, Freddy Wetjen, David Samuel, Peng Liu, Rolv-Arild Braaten, Petter Mæhlum, Magnus Breder Birkenes, Andrey Kutuzov, Tita Enstad, Hans Christian Farsethås, Svein Arne Brygfjeld, Jon Atle Gulla, Stephan Oepen, Erik Velldal, Wilfred Østgulen, Liljia Øvrelid, Aslak Sira Myhre

The use of copyrighted materials in training language models raises critical legal and ethical questions. This paper presents a framework for and the results of empirically assessing the impact of publisher-controlled copyrighted corpora on the performance of generative large language models (LLMs) for Norwegian. When evaluated on a diverse set of tasks, we found that adding both books and newspapers to the data mixture of LLMs tend to improve their performance, while the addition of fiction works seems to be detrimental. Our experiments could inform the creation of a compensation scheme for authors whose works contribute to AI development.

📄 PDF Abstract BibTeX arXiv:2412.09460

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Copyright Violations and Large Language Models

2023-10-20 · Antonia Karamolegkou, Jiaang Li, Li Zhou, Anders Søgaard

Language models may memorize more than just facts, including entire chunks of texts seen during training. Fair use exemptions to copyright laws typically allow for limited use of copyrighted material without permission f…

Memorization

Measuring Copyright Risks of Large Language Model via Partial Information Probing

2024-09-20 · Weijie Zhao, Huajie Shao, Zhaozhuo Xu, Suzhen Duan 외

Exploring the data sources used to train Large Language Models (LLMs) is a crucial direction in investigating potential copyright infringement by these models. While this approach can identify the possible use of copyrig…

Language ModelingLanguage ModellingLarge Language Model

Can Watermarking Large Language Models Prevent Copyrighted Text Generation and Hide Training Data?

2024-07-24 · Michael-Andrei Panaitescu-Liess, Zora Che, Bang An, Yuancheng Xu 외

Large Language Models (LLMs) have demonstrated impressive capabilities in generating diverse and contextually rich text. However, concerns regarding copyright infringement arise as LLMs may inadvertently produce copyrigh…

Text Generation

Beyond English: Unveiling Multilingual Bias in LLM Copyright Compliance

2025-02-14 · Yupeng Chen, XiaoYu Zhang, Yixian Huang, Qian Xie

Large Language Models (LLMs) have raised significant concerns regarding the fair use of copyright-protected content. While prior studies have examined the extent to which LLMs reproduce copyrighted materials, they have p…

NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark

2025-04-10 · Vladislav Mikhailov, Tita Enstad, David Samuel, Hans Christian Farsethås 외

This paper introduces NorEval, a new and comprehensive evaluation suite for large-scale standardized benchmarking of Norwegian generative language models (LMs). NorEval consists of 24 high-quality human-created datasets …

Benchmarking