paper-with-me

Papers

Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing

2026-04-25 · Arthur Amalvy, Vincent Labatut, Xavier Bost, Hen-Hsen Huang arxiv

While annotated corpora are crucial in the field of natural language processing (NLP), those containing copyrighted material are difficult to exchange among researchers. Yet, such corpora are necessary to fully represent the diversity of data found in the wild in the context of NLP tasks. We tackle this issue by proposing a method to lawfully and publicly share the annotations of copyrighted literary texts. The corpus creator shares the annotations in clear, along with a non-reversible hashed version of the source material. The corpus user must own the source material, and apply the same hash function to their own tokens, in order to match them to the shared annotations. Crucially, our method is robust to reasonable divergences in the version of the copyrighted data owned by the user. As an illustration, we present alignment experiments on different editions of novels. Our results show that our method is able to correctly align 98.7 to 99.79% of tokens depending on the novel, provided the user version is sufficiently close to the corpus creator's version. We publicly release novelshare, a Python implementation of our method.

📄 PDF Abstract BibTeX arXiv:2604.23412

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Language Models for German Text Simplification: Overcoming Parallel Data Scarcity through Style-specific Pre-training

2023-05-22 · Miriam Anschütz, Joshua Oehms, Thomas Wimmer, Bartłomiej Jezierski 외

Automatic text simplification systems help to reduce textual information barriers on the internet. However, for languages other than English, only few parallel data to train these systems exists. We propose a two-step ap…

Text Simplification

Training-time Neuron Alignment through Permutation Subspace for Improving Linear Mode Connectivity and Model Fusion

2024-02-02 · Zexi Li, Zhiqi Li, Jie Lin, Tao Shen 외

In deep learning, stochastic gradient descent often yields functionally similar yet widely scattered solutions in the weight space even under the same initialization, causing barriers in the Linear Mode Connectivity (LMC…

Federated LearningLinear Mode Connectivity

DIS-CO: Discovering Copyrighted Content in VLMs Training Data

2025-02-24 · André V. Duarte, Xuandong Zhao, Arlindo L. Oliveira, Lei LI

How can we verify whether copyrighted content was used to train a large vision-language model (VLM) without direct access to its training data? Motivated by the hypothesis that a VLM is able to recognize images from its …

Language ModelingLanguage Modelling

Copyright Violations and Large Language Models

2023-10-20 · Antonia Karamolegkou, Jiaang Li, Li Zhou, Anders Søgaard

Language models may memorize more than just facts, including entire chunks of texts seen during training. Fair use exemptions to copyright laws typically allow for limited use of copyrighted material without permission f…

Memorization

JaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus

2022-11-29 · Tomohiko Nakamura, Shinnosuke Takamichi, Naoko Tanji, Satoru Fukayama 외

We construct a corpus of Japanese a cappella vocal ensembles (jaCappella corpus) for vocal ensemble separation and synthesis. It consists of 35 copyright-cleared vocal ensemble songs and their audio recordings of individ…

Vocal ensemble separation