paper-with-me

홈 › Papers

Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs

2025-02-26 · Christoph Schuhmann, Gollam Rabby, Ameya Prabhu, Tawsif Ahmed, Andreas Hochlehnert, Huu Nguyen, Nick Akinci, Ludwig Schmidt, Robert Kaczmarczyk, Sören Auer, Jenia Jitsev, Matthias Bethge

Paywalls, licenses and copyright rules often restrict the broad dissemination and reuse of scientific knowledge. We take the position that it is both legally and technically feasible to extract the scientific knowledge in scholarly texts. Current methods, like text embeddings, fail to reliably preserve factual content, and simple paraphrasing may not be legally sound. We propose a new idea for the community to adopt: convert scholarly documents into knowledge preserving, but style agnostic representations we term Knowledge Units using LLMs. These units use structured data capturing entities, attributes and relationships without stylistic content. We provide evidence that Knowledge Units (1) form a legally defensible framework for sharing knowledge from copyrighted research texts, based on legal analyses of German copyright law and U.S. Fair Use doctrine, and (2) preserve most (~95\%) factual knowledge from original text, measured by MCQ performance on facts from the original copyrighted text across four research domains. Freeing scientific knowledge from copyright promises transformative benefits for scientific research and education by allowing language models to reuse important facts from copyrighted text. To support this, we share open-source tools for converting research documents into Knowledge Units. Overall, our work posits the feasibility of democratizing access to scientific knowledge while respecting copyright.

📄 PDF Abstract BibTeX arXiv:2502.19413

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Alexandria: Extensible Framework for Rapid Exploration of Social Media

2015-07-23 · Heath Fenno F. III, Hull Richard, Khabiri Elham, Riemer Matthew 외

The Alexandria system under development at IBM Research provides an extensible framework and platform for supporting a variety of big-data analytics and visualizations. The system is currently focused on enabling rapid e…

Enterprise Alexandria: Online High-Precision Enterprise Knowledge Base Construction with Typed Entities

2021-06-22 · AKBC 2021 10 · John Winn, Matteo Venanzi, Tom Minka, Ivan Korostelev 외

We present Enterprise Alexandria, one of the core AI technologies behind Microsoft Viva Topics. Enterprise Alexandria is a new system for automatically constructing a knowledge base with high-precision and typed entities…

Knowledge Base Construction

Alexandria: Unsupervised High-Precision Knowledge Base Construction using a Probabilistic Program

2018-11-17 · AKBC 2019 · John Winn, John Guiver, Sam Webster, Yordan Zaykov 외

Creating a knowledge base that is accurate, up-to-date and complete remains a significant challenge despite substantial efforts in automated knowledge base construction. In this paper, we present Alexandria -- a system …

Knowledge Base ConstructionVocal Bursts Intensity Prediction

Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs

2026-01-19 · Abdellah El Mekki, Samar M. Magdy, Houdaifa Atou, Ruwa AbuHweidi 외 arxiv

Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than Modern Standard Arabic (MSA). Despite this, machine translation (MT) systems often generalize poorly to dialect…

Machine Translation

Echoes from Alexandria: A Large Resource for Multilingual Book Summarization

2023-06-07 · Alessandro Scirè, Simone Conia, Simone Ciciliano, Roberto Navigli

In recent years, research in text summarization has mainly focused on the news domain, where texts are typically short and have strong layout features. The task of full-book summarization presents additional challenges w…

Book summarizationText Summarization