paper-with-me

Papers

Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators

2023-07-08 · Andreas Liesenfeld, Alianda Lopez, Mark Dingemanse

Large language models that exhibit instruction-following behaviour represent one of the biggest recent upheavals in conversational interfaces, a trend in large part fuelled by the release of OpenAI's ChatGPT, a proprietary large language model for text generation fine-tuned through reinforcement learning from human feedback (LLM+RLHF). We review the risks of relying on proprietary software and survey the first crop of open-source projects of comparable architecture and functionality. The main contribution of this paper is to show that openness is differentiated, and to offer scientific documentation of degrees of openness in this fast-moving field. We evaluate projects in terms of openness of code, training data, model weights, RLHF data, licensing, scientific documentation, and access methods. We find that while there is a fast-growing list of projects billing themselves as 'open source', many inherit undocumented data of dubious legality, few share the all-important instruction-tuning (a key site where human annotation labour is involved), and careful scientific documentation is exceedingly rare. Degrees of openness are relevant to fairness and accountability at all points, from data collection and curation to model architecture, and from training and fine-tuning to release and deployment.

📄 PDF Abstract BibTeX arXiv:2307.05532

Code (1)

opening-up-chatgpt/opening-up-chatgpt.github.io 공식 구현

Tasks

FairnessInstruction FollowingLanguage ModelingLanguage ModellingLarge Language ModelText Generation

Similar Papers 제목 키워드 기반

Comprehensive Analysis of Transparency and Accessibility of ChatGPT, DeepSeek, And other SoTA Large Language Models

2025-02-21 · Ranjan Sapkota, Shaina Raza, Manoj Karkee

Despite increasing discussions on open-source Artificial Intelligence (AI), existing research lacks a discussion on the transparency and accessibility of state-of-the-art (SoTA) Large Language Models (LLMs). The Open Sou…

Domain Adaptation

Opening the Scope of Openness in AI

2025-05-09 · Tamara Paris, AJung Moon, Jin Guo

The concept of openness in AI has so far been heavily inspired by the definition and community practice of open source software. This positions openness in AI as having positive connotations; it introduces assumptions of…

MusGO: A Community-Driven Framework For Assessing Openness in Music-Generative AI

2025-07-04 · Roser Batlle-Roca, Laura Ibáñez-Martínez, Xavier Serra, Emilia Gómez 외 arxiv

Since 2023, generative AI has rapidly advanced in the music domain. Despite significant technological advancements, music-generative models raise critical ethical challenges, including a lack of transparency and accounta…

Information Retrieval

Inference is All You Need: Self Example Retriever for Cross-domain Dialogue State Tracking with ChatGPT

2024-09-10 · Jihyun Lee, Gary Geunbae Lee

Traditional dialogue state tracking approaches heavily rely on extensive training data and handcrafted features, limiting their scalability and adaptability to new domains. In this paper, we propose a novel method that l…

AllDialogue State TrackingIn-Context LearningTransfer Learning

The Model Openness Framework: Promoting Completeness and Openness for Reproducibility, Transparency, and Usability in Artificial Intelligence

2024-03-20 · Matt White, Ibrahim Haddad, Cailean Osborne, Xiao-Yang Yanglet Liu 외

Generative artificial intelligence (AI) offers numerous opportunities for research and innovation, but its commercialization has raised concerns about the transparency and safety of frontier AI models. Most models lack t…