paper-with-me

홈 › Papers

Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts

2025-03-23 · Beining Xu, Arkaitz Zubiaga

Large Language Models (LLMs) have demonstrated exceptional performance on a range of downstream NLP tasks by generating text that closely resembles human writing. However, the ease of achieving this similarity raises concerns from potential malicious uses at scale by bad actors, as LLM-generated text becomes increasingly difficult to discern from human text. Although detection methods have been developed to address this issue, bad actors can further manipulate LLM-generated texts to make them less detectable. In this work, we study how further editing texts with Reinforcement Learning from Human Feedback (RLHF), which aligns model outputs with human preferences, affects (a) the quality of generated texts for two tasks, and (b) the performance of LLM-generated text detectors, looking at both training-based and zero-shot detection methods. Although RLHF improves the quality of LLM-generated texts, we find that it also tends to produce more detectable, lengthy, and repetitive outputs. Additionally, we observe that training-based detectors are vulnerable to short texts and to texts that incorporate code, whereas zero-shot detectors exhibit greater robustness.

📄 PDF Abstract BibTeX arXiv:2503.17965

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language Models

2024-02-28 · Mingjia Huo, Sai Ashish Somayajula, Youwei Liang, Ruisi Zhang 외

Large language models generate high-quality responses with potential misinformation, underscoring the need for regulation by distinguishing AI-generated and human-written texts. Watermarking is pivotal in this context, w…

Misinformation

Progressively Label Enhancement for Large Language Model Alignment

2024-08-05 · Biao Liu, Ning Xu, Xin Geng

Large Language Models (LLM) alignment aims to prevent models from producing content that misaligns with human expectations, which can lead to ethical and legal concerns. In the last few years, Reinforcement Learning from…

Language ModelingLanguage ModellingLarge Language Modelmodel

Toward Stronger Code Watermarking: A Grammar-Driven Approach to Optimizing the Trade-off Between Quality and Detectability

2026-07-11 · Licheng Yu, Aiwei Liu, Songze Li arxiv

With the rapid development of Large Language Models (LLMs), text watermarking has emerged as a crucial technique for identifying machine-generated content. However, directly applying existing logits-based watermarking me…

Code Generation

Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization

2025-10-13 · Chenrui Wang, Junyi Shu, Billy Chiu, Yu Li 외 arxiv

The rapid development of LLMs has raised concerns about their potential misuse, leading to various watermarking schemes that typically offer high detectability. However, existing watermarking techniques often face trade-…

Aligning Neural Machine Translation Models: Human Feedback in Training and Inference

2023-11-15 · Miguel Moura Ramos, Patrick Fernandes, António Farinhas, André F. T. Martins

Reinforcement learning from human feedback (RLHF) is a recent technique to improve the quality of the text generated by a language model, making it closer to what humans would generate. A core ingredient in RLHF's succes…

Language ModelingLanguage ModellingMachine TranslationReranking+1