paper-with-me

홈 › Papers

Are Paraphrases Generated by Large Language Models Invertible?

2024-10-29 · Rafael Rivera Soto, Barry Chen, Nicholas Andrews

Large language models can produce highly fluent paraphrases while retaining much of the original meaning. While this capability has a variety of helpful applications, it may also be abused by bad actors, for example to plagiarize content or to conceal their identity. This motivates us to consider the problem of paraphrase inversion: given a paraphrased document, attempt to recover the original text. To explore the feasibility of this task, we fine-tune paraphrase inversion models, both with and without additional author-specific context to help guide the inversion process. We explore two approaches to author-specific inversion: one using in-context examples of the target author's writing, and another using learned style representations that capture distinctive features of the author's style. We show that, when starting from paraphrased machine-generated text, we can recover significant portions of the document using a learned inversion model. When starting from human-written text, the variety of source writing styles poses a greater challenge for invertability. However, even when the original tokens can't be recovered, we find the inverted text is stylistically similar to the original, which significantly improves the performance of plagiarism detectors and authorship identification systems that rely on stylistic markers.

📄 PDF Abstract BibTeX arXiv:2410.21637

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding the Effects of Human-written Paraphrases in LLM-generated Text Detection

2024-11-06 · Hiu Ting Lau, Arkaitz Zubiaga

Natural Language Generation has been rapidly developing with the advent of large language models (LLMs). While their usage has sparked significant attention from the general public, it is important for readers to be awar…

LLM-generated Text DetectionText DetectionText Generation

Efficient Zero-Shot Semantic Parsing with Paraphrasing from Pretrained Language Models

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Building a domain-specific semantic parser with little or no domain-specific training data remains a challenging task. Previous work has shown that crowdsourced paraphrases of synthetic (grammar-generated) utterances can…

Semantic Parsing

Generating Syntactically Controlled Paraphrases without Using Annotated Parallel Pairs

2021-01-26 · EACL 2021 2 · Kuan-Hao Huang, Kai-Wei Chang

Paraphrase generation plays an essential role in natural language process (NLP), and it has many downstream applications. However, training supervised paraphrase models requires many annotated paraphrase pairs, which are…

Data AugmentationDecoderDisentanglementParaphrase Generation+1

Unsupervised Paraphrase Generation using Pre-trained Language Models

2020-06-09 · Chaitra Hegde, Shrikumar Patil

Large scale Pre-trained Language Models have proven to be very powerful approach in various Natural language tasks. OpenAI's GPT-2 \cite{radford2019language} is notable for its capability to generate fluent, well formula…

Data AugmentationParaphrase Generation

How Large Language Models are Transforming Machine-Paraphrased Plagiarism

2022-10-07 · Jan Philip Wahle, Terry Ruas, Frederic Kirstein, Bela Gipp

The recent success of large language models for text generation poses a severe threat to academic integrity, as plagiarists can generate realistic paraphrases indistinguishable from original work. However, the role of la…

ArticlesParaphrase GenerationText Generation