Loading…

A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts

In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research del...

Full description

Saved in:

Bibliographic Details
Published in:	arXiv.org 2024-06
Main Authors:	Tripto, Nafis Irtiza, Venkatraman, Saranya, Macko, Dominik, Moro, Robert, Srba, Ivan, Uchendu, Adaku, Le, Thai, Lee, Dongwon
Format:	Article
Language:	English
Subjects:	Authorship Large language models Questions Texts
Online Access:	Get full text
Tags:	Add Tag No Tags, Be the first to tag this record!

cited_by
cites
container_end_page
container_issue
container_start_page
container_title	arXiv.org
container_volume
creator	Tripto, Nafis Irtiza Venkatraman, Saranya Macko, Dominik Moro, Robert Srba, Ivan Uchendu, Adaku Le, Thai Lee, Dongwon
description	In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research delves into an intriguing question: Does a text retain its original authorship when it undergoes numerous paraphrasing iterations? Specifically, since Large Language Models (LLMs) have demonstrated remarkable proficiency in both the generation of original content and the modification of human-authored texts, a pivotal question emerges concerning the determination of authorship in instances where LLMs or similar paraphrasing tools are employed to rephrase the text--i.e., whether authorship should be attributed to the original human author or the AI-powered tool. Therefore, we embark on a philosophical voyage through the seas of language and authorship to unravel this intricate puzzle. Using a computational approach, we discover that the diminishing performance in text classification models, with each successive paraphrasing iteration, is closely associated with the extent of deviation from the original author's style, thus provoking a reconsideration of the current notion of authorship.
format	article
fullrecord	<record><control><sourceid>proquest</sourceid><recordid>TN_cdi_proquest_journals_2890143213</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2890143213</sourcerecordid><originalsourceid>FETCH-proquest_journals_28901432133</originalsourceid><addsrcrecordid>eNqNyk0KwjAQQOEgCBbtHQZcF9JJq9WdFH8WFQS7LwGnNkWSmmnA46vgAVy9xfcmIkKl0qTIEGciZu6llLhaY56rSBx3cO3MAK6FuiOmwFsogzcuMJSaib9y0V4Pndds7B2Mhao6J0ey5PVIN6jpNfJCTFv9YIp_nYvlYV-Xp2Tw7hmIx6Z3wdsPNVhsZJopTJX673oDPWw6MQ</addsrcrecordid><sourcetype>Aggregation Database</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype><pqid>2890143213</pqid></control><display><type>article</type><title>A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts</title><source>Publicly Available Content Database</source><creator>Tripto, Nafis Irtiza ; Venkatraman, Saranya ; Macko, Dominik ; Moro, Robert ; Srba, Ivan ; Uchendu, Adaku ; Le, Thai ; Lee, Dongwon</creator><creatorcontrib>Tripto, Nafis Irtiza ; Venkatraman, Saranya ; Macko, Dominik ; Moro, Robert ; Srba, Ivan ; Uchendu, Adaku ; Le, Thai ; Lee, Dongwon</creatorcontrib><description>In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research delves into an intriguing question: Does a text retain its original authorship when it undergoes numerous paraphrasing iterations? Specifically, since Large Language Models (LLMs) have demonstrated remarkable proficiency in both the generation of original content and the modification of human-authored texts, a pivotal question emerges concerning the determination of authorship in instances where LLMs or similar paraphrasing tools are employed to rephrase the text--i.e., whether authorship should be attributed to the original human author or the AI-powered tool. Therefore, we embark on a philosophical voyage through the seas of language and authorship to unravel this intricate puzzle. Using a computational approach, we discover that the diminishing performance in text classification models, with each successive paraphrasing iteration, is closely associated with the extent of deviation from the original author's style, thus provoking a reconsideration of the current notion of authorship.</description><identifier>EISSN: 2331-8422</identifier><language>eng</language><publisher>Ithaca: Cornell University Library, arXiv.org</publisher><subject>Authorship ; Large language models ; Questions ; Texts</subject><ispartof>arXiv.org, 2024-06</ispartof><rights>2024. This work is published under http://creativecommons.org/licenses/by/4.0/ (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><linktohtml>$$Uhttps://www.proquest.com/docview/2890143213?pq-origsite=primo$$EHTML$$P50$$Gproquest$$Hfree_for_read</linktohtml><link.rule.ids>776,780,25732,36991,44569</link.rule.ids></links><search><creatorcontrib>Tripto, Nafis Irtiza</creatorcontrib><creatorcontrib>Venkatraman, Saranya</creatorcontrib><creatorcontrib>Macko, Dominik</creatorcontrib><creatorcontrib>Moro, Robert</creatorcontrib><creatorcontrib>Srba, Ivan</creatorcontrib><creatorcontrib>Uchendu, Adaku</creatorcontrib><creatorcontrib>Le, Thai</creatorcontrib><creatorcontrib>Lee, Dongwon</creatorcontrib><title>A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts</title><title>arXiv.org</title><description>In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research delves into an intriguing question: Does a text retain its original authorship when it undergoes numerous paraphrasing iterations? Specifically, since Large Language Models (LLMs) have demonstrated remarkable proficiency in both the generation of original content and the modification of human-authored texts, a pivotal question emerges concerning the determination of authorship in instances where LLMs or similar paraphrasing tools are employed to rephrase the text--i.e., whether authorship should be attributed to the original human author or the AI-powered tool. Therefore, we embark on a philosophical voyage through the seas of language and authorship to unravel this intricate puzzle. Using a computational approach, we discover that the diminishing performance in text classification models, with each successive paraphrasing iteration, is closely associated with the extent of deviation from the original author's style, thus provoking a reconsideration of the current notion of authorship.</description><subject>Authorship</subject><subject>Large language models</subject><subject>Questions</subject><subject>Texts</subject><issn>2331-8422</issn><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2024</creationdate><recordtype>article</recordtype><sourceid>PIMPY</sourceid><recordid>eNqNyk0KwjAQQOEgCBbtHQZcF9JJq9WdFH8WFQS7LwGnNkWSmmnA46vgAVy9xfcmIkKl0qTIEGciZu6llLhaY56rSBx3cO3MAK6FuiOmwFsogzcuMJSaib9y0V4Pndds7B2Mhao6J0ey5PVIN6jpNfJCTFv9YIp_nYvlYV-Xp2Tw7hmIx6Z3wdsPNVhsZJopTJX673oDPWw6MQ</recordid><startdate>20240606</startdate><enddate>20240606</enddate><creator>Tripto, Nafis Irtiza</creator><creator>Venkatraman, Saranya</creator><creator>Macko, Dominik</creator><creator>Moro, Robert</creator><creator>Srba, Ivan</creator><creator>Uchendu, Adaku</creator><creator>Le, Thai</creator><creator>Lee, Dongwon</creator><general>Cornell University Library, arXiv.org</general><scope>8FE</scope><scope>8FG</scope><scope>ABJCF</scope><scope>ABUWG</scope><scope>AFKRA</scope><scope>AZQEC</scope><scope>BENPR</scope><scope>BGLVJ</scope><scope>CCPQU</scope><scope>DWQXO</scope><scope>HCIFZ</scope><scope>L6V</scope><scope>M7S</scope><scope>PIMPY</scope><scope>PQEST</scope><scope>PQQKQ</scope><scope>PQUKI</scope><scope>PRINS</scope><scope>PTHSS</scope></search><sort><creationdate>20240606</creationdate><title>A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts</title><author>Tripto, Nafis Irtiza ; Venkatraman, Saranya ; Macko, Dominik ; Moro, Robert ; Srba, Ivan ; Uchendu, Adaku ; Le, Thai ; Lee, Dongwon</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-proquest_journals_28901432133</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2024</creationdate><topic>Authorship</topic><topic>Large language models</topic><topic>Questions</topic><topic>Texts</topic><toplevel>online_resources</toplevel><creatorcontrib>Tripto, Nafis Irtiza</creatorcontrib><creatorcontrib>Venkatraman, Saranya</creatorcontrib><creatorcontrib>Macko, Dominik</creatorcontrib><creatorcontrib>Moro, Robert</creatorcontrib><creatorcontrib>Srba, Ivan</creatorcontrib><creatorcontrib>Uchendu, Adaku</creatorcontrib><creatorcontrib>Le, Thai</creatorcontrib><creatorcontrib>Lee, Dongwon</creatorcontrib><collection>ProQuest SciTech Collection</collection><collection>ProQuest Technology Collection</collection><collection>Materials Science & Engineering Collection</collection><collection>ProQuest Central (Alumni)</collection><collection>ProQuest Central</collection><collection>ProQuest Central Essentials</collection><collection>ProQuest Central</collection><collection>Technology Collection</collection><collection>ProQuest One Community College</collection><collection>ProQuest Central</collection><collection>SciTech Premium Collection</collection><collection>ProQuest Engineering Collection</collection><collection>Engineering Database</collection><collection>Publicly Available Content Database</collection><collection>ProQuest One Academic Eastern Edition (DO NOT USE)</collection><collection>ProQuest One Academic</collection><collection>ProQuest One Academic UKI Edition</collection><collection>ProQuest Central China</collection><collection>Engineering Collection</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext</fulltext></delivery><addata><au>Tripto, Nafis Irtiza</au><au>Venkatraman, Saranya</au><au>Macko, Dominik</au><au>Moro, Robert</au><au>Srba, Ivan</au><au>Uchendu, Adaku</au><au>Le, Thai</au><au>Lee, Dongwon</au><format>book</format><genre>document</genre><ristype>GEN</ristype><atitle>A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts</atitle><jtitle>arXiv.org</jtitle><date>2024-06-06</date><risdate>2024</risdate><eissn>2331-8422</eissn><abstract>In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research delves into an intriguing question: Does a text retain its original authorship when it undergoes numerous paraphrasing iterations? Specifically, since Large Language Models (LLMs) have demonstrated remarkable proficiency in both the generation of original content and the modification of human-authored texts, a pivotal question emerges concerning the determination of authorship in instances where LLMs or similar paraphrasing tools are employed to rephrase the text--i.e., whether authorship should be attributed to the original human author or the AI-powered tool. Therefore, we embark on a philosophical voyage through the seas of language and authorship to unravel this intricate puzzle. Using a computational approach, we discover that the diminishing performance in text classification models, with each successive paraphrasing iteration, is closely associated with the extent of deviation from the original author's style, thus provoking a reconsideration of the current notion of authorship.</abstract><cop>Ithaca</cop><pub>Cornell University Library, arXiv.org</pub><oa>free_for_read</oa></addata></record>
fulltext	fulltext
identifier	EISSN: 2331-8422
ispartof	arXiv.org, 2024-06
issn	2331-8422
language	eng
recordid	cdi_proquest_journals_2890143213
source	Publicly Available Content Database
subjects	Authorship Large language models Questions Texts
title	A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts
url	http://sfxeu10.hosted.exlibrisgroup.com/loughborough?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-01-23T16%3A59%3A48IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-proquest&rft_val_fmt=info:ofi/fmt:kev:mtx:book&rft.genre=document&rft.atitle=A%20Ship%20of%20Theseus:%20Curious%20Cases%20of%20Paraphrasing%20in%20LLM-Generated%20Texts&rft.jtitle=arXiv.org&rft.au=Tripto,%20Nafis%20Irtiza&rft.date=2024-06-06&rft.eissn=2331-8422&rft_id=info:doi/&rft_dat=%3Cproquest%3E2890143213%3C/proquest%3E%3Cgrp_id%3Ecdi_FETCH-proquest_journals_28901432133%3C/grp_id%3E%3Coa%3E%3C/oa%3E%3Curl%3E%3C/url%3E&rft_id=info:oai/&rft_pqid=2890143213&rft_id=info:pmid/&rfr_iscdi=true