Loading…
A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts
In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research del...
Saved in:
Published in: | arXiv.org 2024-06 |
---|---|
Main Authors: | , , , , , , , |
Format: | Article |
Language: | English |
Subjects: | |
Online Access: | Get full text |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
cited_by | |
---|---|
cites | |
container_end_page | |
container_issue | |
container_start_page | |
container_title | arXiv.org |
container_volume | |
creator | Tripto, Nafis Irtiza Venkatraman, Saranya Macko, Dominik Moro, Robert Srba, Ivan Uchendu, Adaku Le, Thai Lee, Dongwon |
description | In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research delves into an intriguing question: Does a text retain its original authorship when it undergoes numerous paraphrasing iterations? Specifically, since Large Language Models (LLMs) have demonstrated remarkable proficiency in both the generation of original content and the modification of human-authored texts, a pivotal question emerges concerning the determination of authorship in instances where LLMs or similar paraphrasing tools are employed to rephrase the text--i.e., whether authorship should be attributed to the original human author or the AI-powered tool. Therefore, we embark on a philosophical voyage through the seas of language and authorship to unravel this intricate puzzle. Using a computational approach, we discover that the diminishing performance in text classification models, with each successive paraphrasing iteration, is closely associated with the extent of deviation from the original author's style, thus provoking a reconsideration of the current notion of authorship. |
format | article |
fullrecord | <record><control><sourceid>proquest</sourceid><recordid>TN_cdi_proquest_journals_2890143213</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2890143213</sourcerecordid><originalsourceid>FETCH-proquest_journals_28901432133</originalsourceid><addsrcrecordid>eNqNyk0KwjAQQOEgCBbtHQZcF9JJq9WdFH8WFQS7LwGnNkWSmmnA46vgAVy9xfcmIkKl0qTIEGciZu6llLhaY56rSBx3cO3MAK6FuiOmwFsogzcuMJSaib9y0V4Pndds7B2Mhao6J0ey5PVIN6jpNfJCTFv9YIp_nYvlYV-Xp2Tw7hmIx6Z3wdsPNVhsZJopTJX673oDPWw6MQ</addsrcrecordid><sourcetype>Aggregation Database</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype><pqid>2890143213</pqid></control><display><type>article</type><title>A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts</title><source>Publicly Available Content Database</source><creator>Tripto, Nafis Irtiza ; Venkatraman, Saranya ; Macko, Dominik ; Moro, Robert ; Srba, Ivan ; Uchendu, Adaku ; Le, Thai ; Lee, Dongwon</creator><creatorcontrib>Tripto, Nafis Irtiza ; Venkatraman, Saranya ; Macko, Dominik ; Moro, Robert ; Srba, Ivan ; Uchendu, Adaku ; Le, Thai ; Lee, Dongwon</creatorcontrib><description>In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research delves into an intriguing question: Does a text retain its original authorship when it undergoes numerous paraphrasing iterations? Specifically, since Large Language Models (LLMs) have demonstrated remarkable proficiency in both the generation of original content and the modification of human-authored texts, a pivotal question emerges concerning the determination of authorship in instances where LLMs or similar paraphrasing tools are employed to rephrase the text--i.e., whether authorship should be attributed to the original human author or the AI-powered tool. Therefore, we embark on a philosophical voyage through the seas of language and authorship to unravel this intricate puzzle. Using a computational approach, we discover that the diminishing performance in text classification models, with each successive paraphrasing iteration, is closely associated with the extent of deviation from the original author's style, thus provoking a reconsideration of the current notion of authorship.</description><identifier>EISSN: 2331-8422</identifier><language>eng</language><publisher>Ithaca: Cornell University Library, arXiv.org</publisher><subject>Authorship ; Large language models ; Questions ; Texts</subject><ispartof>arXiv.org, 2024-06</ispartof><rights>2024. This work is published under http://creativecommons.org/licenses/by/4.0/ (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><linktohtml>$$Uhttps://www.proquest.com/docview/2890143213?pq-origsite=primo$$EHTML$$P50$$Gproquest$$Hfree_for_read</linktohtml><link.rule.ids>776,780,25732,36991,44569</link.rule.ids></links><search><creatorcontrib>Tripto, Nafis Irtiza</creatorcontrib><creatorcontrib>Venkatraman, Saranya</creatorcontrib><creatorcontrib>Macko, Dominik</creatorcontrib><creatorcontrib>Moro, Robert</creatorcontrib><creatorcontrib>Srba, Ivan</creatorcontrib><creatorcontrib>Uchendu, Adaku</creatorcontrib><creatorcontrib>Le, Thai</creatorcontrib><creatorcontrib>Lee, Dongwon</creatorcontrib><title>A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts</title><title>arXiv.org</title><description>In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research delves into an intriguing question: Does a text retain its original authorship when it undergoes numerous paraphrasing iterations? Specifically, since Large Language Models (LLMs) have demonstrated remarkable proficiency in both the generation of original content and the modification of human-authored texts, a pivotal question emerges concerning the determination of authorship in instances where LLMs or similar paraphrasing tools are employed to rephrase the text--i.e., whether authorship should be attributed to the original human author or the AI-powered tool. Therefore, we embark on a philosophical voyage through the seas of language and authorship to unravel this intricate puzzle. Using a computational approach, we discover that the diminishing performance in text classification models, with each successive paraphrasing iteration, is closely associated with the extent of deviation from the original author's style, thus provoking a reconsideration of the current notion of authorship.</description><subject>Authorship</subject><subject>Large language models</subject><subject>Questions</subject><subject>Texts</subject><issn>2331-8422</issn><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2024</creationdate><recordtype>article</recordtype><sourceid>PIMPY</sourceid><recordid>eNqNyk0KwjAQQOEgCBbtHQZcF9JJq9WdFH8WFQS7LwGnNkWSmmnA46vgAVy9xfcmIkKl0qTIEGciZu6llLhaY56rSBx3cO3MAK6FuiOmwFsogzcuMJSaib9y0V4Pndds7B2Mhao6J0ey5PVIN6jpNfJCTFv9YIp_nYvlYV-Xp2Tw7hmIx6Z3wdsPNVhsZJopTJX673oDPWw6MQ</recordid><startdate>20240606</startdate><enddate>20240606</enddate><creator>Tripto, Nafis Irtiza</creator><creator>Venkatraman, Saranya</creator><creator>Macko, Dominik</creator><creator>Moro, Robert</creator><creator>Srba, Ivan</creator><creator>Uchendu, Adaku</creator><creator>Le, Thai</creator><creator>Lee, Dongwon</creator><general>Cornell University Library, arXiv.org</general><scope>8FE</scope><scope>8FG</scope><scope>ABJCF</scope><scope>ABUWG</scope><scope>AFKRA</scope><scope>AZQEC</scope><scope>BENPR</scope><scope>BGLVJ</scope><scope>CCPQU</scope><scope>DWQXO</scope><scope>HCIFZ</scope><scope>L6V</scope><scope>M7S</scope><scope>PIMPY</scope><scope>PQEST</scope><scope>PQQKQ</scope><scope>PQUKI</scope><scope>PRINS</scope><scope>PTHSS</scope></search><sort><creationdate>20240606</creationdate><title>A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts</title><author>Tripto, Nafis Irtiza ; Venkatraman, Saranya ; Macko, Dominik ; Moro, Robert ; Srba, Ivan ; Uchendu, Adaku ; Le, Thai ; Lee, Dongwon</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-proquest_journals_28901432133</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2024</creationdate><topic>Authorship</topic><topic>Large language models</topic><topic>Questions</topic><topic>Texts</topic><toplevel>online_resources</toplevel><creatorcontrib>Tripto, Nafis Irtiza</creatorcontrib><creatorcontrib>Venkatraman, Saranya</creatorcontrib><creatorcontrib>Macko, Dominik</creatorcontrib><creatorcontrib>Moro, Robert</creatorcontrib><creatorcontrib>Srba, Ivan</creatorcontrib><creatorcontrib>Uchendu, Adaku</creatorcontrib><creatorcontrib>Le, Thai</creatorcontrib><creatorcontrib>Lee, Dongwon</creatorcontrib><collection>ProQuest SciTech Collection</collection><collection>ProQuest Technology Collection</collection><collection>Materials Science & Engineering Collection</collection><collection>ProQuest Central (Alumni)</collection><collection>ProQuest Central</collection><collection>ProQuest Central Essentials</collection><collection>ProQuest Central</collection><collection>Technology Collection</collection><collection>ProQuest One Community College</collection><collection>ProQuest Central</collection><collection>SciTech Premium Collection</collection><collection>ProQuest Engineering Collection</collection><collection>Engineering Database</collection><collection>Publicly Available Content Database</collection><collection>ProQuest One Academic Eastern Edition (DO NOT USE)</collection><collection>ProQuest One Academic</collection><collection>ProQuest One Academic UKI Edition</collection><collection>ProQuest Central China</collection><collection>Engineering Collection</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext</fulltext></delivery><addata><au>Tripto, Nafis Irtiza</au><au>Venkatraman, Saranya</au><au>Macko, Dominik</au><au>Moro, Robert</au><au>Srba, Ivan</au><au>Uchendu, Adaku</au><au>Le, Thai</au><au>Lee, Dongwon</au><format>book</format><genre>document</genre><ristype>GEN</ristype><atitle>A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts</atitle><jtitle>arXiv.org</jtitle><date>2024-06-06</date><risdate>2024</risdate><eissn>2331-8422</eissn><abstract>In the realm of text manipulation and linguistic transformation, the question of authorship has been a subject of fascination and philosophical inquiry. Much like the Ship of Theseus paradox, which ponders whether a ship remains the same when each of its original planks is replaced, our research delves into an intriguing question: Does a text retain its original authorship when it undergoes numerous paraphrasing iterations? Specifically, since Large Language Models (LLMs) have demonstrated remarkable proficiency in both the generation of original content and the modification of human-authored texts, a pivotal question emerges concerning the determination of authorship in instances where LLMs or similar paraphrasing tools are employed to rephrase the text--i.e., whether authorship should be attributed to the original human author or the AI-powered tool. Therefore, we embark on a philosophical voyage through the seas of language and authorship to unravel this intricate puzzle. Using a computational approach, we discover that the diminishing performance in text classification models, with each successive paraphrasing iteration, is closely associated with the extent of deviation from the original author's style, thus provoking a reconsideration of the current notion of authorship.</abstract><cop>Ithaca</cop><pub>Cornell University Library, arXiv.org</pub><oa>free_for_read</oa></addata></record> |
fulltext | fulltext |
identifier | EISSN: 2331-8422 |
ispartof | arXiv.org, 2024-06 |
issn | 2331-8422 |
language | eng |
recordid | cdi_proquest_journals_2890143213 |
source | Publicly Available Content Database |
subjects | Authorship Large language models Questions Texts |
title | A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts |
url | http://sfxeu10.hosted.exlibrisgroup.com/loughborough?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-01-23T16%3A59%3A48IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-proquest&rft_val_fmt=info:ofi/fmt:kev:mtx:book&rft.genre=document&rft.atitle=A%20Ship%20of%20Theseus:%20Curious%20Cases%20of%20Paraphrasing%20in%20LLM-Generated%20Texts&rft.jtitle=arXiv.org&rft.au=Tripto,%20Nafis%20Irtiza&rft.date=2024-06-06&rft.eissn=2331-8422&rft_id=info:doi/&rft_dat=%3Cproquest%3E2890143213%3C/proquest%3E%3Cgrp_id%3Ecdi_FETCH-proquest_journals_28901432133%3C/grp_id%3E%3Coa%3E%3C/oa%3E%3Curl%3E%3C/url%3E&rft_id=info:oai/&rft_pqid=2890143213&rft_id=info:pmid/&rfr_iscdi=true |