Linguistic Fingerprints for Authorship Attribution in Human–AI Collaborative Texts
Abstract views: 63
,
PDF downloads: 21
Abstract
Authors leave distinctive linguistic traces that reflect their identity through consistent writing styles, particularly in morphosyntactic patterns and lexical choices. However, the increasing use of AI writing tools challenges authorship attribution because machine-generated text can imitate or obscure individual writing characteristics. This study investigates linguistic features that effectively identify authors and differentiate human-written from AI-generated texts. A corpus comprising 2,074,125 tokens and 63,414 word types was compiled from collaborative digital platforms, including instant messaging and social media. Lexical and stylistic features were extracted to develop hybrid authorship-classification models, while N-gram tracing was used to identify salient patterns. The findings demonstrate that lexical choice is the most reliable indicator for distinguishing human and AI-generated texts. Character-level N-gram analysis further demonstrates that authorship can be identified through delicate patterns involving letters, capitalization, punctuation, and other non-alphabetic characters. Diction appeared as the strongest factor in differentiating individual authors. These results enhance the reliability of authorship attribution methods and provide valuable insights for forensic investigations of digitally mediated communication involving human–AI interaction.
Downloads
References
Abreu Rodrigues, S., & Sousa Silva, R. (2022). A Forensic Authorship Analysis of Threats. RevSALUS - Revista Científica Da Rede Académica Das Ciências Da Saúde Da Lusofonia, 4(Sup),p. 98–99. https://doi.org/10.51126/revsalus.v4isup.324
Alamleh, H., Alqahtani, A. A. S., & Elsaid, A. (2023). Distinguishing Human-Written and ChatGPT-Generated Text Using Machine Learning. 2023 Systems and Information Engineering Design Symposium, SIEDS 2023. https://doi.org/10.1109/SIEDS58326.2023.10137767
Aribowo, A. S., Basiron, H., Yusof, N. F. A., & Khomsah, S. (2021). Cross-Domain Sentiment Analysis Model on Indonesian YouTube Comment. International Journal of Advances in Intelligent Informatics, 7(1), 12–25. https://doi.org/10.26555/ijain.v7i1.554
Baker, P. (2010). Title of chapter. In J. Smith (Ed.), Sociolinguistics and Corpus Linguistics (pp. 25–42). Edinburgh University Press.
Banga, R., Bhardwaj, A., Peng, S. L., & Shrivastava, G. (2018). Authorship Attribution for Online Social Media. In Social Network Analytics for Contemporary Business Organizations (pp. 141–165). https://doi.org/10.4018/978-1-5225-5097-6.ch008
Belvisi, N. M. S., Muhammad, N., & Alonso-Fernandez, F. (2020). Forensic Authorship Analysis of Microblogging Texts Using N-Grams and Stylometric Features. In 2020, the 8th International Workshop on Biometrics and Forensics (IWBF) (pp. 1–6). IEEE. https://doi.org/10.1109/IWBF49977.2020.9107956
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623. https://doi.org/10.1145/3442188.3445922
Bevendorff, J., Ghanem, B., Giachanou, A., Kestemont, M., Manjavacas, E., Markov, I., Mayerl, M., Potthast, M., Rangel, F., Rosso, P., Specht, G., Stamatatos, E., Stein, B., Wiegmann, M., & Zangerle, E. (2020). Overview of Pan 2020: Authorship Verification, Celebrity Profiling, Profiling Fake News Spreaders on Twitter, and Style Change Detection. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 12260 Lncs. https://doi.org/10.1007/978-3-030-58219-7_25
Beyer, K. (2014). Urban Language Research in South Africa: Achievements and Challenges. Southern African Linguistics and Applied Language Studies, 32(2), 247–254. https://doi.org/10.2989/16073614.2014.992643
BRIN. Cillco Indonesia Corpora: Corpus of Indonesia Language, Literature, and Community 2024. https://cillco-prototype.brin.go.id
Caliskan, A., Bryson, J. J., & Narayanan, A. (2017). Semantics Derived Automatically From Language Corpora Contain Human-Like Biases. Science, 356(6334), 183–186. https://doi.org/10.1126/science.aal4230
Chen, B., Ding, X., Zhao, Y., Fu, B., Lin, T., Qin, B., & Liu, T. (2024). Text Difficulty Study: Do Machines Behave the Same as Humans Regarding Text Difficulty? Machine Intelligence Research, 21(2), 283–293. https://doi.org/10.1007/s11633-023-1424-x
Dou, Y., Forbes, M., Koncel-Kedziorski, R., Smith, N. A., & Choi, Y. (2022). Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1, 7250–7274. https://doi.org/10.18653/v1/2022.acl-long.501
Dugan, L., Ippolito, D., Kirubarajan, A., Shi, S., & Callison-Burch, C. (2023). Real or Fake Text?: Investigating Human Ability to Detect Boundaries Between Human-Written and Machine-Generated Text. Proceedings of the AAAI Conference on Artificial Intelligence, 37(11), 12763–12771. https://doi.org/10.1609/aaai.v37i11.26501
Fedotova, A., Romanov, A., Kurtukova, A., & Shelupanov, A. (2022). Authorship Attribution of Social Media and Literary Russian-Language Texts Using Machine Learning Methods and Feature Selection. Future Internet, 14(1), 4. https://doi.org/10.3390/fi14010004
García-Díaz, J. A., Colomo-Palacios, R., & Valencia-García, R. (2022). Psychographic Traits Identification Based on Political Ideology: An Author Analysis Study on Spanish Politicians’ Tweets Posted in 2020. Future Generation Computer Systems, 130, 59–74. https://doi.org/10.1016/j.future.2021.12.011
Grant, T. (2007). Quantifying Evidence in Forensic Authorship Analysis. International Journal of Speech, Language and the Law, 14(1), 1–25. https://doi.org/10.1558/ijsll.v14i1.1
Grieve, J. (2023). Register Variation Explains Stylometric Authorship Analysis. Corpus Linguistics and Linguistic Theory, 19(1), 47–77. https://doi.org/10.1515/cllt-2022-0040
Grieve, J., Clarke, I., Chiang, E., Gideon, H., Heini, A., Nini, A., & Waibel, E. (2019). Attributing the Bixby Letter Using N-Gram Tracing. Digital Scholarship in the Humanities, 34(3), 493–512. https://doi.org/10.1093/llc/fqy042
Habibzadeh, F. (2023). GPTZero Performance in Identifying Artificial Intelligence-Generated Medical Texts: A Preliminary Study. Journal of Korean Medical Science, 38(38), 1–12. https://doi.org/10.3346/jkms.2023.38.e319
Holmes, D. I. (1998). The Evolution of Stylometry in Humanities Scholarship. Literary and Linguistic Computing, 13(3), 111–117. https://doi.org/10.1093/llc/13.3.111
Hovy, D., & Yang, D. (2021). The Importance of Modeling Social Factors of Language: Theory and Practice. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 588–602. https://doi.org/10.18653/v1/2021.naacl-main.49
Ippolito, D., Duckworth, D., Callison-Burch, C., & Eck, D. (2020). Automatic Detection of Generated Text Is Easiest When Humans Are Fooled. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 1808–1822. https://doi.org/10.18653/v1/2020.acl-main.164
Katib, I., Assiri, F. Y., Abdushkour, H. A., Hamed, D., & Ragab, M. (2023). Differentiating Chat Generative Pretrained Transformer From Humans: Detecting ChatGPT-Generated Text and Human Text Using Machine Learning. Mathematics, 11(15), Article 3400. https://doi.org/10.3390/math11153400
Kestemont, M. (2018). Overview of the Author Identification Task at Pan-2018: Cross-Domain Authorship Attribution and Style Change Detection. CEUR Workshop Proceedings, 2125(Query date: 2024-03-21 10:08:46). https://api.elsevier.com/content/abstract/scopus_id/85051054751
Khalifa, M., & Albadawy, M. (2024). Using Artificial Intelligence in Academic Writing and Research: An Essential Productivity Tool. Computer Methods and Programs in Biomedicine Update, 5, 100145. https://doi.org/10.1016/j.cmpbup.2024.100145
Knowles, A. M. (2024). Machine-In-The-Loop Writing: Optimizing the Rhetorical Load. Computers and Composition, 71, Article 102826. https://doi.org/10.1016/j.compcom.2024.102826
Koppel, M., Schler, J., & Argamon, S. (2011). Authorship Attribution in the Wild. Language Resources and Evaluation, 45(1), 83–94. https://doi.org/10.1007/s10579-009-9111-2
Lestariningsih, Dewi Nastiti; Puspitasari, Devi Ambarwati; Karlina, Yenny; Handoko, Mu'awal Panji; Sukma, Bayu Permana; Sanubarianto, Salimulloh Tegar, 2025, “Eksplorasi Atribusi Kepenulisan Komunitas Multietnis di Kalimantan Barat Sebagai Upaya Pemajuan Kebudayaan”, https://hdl.handle.net/20.500.12690/RIN/JUBPBH, RIN Dataverse, V1
Machová, K., Szabóova, M., Paralič, J., & Mičko, J. (2023). Detection of Emotion by Text Analysis Using Machine Learning. Frontiers in Psychology, 14. https://doi.org/10.3389/fpsyg.2023.1190326
MacLeod, N., & Grant, T. (2012). Whose Tweet? Authorship Analysis of Micro-Blogs and Other Short-Form Messages. In S. Tomblin, N. MacLeod, R. Sousa-Silva, & M. Coulthard (Eds.), Proceedings of the International Association of Forensic Linguists' Tenth Biennial Conference (pp. 210–224). Centre for Forensic Linguistics, Aston University.
McMenamin, G. R. (2002). Forensic Linguistics: Advances in Forensic Stylistics. CRC Press LLC.
Moneus, A. M., & Sahari, Y. (2024). Artificial Intelligence and Human Translation: A Contrastive Study Based on Legal Texts. Heliyon, 10(6), Article e28106. https://doi.org/10.1016/j.heliyon.2024.e28106
Pradita, I., Puspitasari, D. A., Karlina, Y., & Sukma, B. P. (2026). Introducing CILLCO: A Corpus Model of Vernacular Indonesian as a Cultural Capital. Jurnal Komunikasi, 20(1), 161–174. https://doi.org/10.20885/komunikasi.vol20.iss1.art10
Puspitasari, Devi Ambarwati; Karlina, Yenny; Hernina, Hernina; Kurniawan, Kurniawan; Mulyo, Budi Mukhammad, 2025, “Rancang Bangun Text Curation Engine Forensik Kebahasaan”, https://hdl.handle.net/20.500.12690/RIN/B9LJYY, RIN Dataverse, V1
Santos, F. A. O., Macedo, H. T., Bispo, T. D., & Zanchettin, C. (2021). Morphological Skip-Gram: Replacing Fast Text Characters N-Gram with Morphological Knowledge. Inteligencia Artificial, 24(67), 1–17. https://doi.org/10.4114/intartif.vol24iss67pp1-17
Singh, A., Sharma, D., Nandy, A., & Singh, V. K. (2024). Towards a Large-Sized Curated and Annotated Corpus for Discriminating Between Human Written and AI-Generated Texts: A Case Study of Text Sourced From Wikipedia and ChatGPT. Natural Language Processing Journal, 6, 100050. https://doi.org/10.1016/j.nlp.2023.100050
Stamatatos, E. (2009). A Survey of Modern Authorship Attribution Methods. Journal of the American Society for Information Science and Technology, 60(3), 538–556. https://doi.org/10.1002/asi.21001
Stamatatos, E. (2018). Masking Topic-Related Information to Enhance Authorship Attribution. Journal of the Association for Information Science and Technology, 69(3), 461–473. https://doi.org/10.1002/asi.23968
Theophilo, A., Giot, R., & Rocha, A. (2023). Authorship Attribution of Social Media Messages. IEEE Transactions on Computational Social Systems, 10(1), 10–23. https://doi.org/10.1109/TCSS.2021.3123895
Uchendu, A. (2020). Authorship Attribution for Neural Text Generation. EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, Query date: 2024-03-21 10:08:46, 8384–8395. https://api.elsevier.com/content/abstract/scopus_id/85098761250
Yang, L., Jiang, F., & Li, H. (2024). Is ChatGPT Involved in Texts? Measure the Polish Ratio to Detect ChatGPT-Generated Text. APSIPA Transactions on Signal and Information Processing, 13(2), 1–19. https://doi.org/10.1561/116.00000250
The journal operates an Open Access policy under a Creative Commons Attribution-NonCommercial 4.0 International License. Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.


_(1).png)
.png)
.png)
1.png)
.png)

_-_Copy_-_Copy.png)


