Linguistic Fingerprints for Authorship Attribution in Human–AI Collaborative Texts

Abstract views: 63 , PDF downloads: 21
Keywords: authorship analysis, authorship attribution, Indonesian text, Lexical Choice, N-gram tracing

Abstract

Authors leave distinctive linguistic traces that reflect their identity through consistent writing styles, particularly in morphosyntactic patterns and lexical choices. However, the increasing use of AI writing tools challenges authorship attribution because machine-generated text can imitate or obscure individual writing characteristics. This study investigates linguistic features that effectively identify authors and differentiate human-written from AI-generated texts. A corpus comprising 2,074,125 tokens and 63,414 word types was compiled from collaborative digital platforms, including instant messaging and social media. Lexical and stylistic features were extracted to develop hybrid authorship-classification models, while N-gram tracing was used to identify salient patterns. The findings demonstrate that lexical choice is the most reliable indicator for distinguishing human and AI-generated texts. Character-level N-gram analysis further demonstrates that authorship can be identified through delicate patterns involving letters, capitalization, punctuation, and other non-alphabetic characters. Diction appeared as the strongest factor in differentiating individual authors. These results enhance the reliability of authorship attribution methods and provide valuable insights for forensic investigations of digitally mediated communication involving human–AI interaction.

 

Downloads

Download data is not yet available.

Author Biographies

Devi Ambarwati Puspitasari, The National Research and Innovation Agency, Jakarta Selatan 12710

DEVI AMBARWATI PUSPITASARI is a junior researcher at the National Research and Innovation Agency. Her research focuses on forensic linguistics, corpus linguistics, and multilingualism.

Dewi Nastiti Lestariningsih, The National Research and Innovation Agency, Jakarta Selatan 12710

DEWI NASTITI LESTARININGSIH is a junior researcher at the National Research and Innovation Agency. Her research focuses on literacy and corpus linguistics.

Bayu Permana Sukma, Linguistics, Faculty of Humanities, Arts and Social Sciences, Lancaster University, Lancaster, LA1 4YW

BAYU PERMANA SUKMA is a junior researcher at the National Research and Innovation Agency. His research focuses on corpus-assisted discourse analysis, political discourse, and media discourse.

Yenny Karlina, The National Research and Innovation Agency, Jakarta Selatan 12710

YENNY KARLINA is a junior researcher at the National Research and Innovation Agency. His research focuses on applied linguistics.

Salimulloh Tegar Sanubarianto, The National Research and Innovation Agency, Jakarta Selatan 12710

SALIMULLOH TEGAR SANUBARIANTO is a junior researcher at the National Research and Innovation Agency. His research focuses on forensic linguistics.

Mu'awal Panji Handoko, The National Research and Innovation Agency, Jakarta Selatan 12710

MU’AWAL PANJI HANDOKO is a junior researcher at the National Research and Innovation Agency. His research focuses on political communication.

Intan Pradita, English Language Education, Faculty of Social and Cultural Sciences, Universitas Islam Indonesia, Yogyakarta 55584

INTAN PRADITA is a lecturer at English Language Education, Faculty of Social and Cultural Sciences, Universitas Islam Indonesia. Her research focuses on corpus linguistics, multilingualism, and systemic functional linguistics.

References

Abreu Rodrigues, S., & Sousa Silva, R. (2022). A Forensic Authorship Analysis of Threats. RevSALUS - Revista Científica Da Rede Académica Das Ciências Da Saúde Da Lusofonia, 4(Sup),p. 98–99. https://doi.org/10.51126/revsalus.v4isup.324

Alamleh, H., Alqahtani, A. A. S., & Elsaid, A. (2023). Distinguishing Human-Written and ChatGPT-Generated Text Using Machine Learning. 2023 Systems and Information Engineering Design Symposium, SIEDS 2023. https://doi.org/10.1109/SIEDS58326.2023.10137767

Aribowo, A. S., Basiron, H., Yusof, N. F. A., & Khomsah, S. (2021). Cross-Domain Sentiment Analysis Model on Indonesian YouTube Comment. International Journal of Advances in Intelligent Informatics, 7(1), 12–25. https://doi.org/10.26555/ijain.v7i1.554

Baker, P. (2010). Title of chapter. In J. Smith (Ed.), Sociolinguistics and Corpus Linguistics (pp. 25–42). Edinburgh University Press.

Banga, R., Bhardwaj, A., Peng, S. L., & Shrivastava, G. (2018). Authorship Attribution for Online Social Media. In Social Network Analytics for Contemporary Business Organizations (pp. 141–165). https://doi.org/10.4018/978-1-5225-5097-6.ch008

Belvisi, N. M. S., Muhammad, N., & Alonso-Fernandez, F. (2020). Forensic Authorship Analysis of Microblogging Texts Using N-Grams and Stylometric Features. In 2020, the 8th International Workshop on Biometrics and Forensics (IWBF) (pp. 1–6). IEEE. https://doi.org/10.1109/IWBF49977.2020.9107956

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623. https://doi.org/10.1145/3442188.3445922

Bevendorff, J., Ghanem, B., Giachanou, A., Kestemont, M., Manjavacas, E., Markov, I., Mayerl, M., Potthast, M., Rangel, F., Rosso, P., Specht, G., Stamatatos, E., Stein, B., Wiegmann, M., & Zangerle, E. (2020). Overview of Pan 2020: Authorship Verification, Celebrity Profiling, Profiling Fake News Spreaders on Twitter, and Style Change Detection. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 12260 Lncs. https://doi.org/10.1007/978-3-030-58219-7_25

Beyer, K. (2014). Urban Language Research in South Africa: Achievements and Challenges. Southern African Linguistics and Applied Language Studies, 32(2), 247–254. https://doi.org/10.2989/16073614.2014.992643

BRIN. Cillco Indonesia Corpora: Corpus of Indonesia Language, Literature, and Community 2024. https://cillco-prototype.brin.go.id

Caliskan, A., Bryson, J. J., & Narayanan, A. (2017). Semantics Derived Automatically From Language Corpora Contain Human-Like Biases. Science, 356(6334), 183–186. https://doi.org/10.1126/science.aal4230

Chen, B., Ding, X., Zhao, Y., Fu, B., Lin, T., Qin, B., & Liu, T. (2024). Text Difficulty Study: Do Machines Behave the Same as Humans Regarding Text Difficulty? Machine Intelligence Research, 21(2), 283–293. https://doi.org/10.1007/s11633-023-1424-x

Dou, Y., Forbes, M., Koncel-Kedziorski, R., Smith, N. A., & Choi, Y. (2022). Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1, 7250–7274. https://doi.org/10.18653/v1/2022.acl-long.501

Dugan, L., Ippolito, D., Kirubarajan, A., Shi, S., & Callison-Burch, C. (2023). Real or Fake Text?: Investigating Human Ability to Detect Boundaries Between Human-Written and Machine-Generated Text. Proceedings of the AAAI Conference on Artificial Intelligence, 37(11), 12763–12771. https://doi.org/10.1609/aaai.v37i11.26501

Fedotova, A., Romanov, A., Kurtukova, A., & Shelupanov, A. (2022). Authorship Attribution of Social Media and Literary Russian-Language Texts Using Machine Learning Methods and Feature Selection. Future Internet, 14(1), 4. https://doi.org/10.3390/fi14010004

García-Díaz, J. A., Colomo-Palacios, R., & Valencia-García, R. (2022). Psychographic Traits Identification Based on Political Ideology: An Author Analysis Study on Spanish Politicians’ Tweets Posted in 2020. Future Generation Computer Systems, 130, 59–74. https://doi.org/10.1016/j.future.2021.12.011

Grant, T. (2007). Quantifying Evidence in Forensic Authorship Analysis. International Journal of Speech, Language and the Law, 14(1), 1–25. https://doi.org/10.1558/ijsll.v14i1.1

Grieve, J. (2023). Register Variation Explains Stylometric Authorship Analysis. Corpus Linguistics and Linguistic Theory, 19(1), 47–77. https://doi.org/10.1515/cllt-2022-0040

Grieve, J., Clarke, I., Chiang, E., Gideon, H., Heini, A., Nini, A., & Waibel, E. (2019). Attributing the Bixby Letter Using N-Gram Tracing. Digital Scholarship in the Humanities, 34(3), 493–512. https://doi.org/10.1093/llc/fqy042

Habibzadeh, F. (2023). GPTZero Performance in Identifying Artificial Intelligence-Generated Medical Texts: A Preliminary Study. Journal of Korean Medical Science, 38(38), 1–12. https://doi.org/10.3346/jkms.2023.38.e319

Holmes, D. I. (1998). The Evolution of Stylometry in Humanities Scholarship. Literary and Linguistic Computing, 13(3), 111–117. https://doi.org/10.1093/llc/13.3.111

Hovy, D., & Yang, D. (2021). The Importance of Modeling Social Factors of Language: Theory and Practice. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 588–602. https://doi.org/10.18653/v1/2021.naacl-main.49

Ippolito, D., Duckworth, D., Callison-Burch, C., & Eck, D. (2020). Automatic Detection of Generated Text Is Easiest When Humans Are Fooled. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 1808–1822. https://doi.org/10.18653/v1/2020.acl-main.164

Katib, I., Assiri, F. Y., Abdushkour, H. A., Hamed, D., & Ragab, M. (2023). Differentiating Chat Generative Pretrained Transformer From Humans: Detecting ChatGPT-Generated Text and Human Text Using Machine Learning. Mathematics, 11(15), Article 3400. https://doi.org/10.3390/math11153400

Kestemont, M. (2018). Overview of the Author Identification Task at Pan-2018: Cross-Domain Authorship Attribution and Style Change Detection. CEUR Workshop Proceedings, 2125(Query date: 2024-03-21 10:08:46). https://api.elsevier.com/content/abstract/scopus_id/85051054751

Khalifa, M., & Albadawy, M. (2024). Using Artificial Intelligence in Academic Writing and Research: An Essential Productivity Tool. Computer Methods and Programs in Biomedicine Update, 5, 100145. https://doi.org/10.1016/j.cmpbup.2024.100145

Knowles, A. M. (2024). Machine-In-The-Loop Writing: Optimizing the Rhetorical Load. Computers and Composition, 71, Article 102826. https://doi.org/10.1016/j.compcom.2024.102826

Koppel, M., Schler, J., & Argamon, S. (2011). Authorship Attribution in the Wild. Language Resources and Evaluation, 45(1), 83–94. https://doi.org/10.1007/s10579-009-9111-2

Lestariningsih, Dewi Nastiti; Puspitasari, Devi Ambarwati; Karlina, Yenny; Handoko, Mu'awal Panji; Sukma, Bayu Permana; Sanubarianto, Salimulloh Tegar, 2025, “Eksplorasi Atribusi Kepenulisan Komunitas Multietnis di Kalimantan Barat Sebagai Upaya Pemajuan Kebudayaan”, https://hdl.handle.net/20.500.12690/RIN/JUBPBH, RIN Dataverse, V1

Machová, K., Szabóova, M., Paralič, J., & Mičko, J. (2023). Detection of Emotion by Text Analysis Using Machine Learning. Frontiers in Psychology, 14. https://doi.org/10.3389/fpsyg.2023.1190326

MacLeod, N., & Grant, T. (2012). Whose Tweet? Authorship Analysis of Micro-Blogs and Other Short-Form Messages. In S. Tomblin, N. MacLeod, R. Sousa-Silva, & M. Coulthard (Eds.), Proceedings of the International Association of Forensic Linguists' Tenth Biennial Conference (pp. 210–224). Centre for Forensic Linguistics, Aston University.

McMenamin, G. R. (2002). Forensic Linguistics: Advances in Forensic Stylistics. CRC Press LLC.

Moneus, A. M., & Sahari, Y. (2024). Artificial Intelligence and Human Translation: A Contrastive Study Based on Legal Texts. Heliyon, 10(6), Article e28106. https://doi.org/10.1016/j.heliyon.2024.e28106

Pradita, I., Puspitasari, D. A., Karlina, Y., & Sukma, B. P. (2026). Introducing CILLCO: A Corpus Model of Vernacular Indonesian as a Cultural Capital. Jurnal Komunikasi, 20(1), 161–174. https://doi.org/10.20885/komunikasi.vol20.iss1.art10

Puspitasari, Devi Ambarwati; Karlina, Yenny; Hernina, Hernina; Kurniawan, Kurniawan; Mulyo, Budi Mukhammad, 2025, “Rancang Bangun Text Curation Engine Forensik Kebahasaan”, https://hdl.handle.net/20.500.12690/RIN/B9LJYY, RIN Dataverse, V1

Santos, F. A. O., Macedo, H. T., Bispo, T. D., & Zanchettin, C. (2021). Morphological Skip-Gram: Replacing Fast Text Characters N-Gram with Morphological Knowledge. Inteligencia Artificial, 24(67), 1–17. https://doi.org/10.4114/intartif.vol24iss67pp1-17

Singh, A., Sharma, D., Nandy, A., & Singh, V. K. (2024). Towards a Large-Sized Curated and Annotated Corpus for Discriminating Between Human Written and AI-Generated Texts: A Case Study of Text Sourced From Wikipedia and ChatGPT. Natural Language Processing Journal, 6, 100050. https://doi.org/10.1016/j.nlp.2023.100050

Stamatatos, E. (2009). A Survey of Modern Authorship Attribution Methods. Journal of the American Society for Information Science and Technology, 60(3), 538–556. https://doi.org/10.1002/asi.21001

Stamatatos, E. (2018). Masking Topic-Related Information to Enhance Authorship Attribution. Journal of the Association for Information Science and Technology, 69(3), 461–473. https://doi.org/10.1002/asi.23968

Theophilo, A., Giot, R., & Rocha, A. (2023). Authorship Attribution of Social Media Messages. IEEE Transactions on Computational Social Systems, 10(1), 10–23. https://doi.org/10.1109/TCSS.2021.3123895

Uchendu, A. (2020). Authorship Attribution for Neural Text Generation. EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, Query date: 2024-03-21 10:08:46, 8384–8395. https://api.elsevier.com/content/abstract/scopus_id/85098761250

Yang, L., Jiang, F., & Li, H. (2024). Is ChatGPT Involved in Texts? Measure the Polish Ratio to Detect ChatGPT-Generated Text. APSIPA Transactions on Signal and Information Processing, 13(2), 1–19. https://doi.org/10.1561/116.00000250

Published
2026-05-30
How to Cite
Puspitasari, D. A., Lestariningsih, D. N., Sukma, B. P., Karlina, Y., Sanubarianto, S. T., Handoko, M. P., & Pradita, I. (2026). Linguistic Fingerprints for Authorship Attribution in Human–AI Collaborative Texts. OKARA: Jurnal Bahasa Dan Sastra, 20(1), 45–69. https://doi.org/10.19105/ojbs.v20i1.22271