Stylometry Analyzis of Human and Machine Text for Academic Integrity
Abstract
1. Introduction
- The paper proposes a style change analyzis-based NLP framework for tackling some of the key challenges faced by academia in the era of generative AI. These tasks include differentiating between single- and multi-authored documents, classifying machine- and human-generated text, identifying the paragraphs where author changes occur in multi-authored documents, and author recognition in multi-authored documents.
- The paper also analyzes the impact of cleverly crafted prompts on the performance of state-of-the-art NLP models-based tools for plagiarism and document authentication
- Two different sets of machine text are generated and embedded in a human-produced dataset, providing a valuable source for further research in the domain. The dataset and source code are publicly available.
- The paper highlights key challenges associated with the topic and hints at some potential solutions.
2. Related Work
3. Methodology
3.1. Dataset Creation
3.2. Data Pre-Processing
3.3. Text Classification
4. Experiments and Results
4.1. Experimental Setup
- Classification of machine and human-generated text: This is a binary classification task where the models predict whether a document was produced by a human author or a machine.
- Classification of single and multi-authored documents: This is also a binary classification task where the models predict whether a document was produced by a single author or is multi-authored. The single-authored documents include text produced by humans and machines. In the case of multi-authored documents, machine-generated text was randomly embedded in the documents produced by multiple human authors.
- Author change detection: This is also a binary classification task that involves the identification of subsequent paragraphs where the author changes. Each pair of subsequent paragraphs in multi-authored documents is labeled either 0 or 1, representing ”no changes” and ”author changes”, respectively.
- Author recognition: In this experiment, we need to identify the author for each document that could be either a single-authored or collaboratively written document. Our goal is to predict whether a particular author contributed to each document or not. Thus, this task is treated as a multi-label classification task with five different authors, including the machine/AI as one of the authors.
4.2. Experimental Results
4.2.1. Machine and Human-Text Classification
4.2.2. Classification of Single and Multi-Authored Documents
4.2.3. Author Change Detection in Multi-Authored Documents
4.2.4. Author Recognition in Multi-Authored Documents
5. Conclusions and Future Work
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| NLP | Natural Language Processing |
| SMOTE | Synthetic Minority Over-sampling Technique |
| CNN | Convolutional Neural Network |
References
- Fariello, S.; Fenza, G.; Forte, F.; Gallo, M.; Marotta, M. Distinguishing Human from Machine: A Review of Advances and Challenges in AI-Generated Text Detection. Int. J. Interact. Multimed. Artif. Intell. 2025, 9, 6–18. [Google Scholar] [CrossRef]
- Mulenga, R.; Shilongo, H. Academic integrity in higher education: Understanding and addressing plagiarism. Acta Pedagog. Asiana 2024, 3, 30–43. [Google Scholar] [CrossRef]
- Jones, B.; Luger, E.; Jones, R. Generative AI & Journalism: A Rapid Risk-Based Review; University of Edinburgh: Edinburgh, Scotland, 2023. [Google Scholar]
- Škiljić, A. When art meets technology or vice versa: Key challenges at the crossroads of AI-generated artworks and copyright law. IIC-Int. Rev. Intellect. Prop. Compet. Law 2021, 52, 1338–1369. [Google Scholar] [CrossRef]
- Silva, E.C.d.M.; Vaz, J.C. How disinformation and fake news impact public policies?: A review of international literature. arXiv 2024, arXiv:2406.00951. [Google Scholar] [CrossRef]
- Balakrishnan, V.; Ng, W.Z.; Soo, M.C.; Han, G.J.; Lee, C.J. Infodemic and fake news – A comprehensive overview of its global magnitude during the COVID-19 pandemic in 2021: A scoping review. Int. J. Disaster Risk Reduct. 2022, 78, 103144. [Google Scholar] [CrossRef] [PubMed]
- Zhou, Y.; He, B.; Sun, L. Humanizing machine-generated content: Evading AI-text detection through adversarial attack. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Turin, Italy, 20–25 May 2024; pp. 8427–8437. [Google Scholar]
- Manzoor, M.F.; Farooq, M.S.; Abid, A. Stylometry-driven framework for Urdu intrinsic plagiarism detection: A comprehensive analyzis using machine learning, deep learning, and large language models. Neural Comput. Appl. 2025, 37, 6479–6513. [Google Scholar] [CrossRef]
- Manzoor, M.F.; Farooq, M.S.; Haseeb, M.; Farooq, U.; Khalid, S.; Abid, A. Exploring the landscape of intrinsic plagiarism detection: Benchmarks, techniques, evolution, and challenges. IEEE Access 2023, 11, 140519–140545. [Google Scholar] [CrossRef]
- Shapoval, R.V.; Nastyuk, V.Y.; Inshyn, M.I.; Posashkov, A.A. Academic honesty: Current status and ways of improvement. Justicia 2021, 26, 37–46. [Google Scholar] [CrossRef]
- Eaton, S.E. Trust as a foundation for ethics and integrity in educational contexts. Crit. Stud. Teach. Learn. 2025, 13, 4–7. [Google Scholar] [CrossRef]
- Fundamental, T. Values of Academic Integrity; The Center for Academic Integrity: Des Plaines, IL, USA, 1999. [Google Scholar]
- Bretag, T.; Green, M. The role of virtue ethics principles in academic integrity breach decision-making. J. Acad. Ethics 2014, 12, 165–177. [Google Scholar] [CrossRef]
- Balalle, H.; Pannilage, S. Reassessing academic integrity in the age of AI: A systematic literature review on AI and academic integrity. Soc. Sci. Humanit. Open 2025, 11, 101299. [Google Scholar] [CrossRef]
- Baron, P. Are AI detection and plagiarism similarity scores worthwhile in the age of ChatGPT and other Generative AI? Scholarsh. Teach. Learn. South (SOTL) South 2024, 8, 151–179. [Google Scholar] [CrossRef]
- Ardito, C.G. Generative AI detection in higher education assessments. New Dir. Teach. Learn. 2025, 2025, 11–28. [Google Scholar] [CrossRef]
- Elkhatat, A.M.; Elsaid, K.; Almeer, S. Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text. Int. J. Educ. Integr. 2023, 19, 17. [Google Scholar] [CrossRef]
- Malik, M.A.; Amjad, A.I. AI vs AI: How effective are Turnitin, ZeroGPT, GPTZero, and Writer AI in detecting text generated by ChatGPT, Perplexity, and Gemini? J. Appl. Learn. Teach. 2025, 8, 91–101. [Google Scholar]
- Michailidis, P.D. A scientometric study of the stylometric research field. Informatics 2022, 9, 60. [Google Scholar] [CrossRef]
- Zamir, M.T.; Ayub, M.A.; Khan, J.; Ikram, M.J.; Ahmad, N.; Ahmad, K. Document provenance and authentication through authorship classification. In Proceedings of the 2023 1st International Conference on Advanced Innovations in Smart Cities (ICAISC); IEEE: Jeddah, Saudi Arabia, 2023; pp. 1–6. [Google Scholar]
- Oberreuter, G.; Velásquez, J.D. Text mining applied to plagiarism detection: The use of words for detecting deviations in the writing style. Expert Syst. Appl. 2013, 40, 3756–3763. [Google Scholar] [CrossRef]
- Zamir, M.T.; Ayub, M.A.; Gul, A.; Ahmad, N.; Ahmad, K. Stylometry analyzis of multi-authored documents for authorship and author style change detection. arXiv 2024, arXiv:2401.06752. [Google Scholar]
- Zangerle, E.; Mayerl, M.; Tschuggnall, M.; Potthast, M.; Stein, B. Pan21 Authorship analyzis: Style Change Detection. In Proceedings of the CEUR Workshop Proceedings, Bucharest, Romania, 21–24 September 2021. [Google Scholar]
- Strøm, E. Multi-label Style Change Detection by Solving a Binary Classification Problem. In Proceedings of the CLEF 2021—Conference and Labs of the Evaluation Forum, Bucharest, Romania, 21–24 September 2021; pp. 2146–2157. [Google Scholar]
- Kestemont, M.; Tschuggnall, M.; Stamatatos, E.; Daelemans, W.; Specht, G.; Stein, B.; Potthast, M. Overview of the author identification task at PAN-2018: Cross-domain authorship attribution and style change detection. In Proceedings of the Working Notes Papers of the CLEF 2018 Evaluation Labs, Avignon, France, 10–14 September 2018; pp. 1–25. [Google Scholar]
- Zangerle, E.; Mayerl, M.; Potthast, M.; Stein, B. Overview of the Style Change Detection Task at PAN 2020. CLEF (Work. Notes) 2020, 93. [Google Scholar]
- Kumarage, T.; Garland, J.; Bhattacharjee, A.; Trapeznikov, K.; Ruston, S.; Liu, H. Stylometric detection of ai-generated text in twitter timelines. arXiv 2023, arXiv:2303.03697. [Google Scholar] [CrossRef]
- Opara, C. StyloAI: Distinguishing AI-generated content with stylometric analyzis. In Proceedings of the International Conference on Artificial Intelligence in Education; Springer: Berlin/Heidelberg, Germany, 2024; pp. 105–114. [Google Scholar]
- Fabien, M.; Villatoro-Tello, E.; Motlicek, P.; Parida, S. BertAA: BERT fine-tuning for Authorship Attribution. In Proceedings of the 17th International Conference on Natural Language Processing (ICON), Patna, India, 18–21 December 2020; pp. 127–137. [Google Scholar]
- Almutairi, A.; Kang, B.; Al Hashimy, N. Bibert-av: Enhancing authorship verification through siamese networks with pre-trained bert and bi-lstm. In Proceedings of the International Conference on Ubiquitous Security; Springer: Berlin/Heidelberg, Germany, 2023; pp. 17–30. [Google Scholar]
- Grootendorst, M. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv 2022, arXiv:2203.05794. [Google Scholar]
- Chauhan, U.; Shah, A. Topic modeling using latent Dirichlet allocation: A survey. ACM Comput. Surv. (CSUR) 2021, 54, 1–35. [Google Scholar] [CrossRef]
- Lee, D.; Seung, H.S. Algorithms for non-negative matrix factorization. Adv. Neural Inf. Process. Syst. 2000, 13, 535–541. [Google Scholar]
- Chawla, N.V.; Bowyer, K.W.; Hall, L.O.; Kegelmeyer, W.P. SMOTE: Synthetic minority over-sampling technique. J. Artif. Intell. Res. 2002, 16, 321–357. [Google Scholar] [CrossRef]
- Batista, G.E.; Prati, R.C.; Monard, M.C. A study of the behavior of several methods for balancing machine learning training data. ACM SIGKDD Explor. Newsl. 2004, 6, 20–29. [Google Scholar] [CrossRef]
- He, H.; Bai, Y.; Garcia, E.A.; Li, S. ADASYN: Adaptive synthetic sampling approach for imbalanced learning. In Proceedings of the 2008 IEEE International Joint Conference on Neural Networks (IEEE World Congress on Computational Intelligence); IEEE: Hong Kong, China, 2008; pp. 1322–1328. [Google Scholar]
- Nayan, N.M.; Islam, A.; Islam, M.U.; Ahmed, E.; Hossain, M.M.; Alam, M.Z. Smote oversampling and near miss undersampling based diabetes diagnosis from imbalanced dataset with xai visualization. In Proceedings of the 2023 IEEE Symposium on Computers and Communications (ISCC); IEEE: Gammarth, Tunisia, 2023; pp. 1–6. [Google Scholar]
- Konuma, R.; Huaping, Z.; Gao, C.; Wang, J. Japanese Author Attribution Using BERT Finetuning with Stylometric. In Intelligent Multilingual Information Processing: First International Conference, IMLIP 2024, Beijing, China, 16–17 November 2024, Proceedings; Springer Nature: Berlin/Heidelberg, Germany, 2025; Volume 2395, p. 293. [Google Scholar]


| File Systems & Data Management | Programming & Software Development |
| Networking & Cybersecurity | Cloud & Virtualization Technologies |
| AI, ML & Data Science | Hardware & System Performance |
| Software Installation & Troubleshooting | Enterprise IT & DevOps |
| Tech Industry Trends | User Experience & HCI |
| Instructions Used in the Normal Prompt | Instructions Used in the Strict Prompt |
|---|---|
|
|
| Hyper-Parameter | Value | Hyper-Parameter | Value |
|---|---|---|---|
| Batch Size | 32 | Gradient Accumulation | 2 |
| Learning Rate | 0.00001 | Weight Decay | 0.01 |
| Epochs | 5 | Logging | 50 |
| Task | Number of Samples | |||
|---|---|---|---|---|
| Label | Training | Validation | Test | |
| Machine vs. Human Text Classification | 0 | 77,061 | 16,438 | 16,344 |
| 1 | 12,952 | 2751 | 2748 | |
| Single- vs. Multi-authored Classification | 0 | 2800 | 600 | 600 |
| 1 | 8379 | 1797 | 1794 | |
| Author Change Detection | 0 | 27,725 | 5934 | 5814 |
| 1 | 51,138 | 10,887 | 10,886 | |
| Author Recognition | 1 | 41,672 | 8985 | 8825 |
| 2 | 20,605 | 4362 | 4354 | |
| 3 | 10,522 | 2235 | 2238 | |
| 4 | 4291 | 885 | 929 | |
| 5 | 12,952 | 2751 | 2748 | |
| Model | Normal Prompt | Strict Prompt | ||||||
|---|---|---|---|---|---|---|---|---|
| Accuracy | Precision | Recall | F1-Score | Accuracy | Precision | Recall | F1-Score | |
| distilbert-base-uncased | 0.9998 | 0.9998 | 0.9998 | 0.9998 | 0.9477 | 0.9463 | 0.9477 | 0.9463 |
| Albert-base-v2 | 0.9995 | 0.9995 | 0.9995 | 0.9995 | 0.9456 | 0.9440 | 0.9456 | 0.9441 |
| Roberta-base | 0.9997 | 0.9997 | 0.9997 | 0.9997 | 0.9487 | 0.9473 | 0.9487 | 0.9474 |
| Bert-base-uncased | 0.9998 | 0.9998 | 0.9998 | 0.9998 | 0.9483 | 0.9469 | 0.9483 | 0.9467 |
| Class/Author | Dataset 1 (Strict Prompt) | Dataset 2 (Normal Prompt) | ||||
|---|---|---|---|---|---|---|
| Precision | Recall | F1-Score | Precision | Recall | F1-Score | |
| Class 0 (Human Text) | 0.9567 | 0.9795 | 0.9780 | 0.9995 | 1.00 | 0.9998 |
| Class 1 (Machine Text) | 0.8779 | 0.7691 | 0.8199 | 1.00 | 0.9971 | 0.9985 |
| Model | Normal Prompt | Strict Prompt | ||||||
|---|---|---|---|---|---|---|---|---|
| Accuracy | Precision | Recall | F1-Score | Accuracy | Precision | Recall | F1-Score | |
| Distilbert-base-uncased | 0.9866 | 0.9869 | 0.9866 | 0.9866 | 0.9845 | 0.9844 | 0.9845 | 0.9844 |
| Albert-base-v2 | 0.9841 | 0.9850 | 0.9841 | 0.9842 | 0.9853 | 0.9854 | 0.9853 | 0.9853 |
| Roberta-base | 0.9924 | 0.9925 | 0.9924 | 0.9925 | 0.9824 | 0.9825 | 0.9824 | 0.9822 |
| Bert-base-uncased | 0.9857 | 0.9862 | 0.9857 | 0.9858 | 0.9774 | 0.9773 | 0.9774 | 0.9772 |
| Model | Normal Prompt | Strict Prompt | ||||||
|---|---|---|---|---|---|---|---|---|
| Accuracy | Precision | Recall | F1-Score | Accuracy | Precision | Recall | F1-Score | |
| Albert-base-v2 | 0.674 | 0.6982 | 0.674 | 0.6806 | 0.6818 | 0.6964 | 0.6818 | 0.6867 |
| Bert-base-uncased | 0.5878 | 0.6795 | 0.5878 | 0.5928 | 0.64 | 0.6908 | 0.64 | 0.6484 |
| Distilbert-base-uncased | 0.6699 | 0.7023 | 0.6699 | 0.6774 | 0.6973 | 0.7147 | 0.6973 | 0.7026 |
| Roberta-base | 0.6667 | 0.6913 | 0.6667 | 0.6735 | 0.6948 | 0.7149 | 0.6948 | 0.7006 |
| Model | Normal Prompt | Strict Prompt | ||||
|---|---|---|---|---|---|---|
| Precision | Recall | F1-Score | Precision | Recall | F1-Score | |
| Albert-base-v2 | 0.3756 | 0.9106 | 0.5318 | 0.3869 | 0.7433 | 0.5089 |
| bert-base-uncased | 0.3564 | 0.9314 | 0.5155 | 0.3835 | 0.7306 | 0.503 |
| distilbert-base-uncased | 0.355 | 0.9357 | 0.5147 | 0.383 | 0.7094 | 0.4974 |
| Roberta | 0.3561 | 0.9369 | 0.516 | 0.3842 | 0.7344 | 0.5045 |
| Class/Author | Normal Prompt | Strict Prompt | ||||
|---|---|---|---|---|---|---|
| Precision | Recall | F1-Score | Precision | Recall | F1-Score | |
| Author 1 | 0.5396 | 0.999 | 0.70 | 0.6483 | 0.7676 | 0.7029 |
| Author 2 | 0.2674 | 0.9991 | 0.4212 | 0.3005 | 0.7014 | 0.4208 |
| Author 3 | 0.1542 | 0.7297 | 0.2546 | 0.2027 | 0.5624 | 0.2980 |
| Author 4 | 0.0964 | 0.0086 | 0.0158 | 0.0890 | 0.5404 | 0.1528 |
| Author 5 | 0.9990 | 0.9960 | 0.9978 | 0.8516 | 0.7983 | 0.8241 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Albaqami, H.; Ayub, M.A.; Ahmad, N.; Ahmad, Y.; Alqahtani, M.M.; Algamdi, A.M.; Owaidah, A.A.; Ahmad, K. Stylometry Analyzis of Human and Machine Text for Academic Integrity. Computers 2026, 15, 217. https://doi.org/10.3390/computers15040217
Albaqami H, Ayub MA, Ahmad N, Ahmad Y, Alqahtani MM, Algamdi AM, Owaidah AA, Ahmad K. Stylometry Analyzis of Human and Machine Text for Academic Integrity. Computers. 2026; 15(4):217. https://doi.org/10.3390/computers15040217
Chicago/Turabian StyleAlbaqami, Hezam, Muhammad Asif Ayub, Nasir Ahmad, Yaseen Ahmad, Mohammad M. Alqahtani, Abdullah M. Algamdi, Almoaid A. Owaidah, and Kashif Ahmad. 2026. "Stylometry Analyzis of Human and Machine Text for Academic Integrity" Computers 15, no. 4: 217. https://doi.org/10.3390/computers15040217
APA StyleAlbaqami, H., Ayub, M. A., Ahmad, N., Ahmad, Y., Alqahtani, M. M., Algamdi, A. M., Owaidah, A. A., & Ahmad, K. (2026). Stylometry Analyzis of Human and Machine Text for Academic Integrity. Computers, 15(4), 217. https://doi.org/10.3390/computers15040217

