An Empirical Study of Fine-Tuning Pre-Trained Code Models and Adapters for the Classification of Source Code Plagiarism Instances
Abstract
1. Introduction
- 1.
- The work presents a comprehensive empirical study of applying PCMs along with adapters for the task of low-resource SCPD.
- 2.
- This is the first work that reports experiments with PEFT for the task of SCPC with PCMs.
- 3.
- The work compares the supervised fine-tuned models and adapters with the current SOTA unsupervised open-source Source Code Plagiarism Detection Tools (SCPDTs), JPlag and Dolos.
- RQ1
- How does fine-tuning the PCMs perform on the task of SCPC?PCMs are the SOTA for several code tasks. Therefore, the impact of using the PCMs can be investigated in relation to the specific task of classification of source code plagiarism instances.
- RQ2
- What is the impact of merging the adapters on PCMs for the SCPC task?PEFT methods are used to support few-shot learning settings. This research conducts experiments to determine whether they yield better performance, given the task’s low-resource nature.
- RQ3
- How do PEFT and FFT compare regarding training time, inference time, number of trained parameters and GPU usage?There is a trade-off between FFT and using adapters in terms of performance and efficiency. To answer this question, we empirically examine the trade-off using several metrics specified in the research question.
2. Empirical Study
2.1. Methodology Overview
- Training time: the duration required to train the model on the entire dataset for a single dataset version.
- Prediction (inference) time: the time taken for the model to infer the class of the complete testing set within a single dataset version.
- GPU usage percentage: the active GPU usage consumed by the model during the training.
- Number of trainable parameters: the count of parameters adjusted during fine-tuning.
- Model size: the size of the fine-tuned models or adapters.
2.2. Datasets
2.2.1. Plagiarism Levels
- (L1)
- Altering comments and adjusting indentation.
- (L2)
- Renaming or modifying identifiers.
- (L3)
- Modifying declarations, such as adding extra constants or rearranging functions and variables.
- (L4)
- Editing functions, including changing their signature, merging, or creating new ones.
- (L5)
- Replacing program statements with semantic equivalents, such as switching between for and while loops or if and switch.
- (L6)
- Changing decision-making logic or modifying expressions.
2.2.2. Datasets Description
2.2.3. Dataset Sample Examples
2.3. PCMs
2.4. Adapters
Sequential Bottleneck Adapters vs. LoRA
2.5. Experimental Setup
2.5.1. Input Formatting and Tokenisation
2.5.2. Hyperparameter Search
2.6. Evaluation
3. Results
3.1. Random Single Split Results
3.1.1. Full Fine-Tuning of PCMs
3.1.2. PEFT
3.1.3. Error Analysis
3.1.4. FFT vs. PEFT
3.1.5. Comparison with SCPDTs
Single-Split Plagiarism-Level Detection
3.2. Cross-Validation Results
3.2.1. Statistical Testing
3.2.2. Impact of Preprocessing and Longer Context Window
3.2.3. Comparison Against SCPDTs
Cross-Validation Plagiarism-Level Detection
4. Discussion
4.1. Answers to the Research Questions
- RQ1
- How does fine-tuning the PCMs perform on the task of SCPC?The random single split is useful for model selection and diagnostic analysis, but the main conclusion is based on 5-fold cross-validation. Under cross-validation, the selected best fine-tuned configurations achieved a higher than JPlag and Dolos across all datasets when labelled training data were available. The main limitation of the 512-token PCMs is the restricted context window; ModernBERT, with a longer context window and preprocessing, improved the ConPlag results.
- RQ2
- What is the impact of merging the adapters on PCMs for the SCPC task?Adapters with the PCMs had similar performance with several models. The exceptions were CodeBERT in IR-Plag and ConPlag datasets and UniXcoder in ConPlag datasets, which led to lower performance with PEFT. The statistical analysis in Section 3.2.1 indicated no significant difference between FFT and the best PEFT configuration when each dataset–model pair was compared with its strongest adapter result.
- RQ3
- How do Parameter-Efficient Fine-Tuning and Full Fine-Tuning compare in terms of training time, inference time, number of trained parameters, and GPU usage?FFT required the longest training time within the same number of epochs, while the Pfeiffer adapter required the least training time. However, adapters may require more training epochs. Adapters introduced additional inference latency relative to FFT, except for LoRA, which had the lowest inference overhead. Maximum GPU usage occurred with FFT. Full Fine-Tuning adjusts all the weights of the pre-trained model, whereas the adapters adjust around only 0.7% to 2.5% of the trainable parameters. The adapter size is added to the model size for a single-task setting. However, when multiple pre-trained models are fine-tuned on several datasets, adapters can be advantageous in terms of storage.
4.2. Practical Recommendations
4.3. Empirical Study Threats to Validity
4.3.1. Internal
4.3.2. External
4.4. Key Takeaways
- 1.
- The PCMs demonstrate robustness when provided with high-quality training data. PLBART achieved the highest performance among the selected 512-token models, possibly due to its denoising pre-training objective on unlabelled data. By learning to recover original code from corrupted input, the model is better equipped to detect obfuscated or altered plagiarised code. This also makes it well-suited for style-based analysis tasks.
- 2.
- When working with multiple datasets, adapters offer an efficient and modular solution. They enable easy training, merging, and toggling (ON/OFF), thereby increasing flexibility across tasks and datasets.
- 3.
- Smaller models in the study (CodeBERTa and CodeT5 small) benefited more from PEFT. This can be attributed to the model having even fewer trainable parameters, making it more manageable and effective for learning.
- 4.
- Due to their limited token length, the selected PCMs may fail to detect plagiarism when irrelevant or obfuscated code is inserted at the beginning, or when plagiarised content is placed toward the end of longer code files. This was verified using a model with a longer context window, but it still exhibits issues with a high combined token count across pairs. We verified this using ModernBERT with different context lengths.
- 5.
- 6.
- Instructors may occasionally reuse programming assignments across semesters or terms [63]; while reusing assignments is not recommended, anonymised historical submissions can improve the predictive performance of plagiarism detection models when handled responsibly. Adapters may be suitable for this context because they support continual fine-tuning while retaining knowledge of previous adapters.
4.5. Limitations and Directions of Future Work
- 1.
- The maximum length of selected PCMs is 512 tokens. Therefore, plagiarism in lengthy code is harder to detect. As discussed in Section 3.2.2, possible solutions include longer-context models, sparse-attention architectures, chunking with aggregation, hierarchical aggregation, preprocessing, code normalisation, and code summarisation. In this study, we focused on preprocessing and ModernBERT as a longer-context encoder. However, this comes at the cost of greater complexity and increased resource requirements. For practitioners, preprocessing and long-context encoders are the simplest first options when assignments regularly exceed 512 tokens, whereas chunking or hierarchical aggregation may be more suitable when the complete file must be inspected. Future work should compare multiple long-context models and chunk-based strategies rather than relying on a single long-context model.
- 2.
- This work did not cover Code Large Language Models (CodeLLMs) in the evaluation of the models. This is a limitation because CodeLLMs can be beneficial in plagiarism detection [64]. They can be used for dataset creation or annotation, for zero-shot or few-shot pair classification, or fine-tuned with PEFT or instructed to act as plagiarism detectors. They were not included in the empirical evaluation because they require a different prompting-based inference protocol, substantially larger computational resources for fine-tuning, and additional experimental choices such as prompt design, context construction, calibration, and output parsing. Future work should include a dedicated CodeLLM study, starting with zero-shot and few-shot prompting on a representative subset, followed by PEFT-based adaptation and Retrieval-Augmented Generation (RAG) when sufficient compute is available.
- 3.
- Another evaluation approach for PCMs and CodeLLMs is to use a few instances per plagiarism level and formulate the task as few-shot pair classification, as this is more applicable in real-world scenarios.
- 4.
- The ConPlag datasets used in this study are contest-based rather than complete real course cohorts. Contest submissions often involve single-file solutions to certain problems solved individually and may not reflect assignment code transformations. Academic plagiarism may contain partial submissions, multi-file projects, and collaboration patterns. However, they can be used as simulations of plagiarism as recommended by the dataset authors.
- 5.
- Future work should evaluate the same SCPD pipeline on non-Java and multilingual datasets, including Python, C/C++, and JavaScript assignments. Such experiments would clarify whether the observed behaviour of FFT and PEFT transfers to languages with different syntax, typing discipline, idioms, libraries, and common obfuscation strategies. Multilingual code models and language-specific or language-adaptive adapters are promising directions for this validation.
- 6.
- Further empirical studies could explore a broader range of models, additional adapter types, and more extensive hyperparameter optimisation. For example, several variants of LoRA have been proposed, some of which may yield improved performance [65]. Future sensitivity analysis should include adapter reduction factors, LoRA rank and scaling, dropout, batch size, and validation criteria for each model and dataset.
- 7.
- Due to the lack of labelled training data, another way to approach the task is through semi-supervised learning, where both labelled and unlabelled data can be used to enhance the learning in the models.
5. Related Work
5.1. PCMs in SCPD
5.2. Adapters in Software Engineering
| Ref. | Task | Models | Adapters |
|---|---|---|---|
| [70] | Summarisation, defect prediction, translation, clone detection | CodeBERT, GraphCodeBERT, CodeT5, PLBART | Prefix, adapter, IA3, parallel, LoRA, MAM |
| [74] | Clone detection, defect detection, code search, code translation | CodeBERT, RoBERTa, CodeT5, T5, UniXcoder, BART | Houlsby, prefix, LoRA, parallel |
| [75] | Code summarisation, clone detection | CodeBERT, GraphCodeBERT | Adapters |
| [76] | Code change | CodeBERT, GraphCodeBERT, PLBART, UniXcoder, CodeT5 | LoRA, adapters |
| [77] | Defect prediction | CodeBERT, JavaBERT, CodeT5+, CodeReviewer | LoRA |
| [81] | Data race detection | GraphCodeBERT, UniXcoder | Adapters |
| [82] | Code generation | CodeT5+, CodeGen, CodeGen2, CodeLlama | LoRA, prompt, prefix, IA3 |
| [86] | Code review | LLaMAReviewer (LLaMA-based) | Prefix, LoRA |
| [89] | Program repair | CodeLlama, DeepSeek-CoderLlama | LoRA, IA3, prompt, prefix |
| [91] | Program repair | StarCoder2, Granite | LoRA |
| [94] | Code smell detection | CodeBERT, GraphCodeBERT, CodeT5, UniXcoder, DeepSeek-Coder, StarCoder, CodeLlama | Prompt, prefix, LoRA, IA3 |
| [95] | Automatic code repair | InCoder, CodeGen, StarCoder, CodeLlama | LoRA, AdaLoRA, IA3 |
| [98] | Code summarisation | CodeLlama, DeepSeek-Coder, Phi-3-mini | QLoRA |
| This Work | SCPC | CodeBERT, GraphCodeBERT, UniXcoder, CodeT5, CodeBERTa, PLBART | Adapters, LoRA |
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Appendix A. Detailed Error Analysis Examples
Appendix A.1. Dataset Code Pair Examples




Appendix A.2. Local Failure Explanations




Appendix B. JPlag and Dolos Parameters
| Tool | Parameter | Values |
|---|---|---|
| JPlag | Min. matched tokens (t) | 5–20 |
| Dolos | k-gram length (k) | 10, 12, 15, 17, 20, 23, 25 |
| Window size (w) | 14, 17, 20, 25, 30 |
| Dataset | Tool | Hyperparameters | Threshold |
|---|---|---|---|
| Progpedia-19 | JPlag | t = 5; similarity_type = averageSimilarity | 0.83 |
| Dolos | k = 10; w = 14 | 0.89 | |
| ConPlag1 | JPlag | t = 5; similarity_type = maxSimilarity | 0.78 |
| Dolos | k = 25; w = 17 | 0.29 | |
| ConPlag2 | JPlag | t = 8; similarity_type = maxSimilarity | 0.50 |
| Dolos | k = 20; w = 14 | 0.40 | |
| IR-Plag | JPlag | t = 6; similarity_type = averageSimilarity | 0.27 |
| Dolos | k = 17; w = 25 | 0.37 |
Appendix C. Statistical Testing Details
| Dataset | Configuration | Mean | 95% CI |
|---|---|---|---|
| IR-Plag | JPlag | ||
| IR-Plag | Dolos | ||
| IR-Plag | UniXcoder (FFT) | ||
| IR-Plag | UniXcoder (Pfeiffer) | ||
| ConPlag1 | JPlag | ||
| ConPlag1 | Dolos | ||
| ConPlag1 | PLBART (FFT, preprocessing) | ||
| ConPlag1 | PLBART (Pfeiffer, preprocessing) | ||
| ConPlag1 | ModernBERT (Pfeiffer, preprocessing, ) | ||
| ConPlag2 | JPlag | ||
| ConPlag2 | Dolos | ||
| ConPlag2 | PLBART (FFT, preprocessing) | ||
| ConPlag2 | PLBART (Pfeiffer, preprocessing) | ||
| ConPlag2 | ModernBERT (Houlsby, preprocessing, ) | ||
| Progpedia-19 | JPlag | ||
| Progpedia-19 | Dolos | ||
| Progpedia-19 | PLBART (FFT) | ||
| Progpedia-19 | PLBART (Pfeiffer) |
Appendix D. FFT vs. PEFT Efficiency Comparison



References
- Joy, M.; Luck, M. Plagiarism in programming assignments. IEEE Trans. Educ. 1999, 42, 129–133. [Google Scholar] [CrossRef] [Scilit]
- Niu, C.; Li, C.; Ng, V.; Chen, D.; Ge, J.; Luo, B. An Empirical Comparison of Pre-Trained Models of Source Code. arXiv 2023, arXiv:2302.04026. [Google Scholar]
- Pfeiffer, J.; Ruder, S.; Vulić, I.; Ponti, E.M. Modular deep learning. arXiv 2023, arXiv:2302.11529. [Google Scholar]
- Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; Gelly, S. Parameter-efficient transfer learning for NLP. In Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA, 9–15 June 2019; pp. 2790–2799. [Google Scholar]
- Pfeiffer, J.; Rücklé, A.; Poth, C.; Kamath, A.; Vulić, I.; Ruder, S.; Cho, K.; Gurevych, I. Adapterhub: A framework for adapting Transformers. arXiv 2020, arXiv:2007.07779. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. Lora: Low-rank adaptation of large language models. arXiv 2021, arXiv:2106.09685. [Google Scholar]
- Slobodkin, E.; Sadovnikov, A. Towards a Dataset of Programming Contest Plagiarism in Java. arXiv 2023, arXiv:2303.10763. [Google Scholar]
- Karnalim, O.; Budi, S.; Toba, H.; Joy, M. Source Code Plagiarism Detection in Academia with Information Retrieval: Dataset and the Observation. Inform. Educ. 2019, 18, 321–344. [Google Scholar] [CrossRef] [Scilit]
- Paiva, J.C.; Leal, J.P.; Figueira, Á. PROGpedia: Collection of source-code submitted to introductory programming assignments. Data Brief 2023, 46, 108887. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Heneka, N.R. Software Plagiarism Detection on Intermediate Representation. Bachelor’s Thesis, Karlsruhe Institute of Technology (KIT), Karlsruhe, Germany, 2023. [Google Scholar]
- Nayak, A.; Timmapathini, H.P.; Murali, V.; Gohad, A.A. Few-shot learning approaches for classifying low resource domain specific software requirements. arXiv 2023, arXiv:2302.06951. [Google Scholar]
- Zakeri-Nasrabadi, M.; Parsa, S.; Ramezani, M.; Roy, C.; Ekhtiarzadeh, M. A systematic literature review on source code similarity measurement and clone detection: Techniques, applications, and challenges. J. Syst. Softw. 2023, 204, 111796. [Google Scholar] [CrossRef] [Scilit]
- Faidhi, J.A.; Robinson, S.K. An empirical approach for detecting program similarity and plagiarism within a university programming environment. Comput. Educ. 1987, 11, 11–19. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; pp. 5998–6008. [Google Scholar]
- Qiu, X.; Sun, T.; Xu, Y.; Shao, Y.; Dai, N.; Huang, X. Pre-trained models for natural language processing: A survey. Sci. China Technol. Sci. 2020, 63, 1872–1897. [Google Scholar] [CrossRef] [Scilit]
- Mars, M. From Word Embeddings to Pre-Trained Language Models: A State-of-the-Art Walkthrough. Appl. Sci. 2022, 12, 8805. [Google Scholar] [CrossRef] [Scilit]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, Minnesota, June 2019; Association for Computational Linguistics: Kerrville, TX, USA, 2019; pp. 4171–4186. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
- Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; Zettlemoyer, L. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, July 2020; Association for Computational Linguistics: Kerrville, TX, USA, 2020; pp. 7871–7880. [Google Scholar] [CrossRef] [Scilit]
- Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P.J. Exploring the limits of transfer learning with a unified text-to-text Transformer. J. Mach. Learn. Res. 2020, 21, 5485–5551. [Google Scholar]
- Husain, H.; Wu, H.H.; Gazit, T.; Allamanis, M.; Brockschmidt, M. Codesearchnet challenge: Evaluating the state of semantic code search. arXiv 2019, arXiv:1909.09436. [Google Scholar]
- Xu, F.F.; Alon, U.; Neubig, G.; Hellendoorn, V.J. A systematic evaluation of large language models of code. In Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming, San Diego, CA, USA, 13 June 2022; pp. 1–10. [Google Scholar]
- Niu, C.; Li, C.; Luo, B.; Ng, V. Deep learning meets software engineering: A survey on pre-trained models of source code. arXiv 2022, arXiv:2205.11739. [Google Scholar]
- Zeng, Z.; Tan, H.; Zhang, H.; Li, J.; Zhang, Y.; Zhang, L. An extensive study on pre-trained models for program understanding and generation. In Proceedings of the 31st ACM SIGSOFT International Symposium on Software Testing and Analysis, Virtual, Republic of Korea, 18–22 July 2022; pp. 39–51. [Google Scholar]
- Wong, M.F.; Guo, S.; Hang, C.N.; Ho, S.W.; Tan, C.W. Natural Language Generation and Understanding of Big Code for AI-Assisted Programming: A Review. Entropy 2023, 25, 888. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Niu, C.; Li, C.; Ng, V.; Luo, B. Comparing the Pretrained Models of Source Code by Re-pretraining Under a Unified Setup. IEEE Trans. Neural Netw. Learn. Syst. 2024, 35, 17768–17778. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Raihan, N.; Newman, C.; Zampieri, M. Code LLMs: A Taxonomy-based Survey. arXiv 2024, arXiv:2412.08291. [Google Scholar]
- Feng, Z.; Guo, D.; Tang, D.; Duan, N.; Feng, X.; Gong, M.; Shou, L.; Qin, B.; Liu, T.; Jiang, D.; et al. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020, Online, November 2020; Association for Computational Linguistics: Kerrville, TX, USA, 2020; pp. 1536–1547. [Google Scholar] [CrossRef] [Scilit]
- Guo, D.; Ren, S.; Lu, S.; Feng, Z.; Tang, D.; Liu, S.; Zhou, L.; Duan, N.; Svyatkovskiy, A.; Fu, S.; et al. Graphcodebert: Pre-training code representations with data flow. arXiv 2020, arXiv:2009.08366. [Google Scholar]
- Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. Huggingface’s Transformers: State-of-the-art natural language processing. arXiv 2019, arXiv:1910.03771. [Google Scholar] [CrossRef] [Scilit]
- Guo, D.; Lu, S.; Duan, N.; Wang, Y.; Zhou, M.; Yin, J. Unixcoder: Unified cross-modal pre-training for code representation. arXiv 2022, arXiv:2203.03850. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Wang, W.; Joty, S.; Hoi, S.C. Codet5: Identifier-aware unified pre-trained encoder/decoder models for code understanding and generation. arXiv 2021, arXiv:2109.00859. [Google Scholar]
- Ahmad, W.U.; Chakraborty, S.; Ray, B.; Chang, K.W. Unified pre-training for program understanding and generation. arXiv 2021, arXiv:2103.06333. [Google Scholar]
- Fu, Z.; Yang, H.; So, A.M.C.; Lam, W.; Bing, L.; Collier, N. On the Effectiveness of Parameter-Efficient Fine-Tuning. Proc. AAAI Conf. Artif. Intell. 2023, 37, 12799–12807. [Google Scholar] [CrossRef] [Scilit]
- Ding, N.; Qin, Y.; Yang, G.; Wei, F.; Yang, Z.; Su, Y.; Hu, S.; Chen, Y.; Chan, C.M.; Chen, W.; et al. Parameter-efficient fine-tuning of large-scale pre-trained language models. Nat. Mach. Intell. 2023, 5, 220–235. [Google Scholar] [CrossRef] [Scilit]
- Xu, L.; Xie, H.; Qin, S.Z.J.; Tao, X.; Wang, F.L. Parameter-efficient fine-tuning methods for pretrained language models: A critical review and assessment. arXiv 2023, arXiv:2312.12148. [Google Scholar]
- Sabry, M.; Belz, A. Peft-ref: A modular reference architecture and typology for parameter-efficient finetuning techniques. arXiv 2023, arXiv:2304.12410. [Google Scholar]
- Lialin, V.; Deshpande, V.; Rumshisky, A. Scaling down to scale up: A guide to parameter-efficient fine-tuning. arXiv 2023, arXiv:2303.15647. [Google Scholar]
- Han, Z.; Gao, C.; Liu, J.; Zhang, J.; Zhang, S.Q. Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey. arXiv 2024, arXiv:2403.14608. [Google Scholar]
- Pfeiffer, J.; Vulić, I.; Gurevych, I.; Ruder, S. MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, November 2020; Webber, B., Cohn, T., He, Y., Liu, Y., Eds.; Association for Computational Linguistics: Kerrville, TX, USA, 2020; pp. 7654–7673. [Google Scholar] [CrossRef] [Scilit]
- Pfeiffer, J.; Kamath, A.; Rücklé, A.; Cho, K.; Gurevych, I. AdapterFusion: Non-Destructive Task Composition for Transfer Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, Online, April 2021; Merlo, P., Tiedemann, J., Tsarfaty, R., Eds.; Association for Computational Linguistics: Kerrville, TX, USA, 2021; pp. 487–503. [Google Scholar] [CrossRef] [Scilit]
- Poth, C.; Sterz, H.; Paul, I.; Purkayastha, S.; Engländer, L.; Imhof, T.; Vulić, I.; Ruder, S.; Gurevych, I.; Pfeiffer, J. Adapters: A unified library for parameter-efficient and modular transfer learning. arXiv 2023, arXiv:2311.11077. [Google Scholar]
- Zouhar, V.; Meister, C.; Gastaldi, J.; Du, L.; Vieira, T.; Sachan, M.; Cotterell, R. A Formal Perspective on Byte-Pair Encoding. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 2023; Rogers, A., Boyd-Graber, J., Okazaki, N., Eds.; Association for Computational Linguistics: Kerrville, TX, USA, 2023; pp. 598–614. [Google Scholar] [CrossRef] [Scilit]
- Kudo, T.; Richardson, J. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Brussels, Belgium, November 2018; Blanco, E., Lu, W., Eds.; Association for Computational Linguistics: Kerrville, TX, USA, 2018; pp. 66–71. [Google Scholar] [CrossRef] [Scilit]
- Gkouti, N.; Malakasiotis, P.; Toumpis, S.; Androutsopoulos, I. Should I try multiple optimizers when fine-tuning a pre-trained Transformer for NLP tasks? Should I tune their hyperparameters? In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Kerrville, TX, USA, 2024; pp. 2555–2574. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. In Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Joshi, A.V. Machine Learning and Artificial Intelligence; Springer: Cham, Switzerland, 2020. [Google Scholar] [CrossRef] [Scilit]
- Prechelt, L.; Malpohl, G.; Philippsen, M. Finding plagiarisms among a set of programs with JPlag. J. Univers. Comput. Sci. 2002, 8, 1016–1038. [Google Scholar] [CrossRef]
- Maertens, R.; Van Petegem, C.; Strijbol, N.; Baeyens, T.; Jacobs, A.C.; Dawyndt, P.; Mesuere, B. Dolos: Language-agnostic plagiarism detection in source code. J. Comput. Assist. Learn. 2022, 38, 1046–1061. [Google Scholar] [CrossRef] [Scilit]
- Rainio, O.; Teuho, J.; Klén, R. Evaluation metrics and statistical tests for machine learning. Sci. Rep. 2024, 14, 6086. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wilcoxon, F. Individual comparisons by ranking methods. Biom. Bull. 1945, 1, 80–83. [Google Scholar] [CrossRef] [Scilit]
- Friedman, M. The use of ranks to avoid the assumption of normality implicit in the analysis of variance. J. Am. Stat. Assoc. 1937, 32, 675–701. [Google Scholar] [CrossRef]
- Dong, Z.; Tang, T.; Li, J.; Zhao, W.X. A survey on long text modeling with Transformers. arXiv 2023, arXiv:2302.14502. [Google Scholar]
- Huang, Y.; Xu, J.; Lai, J.; Jiang, Z.; Chen, T.; Li, Z.; Yao, Y.; Ma, X.; Yang, L.; Chen, H.; et al. Advancing Transformer architecture in long-context large language models: A comprehensive survey. arXiv 2023, arXiv:2311.12351. [Google Scholar]
- Alva Principe, R.; Chiarini, N.; Viviani, M. Long Document classification in the Transformer era: A survey on challenges, advances, and open issues. Wiley Interdiscip. Rev. Data Min. Knowl. Discov. 2025, 15, e70019. [Google Scholar] [CrossRef] [Scilit]
- Beltagy, I.; Peters, M.E.; Cohan, A. Longformer: The long-document Transformer. arXiv 2020, arXiv:2004.05150. [Google Scholar]
- Zaheer, M.; Guruganesh, G.; Dubey, K.A.; Ainslie, J.; Alberti, C.; Ontanon, S.; Pham, P.; Ravula, A.; Wang, Q.; Yang, L.; et al. Big Bird: Transformers for longer sequences. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 17283–17297. [Google Scholar]
- Kitaev, N.; Kaiser, L.; Levskaya, A. Reformer: The efficient Transformer. In Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia, 26–30 April 2020. [Google Scholar]
- Yang, Z.; Yang, D.; Dyer, C.; He, X.; Smola, A.; Hovy, E. Hierarchical attention networks for document classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Kerrville, TX, USA, 2016; pp. 1480–1489. [Google Scholar]
- Brödel, M. Preventing Automatic Code Plagiarism Generation Through Token String Normalization. Bachelor’s Thesis, Karlsruher Institut für Technologie (KIT), Karlsruhe, Germany, 2023. [Google Scholar]
- Warner, B.; Chaffin, A.; Clavié, B.; Weller, O.; Hallström, O.; Taghadouini, S.; Gallagher, A.; Biswas, R.; Ladhak, F.; Aarsen, T.; et al. Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Kerrville, TX, USA, 2025; pp. 2526–2547. [Google Scholar]
- Biderman, S.; Raff, E. Fooling MOSS detection with pretrained language models. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, Atlanta, GA, USA, 17–21 October 2022; pp. 2933–2943. [Google Scholar]
- Gehringer, E.F. Reuse of homework and test questions: When, why, and how to maintain security? In Proceedings of the 34th Annual Frontiers in Education; IEEE: New York, NY, USA, 2004; p. S1F-24. [Google Scholar]
- Brach, W.; Koš’ál, K.; Ries, M. Can Large Language Model Detect Plagiarism in Source Code? In Proceedings of the 2024 2nd International Conference on Foundation and Large Language Models (FLLM), Dubai, United Arab Emirates, 26–29 November 2024; pp. 370–377. [Google Scholar] [CrossRef] [Scilit]
- Mao, Y.; Ge, Y.; Fan, Y.; Xu, W.; Mi, Y.; Hu, Z.; Gao, Y. A survey on lora of large language models. Front. Comput. Sci. 2025, 19, 197605. [Google Scholar]
- Ebrahim, F.; Joy, M. Source Code Plagiarism Detection with Pre-Trained Model Embeddings and Automated Machine Learning. In Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing, Varna, Bulgaria, 4–6 September 2023; pp. 301–309. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hutter, F.; Kotthoff, L.; Vanschoren, J. Automated Machine Learning: Methods, Systems, Challenges; Springer Nature: Cham, Switzerland, 2019. [Google Scholar]
- Flores, E.; Rosso, P.; Moreno, L.; Villatoro-Tello, E. On the detection of source code re-use. In Proceedings of the 6th Annual Meeting of the Forum for Information Retrieval Evaluation, Bangalore, India, 5–7 December 2014; pp. 21–30. [Google Scholar]
- Ebrahim, F.; Joy, M. Semantic Similarity Search for Source Code Plagiarism Detection: An Exploratory Study. In Proceedings of the 2024 Innovation and Technology in Computer Science Education V. 1 (ITiCSE 2024), Milan, Italy, 8–10 July 2024; p. 7. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Sha, C.; Peng, X. An Empirical Study of Parameter-Efficient Fine-Tuning Methods for Pre-Trained Code Models. In Proceedings of the 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE); IEEE: New York, NY, USA, 2023; pp. 397–408. [Google Scholar]
- Li, X.L.; Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. arXiv 2021, arXiv:2101.00190. [Google Scholar]
- Liu, H.; Tam, D.; Muqeeth, M.; Mohta, J.; Huang, T.; Bansal, M.; Raffel, C.A. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. Adv. Neural Inf. Process. Syst. 2022, 35, 1950–1965. [Google Scholar] [CrossRef] [Scilit]
- He, J.; Zhou, C.; Ma, X.; Berg-Kirkpatrick, T.; Neubig, G. Towards a unified view of parameter-efficient transfer learning. arXiv 2021, arXiv:2110.04366. [Google Scholar]
- Zou, W.; Li, Q.; Ge, J.; Li, C.; Shen, X.; Huang, L.; Luo, B. A Comprehensive Evaluation of Parameter-Efficient Fine-Tuning on Software Engineering Tasks. arXiv 2023, arXiv:2312.15614. [Google Scholar]
- Saberi, I.; Fard, F.; Chen, F. Multilingual Adapter-based Knowledge Aggregation on Code Summarization for Low-Resource Languages. arXiv 2023, arXiv:2307.07854. [Google Scholar]
- Liu, S.; Keung, J.; Yang, Z.; Liu, F.; Zhou, Q.; Liao, Y. Delving into Parameter-Efficient Fine-Tuning in Code Change Learning: An Empirical Study. arXiv 2024, arXiv:2402.06247. [Google Scholar]
- Abu Talib, M.; Bou Nassif, A.; Azzeh, M.; Alesh, Y.; Afadar, Y. Parameter-efficient fine-tuning of pre-trained code models for just-in-time defect prediction. Neural Comput. Appl. 2024, 36, 16911–16940. [Google Scholar] [CrossRef] [Scilit]
- De Sousa, N.T.; Hasselbring, W. Javabert: Training a Transformer-based model for the java programming language. In Proceedings of the 2021 36th IEEE/ACM International Conference on Automated Software Engineering Workshops (ASEW); IEEE: New York, NY, USA, 2021; pp. 90–95. [Google Scholar]
- Wang, Y.; Le, H.; Gotmare, A.D.; Bui, N.D.; Li, J.; Hoi, S.C. Codet5+: Open code large language models for code understanding and generation. arXiv 2023, arXiv:2305.07922. [Google Scholar]
- Li, Z.; Lu, S.; Guo, D.; Duan, N.; Jannu, S.; Jenks, G.; Majumder, D.; Green, J.; Svyatkovskiy, A.; Fu, S.; et al. CodeReviewer: Pre-Training for Automating Code Review Activities. arXiv 2022, arXiv:2203.09095. [Google Scholar]
- Shen, Y.; Peng, M.; Zhang, F.; Wu, Q. Data race detection via few-shot parameter-efficient fine-tuning. J. Syst. Softw. 2025, 222, 112289. [Google Scholar] [CrossRef] [Scilit]
- Weyssow, M.; Zhou, X.; Kim, K.; Lo, D.; Sahraoui, H. Exploring parameter-efficient fine-tuning techniques for code generation with large language models. arXiv 2023, arXiv:2308.10462. [Google Scholar]
- Nijkamp, E.; Pang, B.; Hayashi, H.; Tu, L.; Wang, H.; Zhou, Y.; Savarese, S.; Xiong, C. Codegen: An open large language model for code with multi-turn program synthesis. arXiv 2022, arXiv:2203.13474. [Google Scholar]
- Roziere, B.; Gehring, J.; Gloeckle, F.; Sootla, S.; Gat, I.; Tan, X.E.; Adi, Y.; Liu, J.; Sauvestre, R.; Remez, T.; et al. Code llama: Open foundation models for code. arXiv 2023, arXiv:2308.12950. [Google Scholar]
- Lester, B.; Al-Rfou, R.; Constant, N. The power of scale for parameter-efficient prompt tuning. arXiv 2021, arXiv:2104.08691. [Google Scholar]
- Lu, J.; Yu, L.; Li, X.; Yang, L.; Zuo, C. LLaMA-Reviewer: Advancing Code Review Automation with Large Language Models through Parameter-Efficient Fine-Tuning. In Proceedings of the 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE); IEEE: New York, NY, USA, 2023; pp. 647–658. [Google Scholar]
- Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. Llama: Open and efficient foundation language models. arXiv 2023, arXiv:2302.13971. [Google Scholar]
- Hong, Y.; Tantithamthavorn, C.; Thongtanunam, P.; Aleti, A. Commentfinder: A simpler, faster, more accurate code review comments recommendation. In Proceedings of the 30th ACM joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, Singapore, 14–18 November 2022; pp. 507–519. [Google Scholar]
- Li, G.; Zhi, C.; Chen, J.; Han, J.; Deng, S. A Comprehensive Evaluation of Parameter-Efficient Fine-Tuning on Automated Program Repair. arXiv 2024, arXiv:2406.05639. [Google Scholar]
- Guo, D.; Zhu, Q.; Yang, D.; Xie, Z.; Dong, K.; Zhang, W.; Chen, G.; Bi, X.; Wu, Y.; Li, Y.; et al. DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence. arXiv 2024, arXiv:2401.14196. [Google Scholar]
- Dehghan, M.; Wu, J.J.; Fard, F.H.; Ouni, A. MergeRepair: An Exploratory Study on Merging Task-Specific Adapters in Code LLMs for Automated Program Repair. arXiv 2024, arXiv:2408.09568. [Google Scholar]
- Lozhkov, A.; Li, R.; Allal, L.B.; Cassano, F.; Lamy-Poirier, J.; Tazi, N.; Tang, A.; Pykhtar, D.; Liu, J.; Wei, Y.; et al. Starcoder 2 and the stack v2: The next generation. arXiv 2024, arXiv:2402.19173. [Google Scholar]
- Mishra, M.; Stallone, M.; Zhang, G.; Shen, Y.; Prasad, A.; Soria, A.M.; Merler, M.; Selvam, P.; Surendran, S.; Singh, S.; et al. Granite code models: A family of open foundation models for code intelligence. arXiv 2024, arXiv:2405.04324. [Google Scholar]
- Zhang, B.; Liang, P.; Zhou, X.; Zhou, X.; Lo, D.; Feng, Q.; Li, Z.; Li, L. A Comprehensive Evaluation of Parameter-Efficient Fine-Tuning on Method-Level Code Smell Detection. arXiv 2024, arXiv:2412.13801. [Google Scholar]
- Huang, K.; Zhang, J.; Bao, X.; Wang, X.; Liu, Y. Comprehensive Fine-Tuning Large Language Models of Code for Automated Program Repair. IEEE Trans. Softw. Eng. 2025, 51, 904–928. [Google Scholar] [CrossRef] [Scilit]
- Fried, D.; Aghajanyan, A.; Lin, J.; Wang, S.; Wallace, E.; Shi, F.; Zhong, R.; Yih, W.t.; Zettlemoyer, L.; Lewis, M. Incoder: A generative model for code infilling and synthesis. arXiv 2022, arXiv:2204.05999. [Google Scholar]
- Zhang, Q.; Chen, M.; Bukharin, A.; Karampatziakis, N.; He, P.; Cheng, Y.; Chen, W.; Zhao, T. Adalora: Adaptive budget allocation for parameter-efficient fine-tuning. arXiv 2023, arXiv:2303.10512. [Google Scholar]
- Afrin, S.; Call, J.; Nguyen, K.N.; Chaparro, O.; Mastropaolo, A. Resource-Efficient & Effective Code Summarization. arXiv 2025, arXiv:2502.03617. [Google Scholar]
- Abdin, M.; Aneja, J.; Awadalla, H.; Awadallah, A.; Awan, A.A.; Bach, N.; Bahree, A.; Bakhtiari, A.; Bao, J.; Behl, H.; et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv 2024, arXiv:2404.14219. [Google Scholar]
- Dettmers, T.; Pagnoni, A.; Holtzman, A.; Zettlemoyer, L. Qlora: Efficient finetuning of quantized llms. Adv. Neural Inf. Process. Syst. 2023, 36, 10088–10115. [Google Scholar] [CrossRef] [Scilit]
- Haque, M.Z.; Afrin, S.; Mastropaolo, A. A Systematic Literature Review of Parameter-Efficient Fine-Tuning for Large Code Models. arXiv 2025, arXiv:2504.21569. [Google Scholar]
- Sağlam, T. Mitigating Automated Obfuscation Attacks on Software Plagiarism Detection Systems. Ph.D. Thesis, Karlsruher Institut für Technologie (KIT), Karlsruhe, Germany, 2025. [Google Scholar] [CrossRef]
- Mariani, L.; Micucci, D. AuDeNTES: Automatic Detection of teNtative plagiarism according to a rEference Solution. ACM Trans. Comput. Educ. 2012, 12, 1–26. [Google Scholar] [CrossRef] [Scilit]
- Svajlenko, J.; Islam, J.F.; Keivanloo, I.; Roy, C.K.; Mia, M.M. Towards a big data curated benchmark of inter-project code clones. In Proceedings of the 2014 IEEE International Conference on Software Maintenance and Evolution; IEEE: New York, NY, USA, 2014; pp. 476–480. [Google Scholar]
- Krinke, J.; Ragkhitwetsagul, C. Bigclonebench considered harmful for machine learning. In Proceedings of the 2022 IEEE 16th International Workshop on Software Clones (IWSC); IEEE: New York, NY, USA, 2022; pp. 1–7. [Google Scholar]
- Krinke, J.; Ragkhitwetsagul, C. How the Misuse of a Dataset Harmed Semantic Clone Detection. arXiv 2025, arXiv:2505.04311. [Google Scholar]
- Mersha, M.; Lam, K.; Wood, J.; AlShami, A.; Kalita, J. Explainable artificial intelligence: A survey of needs, techniques, applications, and future direction. Neurocomputing 2024, 599, 128111. [Google Scholar] [CrossRef] [Scilit]
- Cao, S.; Sun, X.; Widyasari, R.; Lo, D.; Wu, X.; Bo, L.; Zhang, J.; Li, B.; Liu, W.; Wu, D.; et al. A systematic literature review on explainability for machine/deep learning-based software engineering research. arXiv 2024, arXiv:2401.14617. [Google Scholar]
- Lundberg, S.M.; Lee, S.I. A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems 30; Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2017; pp. 4765–4774. [Google Scholar]







| Level | Code 1 | Code 2 | Notes |
|---|---|---|---|
| L0 | int p2(int n){ return n * n; } | int p2(int n){ return n * n; } | The two code snippets are identical. |
| L1 | int p2(int n){ return n * n; } | // Function p2 int p2(int n){ return n * n; } | Code 2 has an additional comment and extra white spaces. |
| L2 | int p2(int n){ return n * n; } | int pow2(int n){ return n * n; } | Code 2 changes the function name from p2 to pow2. |
| L3 | int p2(int n){ return n * n; } | int p2(int n){ int a = n * n; return a; } | Code 2 creates a new variable a to represent n * n and then returns it. |
| L4 | int p2(int n){ return n * n; } | int p2(int n){ return nn(n); } int nn(int n){ return n * n; } | Code 2 creates an additional function to compute the power and calls it from p2. |
| L5 | int sumUpTo(int n){ int sum = 0; for (int i = 1; i <= n; i++){ sum += i; } return sum; } | int sumUpTo(int n){ int sum = 0, i = 1; while (i <= n){ sum += i; i++; } return sum; } | Code 2 changes the for loop into a while loop. |
| L6 | boolean isEven(int n){ return n % 2 == 0; } | boolean isEven(int n){ return (n & 1) == 0; } | Code 2 uses a different logic to decide whether a number is even or odd. |
| Dataset | Language | Plagiarised | Non-Plagiarised | Total Pairs |
|---|---|---|---|---|
| ConPlag1/ConPlag2 | Java | 256 | 655 | 911 |
| IR-Plag | Java | 365 | 95 | 460 |
| Progpedia-19 | Java | 91 | 2054 | 2145 |
| Dataset | Non-Plag. | L1 | L2 | L3 | L4 | L5 | L6 | Total |
|---|---|---|---|---|---|---|---|---|
| ConPlag1/ConPlag2 | 655 | 68 | 5 | 11 | 63 | 70 | 39 | 911 |
| IR-Plag | 95 | 61 | 57 | 65 | 60 | 59 | 63 | 460 |
| Model | Number of Parameters | Model Size | Base Architecture |
|---|---|---|---|
| CodeBERT | 125 M | 499 MB | RoBERTa |
| GraphCodeBERT | 125 M | 499 MB | RoBERTa |
| UniXcoder | 125 M | 504 MB | RoBERTa |
| PLBART | 140 M | 557 MB | BART |
| CodeT5-Small | 60 M | 242 MB | T5 |
| CodeBERTa | 84 M | 336 MB | RoBERTa |
| Method | Number of Parameters | Inference Cost | Complexity |
|---|---|---|---|
| Sequential Adapters | Extra bottleneck FFN | ||
| LoRA | None if merged |
| Dataset | Split | Plagiarised | Non-Plagiarised | Total |
|---|---|---|---|---|
| ConPlag1/ConPlag2 | Train | 176 | 461 | 637 |
| Validation | 38 | 99 | 137 | |
| Test | 42 | 95 | 137 | |
| IR-Plag | Train | 257 | 65 | 322 |
| Validation | 55 | 14 | 69 | |
| Test | 54 | 15 | 69 | |
| Progpedia-19 | Train | 64 | 1437 | 1501 |
| Validation | 12 | 310 | 322 | |
| Test | 15 | 307 | 322 |
| Dataset | Model | FFT | Houlsby | Pfeiffer | LoRA |
|---|---|---|---|---|---|
| ConPlag1 | CodeBERT | 0.8046 | 0.7179 | 0.6098 | 0.6250 |
| GraphCodeBERT | 0.8235 | 0.8049 | 0.8333 | 0.6869 | |
| UniXcoder | 0.8810 | 0.8276 | 0.7619 | 0.7736 | |
| CodeBERTa | 0.7857 | 0.7742 | 0.7586 | 0.7660 | |
| PLBART | 0.9176 | 0.8000 | 0.9398 | 0.7907 | |
| CodeT5 | 0.7959 | 0.9000 | 0.7778 | 0.7723 | |
| ConPlag2 | CodeBERT | 0.8293 | 0.7234 | 0.7527 | 0.6512 |
| GraphCodeBERT | 0.7654 | 0.7727 | 0.7234 | 0.6508 | |
| UniXcoder | 0.8387 | 0.8095 | 0.7917 | 0.7835 | |
| CodeBERTa | 0.8049 | 0.8095 | 0.8101 | 0.8293 | |
| PLBART | 0.9268 | 0.8966 | 0.9195 | 0.8864 | |
| CodeT5 | 0.7451 | 0.8000 | 0.7473 | 0.8000 | |
| IR-Plag | CodeBERT | 0.9908 | 0.9725 | 0.9636 | 0.9550 |
| GraphCodeBERT | 0.9907 | 1.0000 | 0.9818 | 0.9908 | |
| UniXcoder | 1.0000 | 1.0000 | 1.0000 | 0.9818 | |
| CodeBERTa | 0.9818 | 0.9818 | 0.9818 | 0.9908 | |
| PLBART | 1.0000 | 0.9908 | 1.0000 | 0.9908 | |
| CodeT5 | 0.9286 | 0.8852 | 0.9076 | 0.9725 | |
| Progpedia-19 | CodeBERT | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| GraphCodeBERT | 1.0000 | 1.0000 | 1.0000 | 1.0000 | |
| UniXcoder | 1.0000 | 0.9677 | 1.0000 | 1.0000 | |
| CodeBERTa | 1.0000 | 0.9677 | 0.9677 | 0.9677 | |
| PLBART | 1.0000 | 1.0000 | 1.0000 | 1.0000 | |
| CodeT5 | 0.9677 | 0.9091 | 0.9677 | 0.9375 |
| IR-Plag | |||||
| Tool/Model | P | R | # FPs | # FNs | |
| JPlag | 0.7925 | 0.7778 | 0.7850 | 11 | 12 |
| Dolos | 0.8136 | 0.8889 | 0.8496 | 11 | 6 |
| PLBART (FFT) | 1.0000 | 1.0000 | 1.0000 | 0 | 0 |
| PLBART (Pfeiffer) | 1.0000 | 1.0000 | 1.0000 | 0 | 0 |
| ConPlag1 | |||||
| Tool/Model | P | R | # FPs | # FNs | |
| JPlag | 0.6939 | 0.8095 | 0.7473 | 15 | 8 |
| Dolos | 0.5000 | 0.9286 | 0.6500 | 39 | 3 |
| PLBART (FFT) | 0.9070 | 0.9286 | 0.9176 | 4 | 3 |
| PLBART (Pfeiffer) | 0.9512 | 0.9286 | 0.9398 | 2 | 3 |
| ConPlag2 | |||||
| Tool/Model | P | R | # FPs | # FNs | |
| JPlag | 0.6379 | 0.8810 | 0.7400 | 21 | 5 |
| Dolos | 0.6786 | 0.9048 | 0.7755 | 18 | 4 |
| PLBART (FFT) | 0.9500 | 0.9048 | 0.9268 | 2 | 4 |
| PLBART (Pfeiffer) | 0.8889 | 0.9524 | 0.9195 | 5 | 2 |
| Progpedia-19 | |||||
| Tool/Model | P | R | # FPs | # FNs | |
| JPlag | 0.8824 | 1.0000 | 0.9375 | 2 | 0 |
| Dolos | 0.9375 | 1.0000 | 0.9677 | 1 | 0 |
| PLBART (FFT) | 1.0000 | 1.0000 | 1.0000 | 0 | 0 |
| PLBART (Pfeiffer) | 1.0000 | 1.0000 | 1.0000 | 0 | 0 |
| Dataset | Model | FFT | Houlsby | Pfeiffer | LoRA |
|---|---|---|---|---|---|
| ConPlag1 | CodeBERT | ||||
| GraphCodeBERT | |||||
| UniXcoder | |||||
| CodeBERTa | |||||
| PLBART | |||||
| CodeT5 | |||||
| ConPlag2 | CodeBERT | ||||
| GraphCodeBERT | |||||
| UniXcoder | |||||
| CodeBERTa | |||||
| PLBART | |||||
| CodeT5 | |||||
| IR-Plag | CodeBERT | ||||
| GraphCodeBERT | |||||
| UniXcoder | |||||
| CodeBERTa | |||||
| PLBART | |||||
| CodeT5 | |||||
| Progpedia-19 | CodeBERT | ||||
| GraphCodeBERT | |||||
| UniXcoder | |||||
| CodeBERTa | |||||
| PLBART | |||||
| CodeT5 |
| Comparison | Test Result | Interpretation |
|---|---|---|
| FFT vs. best PEFT | Wilcoxon: , 95% CI , | No significant difference. The strongest adapter choice is comparable to FFT in . |
| Houlsby vs. Pfeiffer vs. LoRA | Friedman: , | No significant overall difference among the three PEFT methods; no post hoc pairwise test is required. |
| Dataset | Model | FFT | Houlsby | Pfeiffer | LoRA |
|---|---|---|---|---|---|
| ConPlag1 | CodeBERT | ||||
| GraphCodeBERT | |||||
| UniXcoder | |||||
| CodeBERTa | |||||
| PLBART | |||||
| CodeT5 | |||||
| ConPlag2 | CodeBERT | ||||
| GraphCodeBERT | |||||
| UniXcoder | |||||
| CodeBERTa | |||||
| PLBART | |||||
| CodeT5 |
| (a) Without preprocessing | |||||
| Dataset | L | FFT | Houlsby | Pfeiffer | LoRA |
| ConPlag1 | 512 | ||||
| 768 | |||||
| 1024 | |||||
| 1280 | |||||
| ConPlag2 | 512 | ||||
| 768 | |||||
| 1024 | |||||
| 1280 | |||||
| (b) With preprocessing | |||||
| Dataset | L | FFT | Houlsby | Pfeiffer | LoRA |
| ConPlag1 | 512 | ||||
| 768 | |||||
| 1024 | |||||
| 1280 | |||||
| ConPlag2 | 512 | ||||
| 768 | |||||
| 1024 | |||||
| 1280 | |||||
| IR-Plag | |||
| Tool/Model | P | R | |
| JPlag | |||
| Dolos | |||
| UniXcoder (FFT) | |||
| UniXcoder (Pfeiffer) | |||
| ConPlag1 (Raw) | |||
| Tool/Model | P | R | |
| JPlag | |||
| Dolos | |||
| PLBART (FFT) | |||
| PLBART (Pfeiffer) | |||
| ModernBERT (Pfeiffer, pre., ) | |||
| ConPlag2 (Template-Free) | |||
| Tool/Model | P | R | |
| JPlag | |||
| Dolos | |||
| PLBART (FFT) | |||
| PLBART (Pfeiffer) | |||
| ModernBERT (Houlsby, pre., ) | |||
| Progpedia-19 | |||
| Tool/Model | P | R | |
| JPlag | |||
| Dolos | |||
| PLBART (FFT) | |||
| PLBART (Pfeiffer) | |||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ebrahim, F.; Joy, M. An Empirical Study of Fine-Tuning Pre-Trained Code Models and Adapters for the Classification of Source Code Plagiarism Instances. Appl. Sci. 2026, 16, 7156. https://doi.org/10.3390/app16147156
Ebrahim F, Joy M. An Empirical Study of Fine-Tuning Pre-Trained Code Models and Adapters for the Classification of Source Code Plagiarism Instances. Applied Sciences. 2026; 16(14):7156. https://doi.org/10.3390/app16147156
Chicago/Turabian StyleEbrahim, Fahad, and Mike Joy. 2026. "An Empirical Study of Fine-Tuning Pre-Trained Code Models and Adapters for the Classification of Source Code Plagiarism Instances" Applied Sciences 16, no. 14: 7156. https://doi.org/10.3390/app16147156
APA StyleEbrahim, F., & Joy, M. (2026). An Empirical Study of Fine-Tuning Pre-Trained Code Models and Adapters for the Classification of Source Code Plagiarism Instances. Applied Sciences, 16(14), 7156. https://doi.org/10.3390/app16147156

