Comparative Evaluation of GPT-4o, GPT-OSS-120B and Llama-3.1-8B-Instruct Language Models in a Reproducible CV-to-JSON Extraction Pipeline
Abstract
1. Introduction
Related Work
- We introduce a reproducible CV-to-JSON extraction pipeline that converts heterogeneous recruitment documents into structured, anonymized data while ensuring GDPR compliance.
- We conduct a controlled comparative evaluation of GPT-4o, GPT-OSS-120B and Llama-3.1-8B-Instruct under identical prompt templates and deterministic decoding conditions.
- We propose two novel evaluation metrics—schema completeness and content similarity—that jointly assess structural coverage and semantic fidelity of extracted outputs.
- We incorporate a deterministic anonymization layer operating directly on Markdown-formatted CVs, preventing any leakage of personal identifiers.
- We ensure full auditability and transparency, with fixed model versions, version-controlled prompts and byte-identical re-executions ensuring verifiable reproducibility.
2. Materials and Methods
2.1. Dataset Description
2.2. Conversion PDF to Markdown Pipeline
2.2.1. Layout Analysis Model
2.2.2. Table Structure Recognition
2.2.3. Optical Character Recognition (OCR)
2.2.4. Reading Order and Post-Processing
2.2.5. Serialization and Extensibility
2.3. CV Markdown Anonymization Protocol
2.3.1. Detection Strategy
2.3.2. Replacement Strategy
Two-Pass Sanitation and Quality Assurance (QA)
| Algorithm 1 Anonymize-Markdown for One CV |
| Require: : Markdown from Docling; salt ← SHA256(checksum) Ensure: : GDPR-compliant anonymized Markdown 1: segment() ▷ Divide into sections: Contact, Skills, Experience, Education, etc. 2: detect_pii(S) ▷ Detect emails, phones, URLs, IDs + light NER for PERSON, ORG, GPE 3: for entity in order: email → phone → url/id → PERSON/ORG/GPE do 4: make_replacement(e, ) ▷ According to Table 3; preserve case and format 5: replace_all(, e, r) 6: end for 7: detect_pii() 8: if then 9: harden(, ) ▷ Fallback masking and conservative redaction 10: end if 11: return |
2.3.3. Impact of De-Identification on Extraction Accuracy and Inference Efficiency
Extraction Quality
Inference Efficiency
2.4. Anonymized Markdown Sample
2.5. Methods: Language-Model Back-Ends
2.5.1. Shared Prompt Template and System Instructions
Determinism and Guardrails
2.5.2. GPT-4o (Azure)
2.5.3. Compute Environment for Open-Weight Models (Llama-3.1-8B-Instruct and GPT-OSS-120B)
2.5.4. GPT-OSS-120B (Open Source)
- , , : bytes MiB per token;
- , , : bytes MiB per token.
| Algorithm 2 Deterministic CV-to-JSON extraction with GPT-OSS-120B |
| Require: anonymized Markdown m, prompt template p, schema s Ensure: single JSON object j conforming to s 1: tokenize(p with m); shard model across n GPUs via TP; pre-allocate paged KV 2: set decoding: temperature , top-, nucleus off; disable speculative/beam decoding 3: generate(x) ▷ single pass; no streaming 4: longest_json_substring(); if parse(j) fails or then repair_and_retry() 5: return canonicalize(j) |
2.5.5. Llama-3.1-8B-Instruct (Open Source)
2.5.6. Comparative Summary of the Three LLM Models
2.6. Automated Evaluation of Extracted JSON
2.6.1. Reference-Based Completeness
- Unknown element type (empty example in the schema): We score the presence of any items in the reference list. The candidate matches if it also has a non-empty list.
- List of scalars: One point if the reference list contains at least one filled scalar. The candidate matches if it also contains any filled scalar (order and multiplicity are ignored).
- List of objects: We align by index over the first elements of the reference list and recurse element-wise.
2.6.2. Content Similarity
3. Results
3.1. Completeness of Extracted Sections
3.2. Content Similarity Against the Reference
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| API | Application Programming Interface |
| ARD-LoRA | Adaptive Rank-Determined Low-Rank Adaptation |
| AUC-PR | Area Under the Precision–Recall Curve |
| BERT | Bidirectional Encoder Representations from Transformers |
| BoW | Bag of Words |
| CEFR | Common European Framework of Reference for Languages |
| CI/CD | Continuous Integration/Continuous Deployment |
| CPU | Central Processing Unit |
| CTC | Connectionist Temporal Classification |
| CV | Curriculum Vitae |
| DETR | Detection Transformer |
| DocLayNet | Document Layout Network dataset |
| ETL | Extract–Transform–Load |
| F1 | F1-score (harmonic mean of precision and recall) |
| GDPR | General Data Protection Regulation |
| GPU | Graphics Processing Unit |
| JSON | JavaScript Object Notation |
| LLM | Large Language Model |
| LoRA | Low-Rank Adaptation |
| LoRA/QLoRA | Quantized Low-Rank Adaptation |
| MDPI | Multidisciplinary Digital Publishing Institute |
| NER | Named Entity Recognition |
| OCR | Optical Character Recognition |
| PII | Personally Identifiable Information |
| RAM | Random Access Memory |
| RBAC | Role-Based Access Control |
| RoLoRA | Rotated Low-Rank Adaptation |
| SSL | Secure Sockets Layer |
| t-SNE | t-distributed Stochastic Neighbor Embedding |
| TEDS | Tree Edit Distance Similarity |
| URL | Uniform Resource Locator |
| VM | Virtual Machine |
| YAML | Yet Another Markup Language |
| AIFI | Attention-based Intra-Scale Feature Interaction |
| CCFF | Cross-Scale Convolutional Feature Fusion |
| CRNN | Convolutional Recurrent Neural Network |
| FPGA | Field Programmable Gate Array |
| JSON-LD | JSON for Linked Data |
| MLP | Multi-Layer Perceptron |
| OCRopus | Open Source OCR System (Google) |
| Portable Document Format | |
| RT-DETR | Real-Time Detection Transformer |
| T5 | Text-to-Text Transfer Transformer |
| ViT | Vision Transformer |
Appendix A. System Instruction Listing
| Listing A1. System instruction used for all model back-ends. |
![]() |
Appendix B. User Prompt Template Listing
| Listing A2. Unified user prompt template (abridged, English) used across all model back-ends. |
![]() ![]() ![]() ![]() ![]() |
Appendix C. Dummy Markdown CV Listing
| Listing A3. Dummy Markdown CV (fictional data). |
![]() ![]() ![]() |
References
- Kurek, J.; Latkowski, T.; Bukowski, M.; Świderski, B.; Łępicki, M.; Baranik, G.; Nowak, B.; Zakowicz, R.; Dobrakowski, Ł. Zero-Shot Recommendation AI Models for Efficient Job–Candidate Matching in Recruitment Process. Appl. Sci. 2024, 14, 2601. [Google Scholar] [CrossRef]
- Łępicki, M.; Latkowski, T.; Antoniuk, I.; Bukowski, M.; Świderski, B.; Baranik, G.; Nowak, B.; Zakowicz, R.; Dobrakowski, Ł.; Act, B.; et al. Comparative Evaluation of Sequential Neural Network (GRU, LSTM, Transformer) Within Siamese Networks for Enhanced Job–Candidate Matching in Applied Recruitment Systems. Appl. Sci. 2025, 15, 5988. [Google Scholar] [CrossRef]
- Polat, F.; Tiddi, I.; Groth, P. Testing prompt engineering methods for knowledge extraction from text. Semant. Web 2025, 16, SW-243719. [Google Scholar] [CrossRef]
- Hu, Y.; Chen, Q.; Du, J.; Peng, X.; Keloth, V.K.; Zuo, X.; Zhou, Y.; Li, Z.; Jiang, X.; Lu, Z.; et al. Improving large language models for clinical named entity recognition via prompt engineering. J. Am. Med. Inform. Assoc. 2024, 31, 1812–1820. [Google Scholar] [CrossRef]
- Kumar, A.; Sharma, R.; Bedi, P. Towards Optimal NLP Solutions: Analyzing GPT and LLaMA-2 Models Across Model Scale, Dataset Size, and Task Diversity. Eng. Technol. Appl. Sci. Res. 2024, 14, 14219–14224. [Google Scholar] [CrossRef]
- Khairat, S.; Niu, T.; Geracitano, J.; Zhou, Z. Performance Evaluation of Popular Open-Source Large Language Models in Healthcare. Stud. Health Technol. Inform. 2025, 328, 215–219. [Google Scholar] [CrossRef]
- Chi, J.; Rouphail, Y.; Hillis, E.; Ma, N.; Nguyen, A.; Wang, J.; Hofford, M.; Gupta, A.; Lyons, P.G.; Wilcox, A.; et al. EchoLLM: Extracting echocardiogram entities with light-weight, open-source large language models. JAMIA Open 2025, 8, 92. [Google Scholar] [CrossRef]
- Ntinopoulos, V.; Rodriguez Cetina Biefer, H.; Tudorache, I.; Papadopoulos, N.; Odavic, D.; Risteski, P.; Haeussler, A.; Dzemali, O. Large language models for data extraction from unstructured and semi-structured electronic health records: A multiple model performance evaluation. BMJ Health Care Inform. 2025, 32, 101139. [Google Scholar] [CrossRef]
- Kim, M.S.; Chung, P.; Aghaeepour, N.; Kim, N. Information Extraction from Clinical Texts with Generative Pre-trained Transformer Models. Int. J. Med. Sci. 2025, 22, 1015–1028. [Google Scholar] [CrossRef]
- Biderman, D.; Portes, J.; Ortiz, J.J.G.; Paul, M.; Greengard, P.; Jennings, C.; King, D.; Havens, S.; Chiley, V.; Frankle, J.; et al. LoRA Learns Less and Forgets Less. arXiv 2024, arXiv:2405.09673. [Google Scholar] [CrossRef]
- Shinwari, H.U.K.; Usama, M. ARD-LoRA: Dynamic Rank Allocation for Parameter-Efficient Fine-Tuning of Foundation Models with Heterogeneous Adaptation Needs. IEEE Trans. Artif. Intell. 2025, 2025, 1–10. [Google Scholar] [CrossRef]
- Sheng, S.; Xu, Y.; Zhang, T.; Shen, Z.; Fu, L.; Ding, J.; Zhou, L.; Gan, X.; Wang, X.; Zhou, C. RepEval: Effective Text Evaluation with LLM Representation; Association for Computational Linguistics: Miami, FL, USA, 2024; pp. 7019–7033. [Google Scholar] [CrossRef]
- Jiang, Z.; Wu, P.; Liang, Z.; Chen, P.Q.; Yuan, X.; Jia, Y.; Tu, J.; Li, C.; Ng, P.H.F.; Li, Q. HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning. In Proceedings of the ACM Web Conference 2025; Association for Computing Machinery: New York, NY, USA, 2025; pp. 5505–5515. [Google Scholar] [CrossRef]
- Tahmasebi, S.; Nikzad, N.; Payberah, A.H.; Asgari-chenaghlu, M.; Matskin, M. Fact vs. Fiction: Are the Reportedly “Magical” LLM-Based Recommenders Reproducible? Lect. Notes Comput. Sci. 2025, 15575 LNCS, 79–94. [Google Scholar] [CrossRef]
- Digan, W.; Névéol, A.; Neuraz, A.; Wack, M.; Baudoin, D.; Burgun, A.; Rance, B. Can reproducibility be improved in clinical natural language processing? A study of 7 clinical NLP suites. J. Am. Med. Inform. Assoc. 2021, 28, 504–515. [Google Scholar] [CrossRef] [PubMed]
- Yoon, C.O.; Lee, W.; Jang, S.; Choi, K.; Jung, M.; Choi, D. Language, OCR, Form Independent (LOFI) Pipeline for Industrial Document Information Extraction; Association for Computational Linguistics: Miami, FL, USA, 2024; pp. 1056–1067. [Google Scholar] [CrossRef]
- Shrikaran, P.; Mamalaivasan, H.; Chitradevi, D.; Kumar, S.T.; Abishekwoolridge. A Secure On-Premises ETL Pipeline for Enterprise Data Warehousing: Integrating OCR and Local LLMs. In Proceedings of the 2025 8th International Conference on Computing Methodologies and Communication (ICCMC), Erode, India, 23–25 July 2025; pp. 1859–1864. [Google Scholar] [CrossRef]
- Xin, C.; Lu, Y.; Lin, H.; Zhou, S.; Zhu, H.; Wang, W.; Liu, Z.; Han, X.; Sun, L. Beyond Full Fine-tuning: Harnessing the Power of LoRA for Multi-Task Instruction Tuning; ELRA and ICCL: Torino, Italy, 2024; pp. 2307–2317. [Google Scholar]
- Huang, X.; Liu, Z.; Liu, S.Y.; Cheng, K.T. RoLoRA: Fine-Tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization. In Findings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP); Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 7563–7576. [Google Scholar] [CrossRef]
- Wang, Y.; Ivison, H.; Dasigi, P.; Hessel, J.; Khot, T.; Chandu, K.R.; Wadden, D.; MacMillan, K.; Smith, N.A.; Beltagy, I.; et al. How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources. In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023) — Datasets and Benchmarks Track; Curran Associates, Inc.: Red Hook, NY, USA, 2023; Available online: https://proceedings.neurips.cc/paper/2023/hash/ec6413875e4ab08d7bc4d8e225263398-Abstract.html (accessed on 22 December 2025).
- Uluirmak, B.A.; Kurban, R. Fine Tuning DeepSeek and Llama Large Language Models with LoRA. In Proceedings of the 2025 33rd Signal Processing and Communications Applications Conference (SIU), Istanbul, Turkiye, 25–28 June 2025. [Google Scholar] [CrossRef]
- Loukil, F.; Cadereau, S.; Verjus, H.; Galfre, M.; Salamatian, K.; Telisson, D.; Kembellec, Q.; Le Van, O. LLM-centric pipeline for information extraction from invoices. In Proceedings of the 2024 2nd International Conference on Foundation and Large Language Models (FLLM), Dubai, United Arab Emirates, 26–29 November 2024; pp. 569–575. [Google Scholar] [CrossRef]
- Zambrano, G. Case Law as Data: Prompt Engineering Strategies for Case Outcome Extraction with Large Language Models in a Zero-Shot Setting. Law Technol. Hum. 2024, 6, 80–101. [Google Scholar] [CrossRef]
- Beri, G.; Srivastava, V. Advanced Techniques in Prompt Engineering for Large Language Models: A Comprehensive Study. In Proceedings of the 2024 IEEE 4th International Conference on ICT in Business Industry & Government (ICTBIG), Indore, India, 13–14 December 2024. [Google Scholar] [CrossRef]
- Nazir, A.; Chakravarthy, T.K.; Cecchini, D.A.; Khajuria, R.; Sharma, P.; Mirik, A.T.; Kocaman, V.; Talby, D. LangTest: A comprehensive evaluation library for custom LLM and NLP models. Softw. Impacts 2024, 19, 100619. [Google Scholar] [CrossRef]
- Anghel, C.; Anghel, A.A.; Pecheanu, E.; Cocu, A.; Istrate, A.; Andrei, C.A. Diagnosing Bias and Instability in LLM Evaluation: A Scalable Pairwise Meta-Evaluator. Information 2025, 16, 652. [Google Scholar] [CrossRef]
- Ma, Y.; Qing, L.; Liu, J.; Kang, Y.; Zhang, Y.; Lu, W.; Liu, X.; Cheng, Q. From Model-Centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications; Association for Computational Linguistics: Bangkok, Thailand, 2024; pp. 2127–2137. [Google Scholar] [CrossRef]
- Li, Y.; Sun, Y. A fine-grained evaluation framework for language models: Combining pointwise grading and pairwise comparison. Inf. Process. Manag. 2026, 63, 104270. [Google Scholar] [CrossRef]
- Hsu, E.; Roberts, K. LLM-IE: A python package for biomedical generative information extraction with large language models. JAMIA Open 2025, 8, ooaf012. [Google Scholar] [CrossRef]
- Hein, D.; Christie, A.; Holcomb, M.; Xie, B.; Jain, A.; Vento, J.; Rakheja, N.; Shakur, A.H.; Christley, S.; Cowell, L.G.; et al. Iterative refinement and goal articulation to optimize large language models for clinical information extraction. npj Digit. Med. 2025, 8, 301. [Google Scholar] [CrossRef]
- Shanmugavelu, S.; Taillefumier, M.; Culver, C.; Hernandez, O.; Coletti, M.; Sedova, A. Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications. In Proceedings of the SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, Atlanta, GA, USA, 17–22 November 2024; pp. 170–179. [Google Scholar] [CrossRef]
- Sethi, R.J.; Gil, Y. Reproducibility in computer vision: Towards open publication of image analysis experiments as semantic workflows. In Proceedings of the 2016 IEEE 12th International Conference on e-Science (e-Science), Baltimore, MD, USA, 23–27 October 2017; pp. 343–348. [Google Scholar] [CrossRef]
- Ahn, D.H.; Lee, G.L.; Gopalakrishnan, G.; Rakamarić, Z.; Schulz, M.; Laguna, I. Overcoming extreme-scale reproducibility challenges through a unified, targeted, and multilevel toolset. In Proceedings of the 1st International Workshop on Software Engineering for High Performance Computing in Computational Science and Engineering, Halifax, NS, Canada, 7 November 2013; pp. 41–44. [Google Scholar] [CrossRef]
- Vrdoljak, J.; Boban, Z.; Vilović, M.; Kumrić, M.; Božić, J. A Review of Large Language Models in Medical Education, Clinical Decision Support, and Healthcare Administration. Healthcare 2025, 13, 603. [Google Scholar] [CrossRef]
- Schur, A.; Groenjes, S. Comparative Analysis for Open-Source Large Language Models. Commun. Comput. Inf. Sci. 2024, 1958 CCIS, 48–54. [Google Scholar] [CrossRef]
- Lehnen, N.C.; Kürsch, J.; Wichtmann, B.D.; Wolter, M.; Bendella, Z.; Bode, F.J.; Zimmermann, H.; Radbruch, A.; Vollmuth, P.; Dorn, F. Llama 3.1 405B Is Comparable to GPT-4 for Extraction of Data from Thrombectomy Reports—A Step Towards Secure Data Extraction. Clin. Neuroradiol. 2025, 35, 495–510. [Google Scholar] [CrossRef] [PubMed]
- Can, E.; Uller, W.; Vogt, K.; Doppler, M.C.; Busch, F.; Bayerl, N.; Ellmann, S.; Kader, A.; Elkilany, A.; Makowski, M.R.; et al. Large Language Models for Simplified Interventional Radiology Reports: A Comparative Analysis. Acad. Radiol. 2025, 32, 888–898. [Google Scholar] [CrossRef] [PubMed]
- Elnashar, A.; White, J.; Schmidt, D.C. Enhancing structured data generation with GPT-4o evaluating prompt efficiency across prompt styles. Front. Artif. Intell. 2025, 8, 1558938. [Google Scholar] [CrossRef]
- Liu, Y.; Duan, H.; Zhang, Y.; Li, B.; Zhang, S.; Zhao, W.; Yuan, Y.; Wang, J.; He, C.; Liu, Z.; et al. MMBench: Is Your Multi-modal Model an All-Around Player? Lect. Notes Comput. Sci. 2025, 15064 LNCS, 216–233. [Google Scholar] [CrossRef]
- Wang, Z.; Zhou, Y.; Wei, W.; Lee, C.Y.; Tata, S. VRDU: A Benchmark for Visually-rich Document Understanding. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Long Beach, CA, USA, 4 August 2023; pp. 5184–5193. [Google Scholar] [CrossRef]
- Lee, T.; Tu, H.; Wong, C.H.; Zheng, W.; Zhou, Y.; Mai, Y.; Roberts, J.S.; Yasunaga, M.; Yao, H.; Xie, C.; et al. VHELM: A Holistic Evaluation of Vision Language Models. arXiv 2024, arXiv:2410.07112. [Google Scholar] [CrossRef]
- Yao, Y.; Lin, Z.; Liu, X.; Li, Y. Document Information Extraction in Engineering Reports through Prompt Engineering with Large Language Models. In Proceedings of the 2025 IEEE 6th International Seminar on Artificial Intelligence, Networking and Information Technology (AINIT), Shenzhen, China, 11–13 April 2025; pp. 1969–1972. [Google Scholar] [CrossRef]
- Team, D.S. Docling Technical Report. Technical report. arXiv 2024, arXiv:2408.09869. [Google Scholar] [CrossRef]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs Beat YOLOs on Real-time Object Detection. arXiv 2023, arXiv:2304.08069. [Google Scholar]
- Lv, W.; Zhao, Y.; Chang, Q.; Huang, K.; Wang, G.; Liu, Y. RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer. arXiv 2024. [Google Scholar] [CrossRef]
- Maitin, A.M.; Nogales, A.; Fernández-Rincón, S.; Aranguren, E.; Cervera-Barba, E.; Denizon-Arranz, S.; Mateos-Rodríguez, A.; García-Tejedor, A.J. Application of large language models in clinical record correction: A comprehensive study on various retraining methods. J. Am. Med. Inform. Assoc. 2025, 32, 341–348. [Google Scholar] [CrossRef]
- Gogani-Khiabani, S.; Trivedi, A.; Chyi, S.; Tizpaz-Niari, S. Performance of LLMs on VITA test: Potential for AI-assisted tax returns for low income taxpayers. Artif. Intell. Law, 2025; in press. [Google Scholar] [CrossRef]
- Yin, J.; Bose, A.; Cong, G.; Lyngaas, I.; Anthony, Q. Comparative Study of Large Language Model Architectures on Frontier. In Proceedings of the 2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS), San Francisco, CA, USA, 27–31 May 2024; pp. 556–569. [Google Scholar] [CrossRef]
- Yang, W.; Some, L.; Bain, M.; Kang, B. A comprehensive survey on integrating large language models with knowledge-based methods. Knowl.-Based Syst. 2025, 318, 113503. [Google Scholar] [CrossRef]
- Gajulamandyam, D.K.; Veerla, S.; Emami, Y.; Lee, K.; Li, Y.; Mamillapalli, J.S.; Shim, S. Domain Specific Finetuning of LLMs Using PEFT Techniques. In Proceedings of the 2025 IEEE 15th Annual Computing and Communication Workshop and Conference (CCWC), Las Vegas, NV, USA, 6–8 January 2025; pp. 484–490. [Google Scholar] [CrossRef]
- Guo, Y.; Mousavi, S.S.; Ge, Y.; Baskaran, M.; Sameni, R.; Sarker, A. Leveraging few-shot learning and large language models for analyzing blood pressure variations across biological sex from scientific literature. Comput. Biol. Med. 2025, 198, 111128. [Google Scholar] [CrossRef]

| ISO Code | Language | Count | Percent (%) |
|---|---|---|---|
| en | English | 2041 | 89.52 |
| pl | Polish | 239 | 10.48 |
| Total | 2280 | 100.00 |
| Model/Component | Primary Function |
|---|---|
| RT-DETR/RT-DETRv2 | Transformer-based detector for page layout analysis; identifies paragraphs, figures, captions, tables. |
| TableFormer | Vision Transformer that predicts logical table structure (rows, columns, merged cells) aligned with geometric tokens. |
| EasyOCR/Tesseract | OCR back-ends for scanned or image-based pages. |
| Reading-Order Model | Determines the reading sequence using spatial/geometric relationships. |
| Post-Processing Heuristics | Caption linkage, heading hierarchy, language detection, text normalization. |
| Serialization Module | Converts the assembled document to Markdown and JSON. |
| Category | Policy (Deterministic Within a CV) |
|---|---|
| Personal names (PERSON) | Replace with synthetic, human-readable pseudonyms drawn from a fixed pool; stable per-entity within the CV. Preserve capitalization (e.g., Alex Novak). No mapping retained. |
| Emails | Replace local-part with a token derived from salted hash; domain forced to a documentation-only domain (example.com); e.g., first.last@example.com. |
| Phone numbers | Preserve country code if present; replace subscriber number with a non-dialable digit pattern of the same length (e.g., +48 600 000 000). Spacing retained. |
| Online profiles (LinkedIn, GitHub, personal sites) | Normalize to documentation/example endpoints while keeping labels (e.g., https://example.com/in/<token>, https://github.com/example-user). |
| Postal addresses | Remove street/number; keep coarse city/country granularity with synthetic city names (e.g., Warsaw, PL). |
| Employers/Organizations (ORG) | Replace with synthetic but domain-plausible names (e.g., EuroFin Bank, ThermoTech R&D Center, SkyTrip Group); stable within the CV. |
| Project/product names | Replace with neutral descriptors (e.g., Mobile Banking App, Travel Booking Platform) or synthetic names; maintain role context. |
| Dates | Keep month–year granularity (MM.YYYY or Month YYYY). Day-of-month and exact durations are not increased in precision. |
| Government/unique IDs (e.g., PESEL, passport) | Fully redact and replace with fixed-length masks (XXXXXXXXXXXX). |
| Images/photos | Remove and replace with an HTML comment placeholder <!– image –> (see Listing A3 in Appendix C). |
| Free-text quotations | No change unless direct identifiers occur; if so, apply the same policies as above. |
| Component | Value |
|---|---|
| GPUs (count × model) | NVIDIA A100-SXM4-40GB |
| Per-GPU framebuffer (reported) | 40,960 MiB (=40 GiB ) |
| Total framebuffer across 8 GPUs | 40,960 MiB MiB (=320 GiB) |
| Driver version | 570.195.03 |
| CUDA toolkit/runtime | 12.8 |
| Persistence mode | On (all GPUs) |
| Acronym | Meaning/Role |
|---|---|
| TP | Tensor Parallelism: shard large weight tensors of a layer across GPUs; all GPUs execute the same layer in lockstep. |
| DP | Data Parallelism: replicate the whole model on multiple workers and split the batch; used only for throughput, not within a single decode step. |
| PP | Pipeline Parallelism: place consecutive layer blocks on different devices; increases inter-stage latency; not used in our default runs. |
| SP | Sequence Parallelism: split the sequence dimension to reduce activation memory; optional, not required for our batch sizes. |
| KV cache | Key/Value cache: per-layer per-token tensors retained to enable incremental decoding. |
| RoPE | Rotary Position Embedding: relative position encoding commonly used in modern decoder-only LLMs. |
| MQA/GQA | Multi-/Grouped-Query Attention: share K/V projections across attention heads to reduce KV memory; support is runtime-dependent. |
| PagedAttention | Paged KV allocator: manages KV cache in fixed-size pages to limit fragmentation under long contexts. |
| EOS | End-of-Sequence token used to terminate generation. |
| VRAM | GPU device memory (a.k.a. framebuffer memory) holding weights and caches. |
| INTk | k-bit integer quantization of weights (e.g., INT8, INT4); activations remained FP16 in our setup. |
| Precision | Bytes/Param | Total Weights | Per-GPU (TP = 8) |
|---|---|---|---|
| FP16 (half) | 2 | 240 GB GiB | GiB |
| INT8 (weight-only) | 1 | 120 GB GiB | GiB |
| INT4 (weight-only) | 0.5 | 60 GB GiB | GiB |
| Aspect | GPT-4o | GPT-OSS-120B | Llama-3.1-8B-Instruct |
|---|---|---|---|
| Architecture | Proprietary, large, autoregressive transformer | Open-source, large, transformer | Open-source, smaller, transformer |
| Accuracy (CV-to-JSON extraction) | Highest, robust | High, improvable with fine-tuning | Competitive with prompt engineering |
| Reproducibility | Challenging (closed) | Moderate (open, containerizable) | High (open, local deployment) |
| Fine-tuning | Flexible, effective | Effective with LoRA/QLoRA | Highly effective, especially with LoRA |
| Prompt Engineering Impact | Moderate | Moderate | High (critical for performance) |
| Cost and Deployment | Expensive, cloud | Moderate, local/cloud | Low, local possible |
| Tokenization Bias | Present, mitigated by scale | Present, more pronounced | Present, can be mitigated with adaptation |
| Best Use Case | Highest accuracy, less privacy concern | Customizable, large-scale, privacy-aware | Cost-effective, privacy-critical, adaptable |
| JSON Section | GPT-4o | GPT-OSS-120B | Llama-3.1-8B-Instruct |
|---|---|---|---|
| analytics | 100 | 25.97 | 03.77 |
| awards | 100 | 69.67 | 51.75 |
| certifications | 100 | 83.64 | 69.91 |
| compatibility | 100 | 45.68 | 86.69 |
| conferences_talks | 100 | 43.83 | 28.33 |
| contact | 100 | 94.36 | 94.00 |
| education | 100 | 91.40 | 83.33 |
| languages | 100 | 85.10 | 91.50 |
| legal | 100 | 73.70 | 91.93 |
| memberships | 100 | 50.91 | 14.32 |
| patents | 100 | 20.00 | 16.67 |
| person | 100 | 94.09 | 90.03 |
| preferences | 100 | 97.24 | 98.95 |
| projects | 100 | 74.27 | 73.10 |
| publications | 100 | 66.13 | 48.20 |
| references | 100 | 48.04 | 39.85 |
| skills | 100 | 42.71 | 68.17 |
| source | 100 | 74.10 | 00.00 |
| volunteering | 100 | 63.22 | 50.99 |
| work_experience | 100 | 86.48 | 84.55 |
| Model | Avg. Rel. Completeness (%) | vs. Ref (pp) | Files (n) |
|---|---|---|---|
| GPT-4o | 100.00 | 0.00 | 2280 |
| GPT-OSS-120B | 73.52 | −26.48 | 2280 |
| Llama-3.1-8B-Instruct | 79.21 | −20.79 | 2280 |
| Section | GPT-4o | GPT-OSS-120B | Llama-3.1-8B-Instruct |
|---|---|---|---|
| analytics | 100 | 23.21 | 01.61 |
| awards | 100 | 62.36 | 51.75 |
| certifications | 100 | 78.46 | 70.06 |
| compatibility | 100 | 40.78 | 85.93 |
| conferences_talks | 100 | 35.93 | 25.56 |
| contact | 100 | 90.69 | 90.37 |
| education | 100 | 84.59 | 83.92 |
| languages | 100 | 83.50 | 79.96 |
| legal | 100 | 76.95 | 91.73 |
| memberships | 100 | 46.85 | 14.54 |
| patents | 100 | 1.75 | 6.10 |
| person | 100 | 92.54 | 65.56 |
| preferences | 100 | 99.61 | 99.50 |
| projects | 100 | 62.34 | 56.52 |
| publications | 100 | 58.28 | 44.06 |
| references | 100 | 40.20 | 34.74 |
| skills | 100 | 18.88 | 42.19 |
| source | 100 | 75.95 | 0.00 |
| volunteering | 100 | 51.47 | 42.66 |
| work_experience | 100 | 71.21 | 61.91 |
| Model | Avg. Rel. Content Similarity (%) | vs. Ref (pp) | Files (n) |
|---|---|---|---|
| GPT-4o | 100.00 | 0.00 | 2280 |
| GPT-OSS-120B | 72.44 | −27.56 | 2280 |
| Llama-3.1-8B-Instruct | 58.66 | −41.34 | 2280 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Nawalny, M.; Łępicki, M.; Latkowski, T.; Bujak, S.; Bukowski, M.; Świderski, B.; Baranik, G.; Nowak, B.; Zakowicz, R.; Dobrakowski, Ł.; et al. Comparative Evaluation of GPT-4o, GPT-OSS-120B and Llama-3.1-8B-Instruct Language Models in a Reproducible CV-to-JSON Extraction Pipeline. Appl. Sci. 2026, 16, 217. https://doi.org/10.3390/app16010217
Nawalny M, Łępicki M, Latkowski T, Bujak S, Bukowski M, Świderski B, Baranik G, Nowak B, Zakowicz R, Dobrakowski Ł, et al. Comparative Evaluation of GPT-4o, GPT-OSS-120B and Llama-3.1-8B-Instruct Language Models in a Reproducible CV-to-JSON Extraction Pipeline. Applied Sciences. 2026; 16(1):217. https://doi.org/10.3390/app16010217
Chicago/Turabian StyleNawalny, Marcin, Mateusz Łępicki, Tomasz Latkowski, Sebastian Bujak, Michał Bukowski, Bartosz Świderski, Grzegorz Baranik, Bogusz Nowak, Robert Zakowicz, Łukasz Dobrakowski, and et al. 2026. "Comparative Evaluation of GPT-4o, GPT-OSS-120B and Llama-3.1-8B-Instruct Language Models in a Reproducible CV-to-JSON Extraction Pipeline" Applied Sciences 16, no. 1: 217. https://doi.org/10.3390/app16010217
APA StyleNawalny, M., Łępicki, M., Latkowski, T., Bujak, S., Bukowski, M., Świderski, B., Baranik, G., Nowak, B., Zakowicz, R., Dobrakowski, Ł., Oczeretko, A., Sadowski, P., Szlaga, K., Kubica, B., & Kurek, J. (2026). Comparative Evaluation of GPT-4o, GPT-OSS-120B and Llama-3.1-8B-Instruct Language Models in a Reproducible CV-to-JSON Extraction Pipeline. Applied Sciences, 16(1), 217. https://doi.org/10.3390/app16010217










