A Modular Evaluation of AI-Assisted Clinical Documentation
Abstract
1. Introduction
- 1.
- RQ1. How do contemporary ASR–LLM configurations differ in transcription accuracy, medical terminology preservation, and downstream report-generation performance in a controlled multilingual benchmark?
- 2.
- RQ2. What patterns of operational use are observed in finalized-export traces from routine, clinician-supervised deployments?
- 3.
- RQ3. What practical evaluation considerations emerge when controlled benchmarking and deployment-grounded usage traces are interpreted as complementary, but distinct, forms of evidence?
2. System Architecture
2.1. Conversion of Audio to Text
2.2. Generation of the Clinical Report
3. Related Work
3.1. Evolution of Clinical Documentation and AI Scribes
3.2. Core Technologies: ASR and LLMs
3.3. Existing Systems and Landscape
3.4. Challenges and Research Gaps
4. Experiments
4.1. Dataset Generation
4.1.1. Dialogue Generation
4.1.2. Audio Generation
4.1.3. Dataset Characteristics
4.2. Experimental Setup
4.2.1. ASR Model Comparison
4.2.2. LLM Post-Processing
4.2.3. Prompting and Structured Evaluation
4.2.4. Multilingual and Length Variations
5. Results
5.1. Speech Recognition Results
5.2. Language Model Results for Report Generation
5.3. Safety-Oriented Quality Metrics
5.4. Economic Results Under Study Assumptions
5.5. Performance Visualizations and Trade-Offs
5.6. Cross-Language and Medical Specialty Results
5.7. Statistical Validation and Pilot-Stage Feasibility
6. Discussion
6.1. Key Findings and Clinical Implications
6.2. Multilingual Capabilities and Global Healthcare Implications
6.3. Clinical Safety and Quality Assurance
6.4. Economic and Scalability Considerations
6.5. Methodological Considerations for Medical NLP
7. Real-World Context and Implications
7.1. Clinical Relevance and Motivation
7.2. Evidence Gaps in Current Literature
7.3. Analytical Framework
7.4. Contribution of the Present Study
7.5. Observed Usage Patterns During the 120-Day Deployment Period (November 2025–February 2026)
7.5.1. Specialty-Level Usage Patterns
7.5.2. Concentration Effects and High-Intensity Usage
7.5.3. Sensitivity Analysis and Observational Bias
7.5.4. Interpretation of Finalized Exports as a Behavioral Signal
7.6. Summary Interpretation and Human-in-the-Loop Framing
8. Future Work and Limitations
8.1. Future Research Directions
8.1.1. Technical Robustness and Real-World Validation
8.1.2. Advanced Model Development
8.1.3. Voice as a Clinical Signal Beyond Transcription
8.1.4. Multilingual and Cross-Cultural Expansion
8.2. Limitations and Methodological Considerations
8.2.1. Dataset, Evaluation, and Generalizability Limits
8.2.2. Operational, Data, and Implementation Constraints
9. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| API | Application Programming Interface |
| ASR | Automatic Speech Recognition |
| BERT | Bidirectional Encoder Representations from Transformers |
| EHR | Electronic Health Record |
| FHIR | Fast Healthcare Interoperability Resources |
| GCP | Google Cloud Platform |
| HL7 | Health Level Seven |
| IRB | Institutional Review Board |
| LATAM | Latin America |
| LLM | Large Language Model |
| M-WER | Medical Word Error Rate |
| ML | Machine Learning |
| ModernBERT | Modern Bidirectional Encoder Representations from Transformers model family |
| NLP | Natural Language Processing |
| ROUGE-L | Recall-Oriented Understudy for Gisting Evaluation–Longest Common Subsequence |
| SNOMED CT | Systematized Nomenclature of Medicine–Clinical Terms |
| SOAP | Subjective, Objective, Assessment, and Plan |
| TTS | Text-to-Speech |
| WER | Word Error Rate |
Appendix A. Experimental Methodology and Implementation Details
Appendix A.1. Dataset Composition and Medical Specialties
| Medical Specialty | Medical Condition | Clinical Context |
|---|---|---|
| Cardiology | Chest pain and shortness of breath | Heart health evaluation |
| Dentistry | Tooth sensitivity and bleeding gums | Regular dental checkup and cleaning |
| Dermatology | Unusual mole and skin rash | Skin cancer screening |
| ENT (Otolaryngology) | Chronic sinusitis and hearing loss | ENT evaluation |
| General Practice | Persistent cough and fever | Follow-up visit for respiratory symptoms |
| Neurology | Recurring migraines and dizziness | Neurological consultation |
| Ophthalmology | Gradual vision loss and eye strain | Annual eye examination |
| Orthopedics | Chronic knee pain after sports injury | Post-injury assessment |
| Pediatrics | Child with recurring ear infections | Pediatric checkup |
| Psychiatry | Anxiety and sleep disorders | Mental health consultation |
Appendix A.2. Example Structure of a Generated Clinical Dialogue
Scenario: Specialty: Cardiology Condition: Chest pain and shortness of breath Context: Heart health evaluation Language: English
Dialogue: [neutral] DOCTOR: Good morning, Mr. Davies. I’m Dr. Anya Sharma, your cardiologist. Please, have a seat. What brings you in today?
[concerned] PATIENT: Good morning, Doctor. I have been experiencing chest pain and shortness of breath for the past few weeks, and it’s getting worse.
[professional] DOCTOR: I understand. Can you describe the chest pain? Where exactly do you feel it, and what does it feel like? Is it sharp, dull, or a pressure?
[anxious] PATIENT: It’s mostly in the center of my chest, like a heavy pressure. Sometimes it radiates to my left arm. It feels like someone is sitting on my chest. The shortness of breath comes on with even mild exertion, like walking up the stairs.
[professional] DOCTOR: Okay. Furthermore, how long does the chest pain typically last? Are there any specific activities that trigger it?
[anxious] PATIENT: It usually lasts for about 10-15 min. It seems to come on when I’m stressed or after I have eaten a big meal. Sometimes it happens just randomly, too.
[professional] DOCTOR: Do you have any other symptoms, such as palpitations, dizziness, sweating, or nausea?
[anxious] PATIENT: Yes, sometimes I feel a bit dizzy, and I have had some episodes of sweating, especially at night.
[professional] DOCTOR: Alright. What is your medical history? Do you have any pre-existing conditions, such as hypertension, diabetes, or high cholesterol? Furthermore, what medications are you currently taking?
[anxious] PATIENT: I was diagnosed with hypertension about five years ago, and I’m taking Lisinopril for it. My cholesterol is also a bit high, but I’m not on any medication for that yet. My father had a heart attack when he was in their early 60s.
[professional] DOCTOR: Thank you for that information. Now, I’m going to perform a physical examination. I will listen to your heart and lungs, check your blood pressure, and feel your pulse.
[neutral] DOCTOR: (Performing the examination) Okay, your blood pressure is a bit elevated, 150/90. Your heart sounds are regular, but I do hear a slight murmur. Your lungs are clear. There is no edema in your lower extremities. Your pulse is 88 bpm and regular.
[professional] DOCTOR: Based on your symptoms, medical history, and the physical examination, I suspect you may have angina, possibly related to coronary artery disease. However, we need further testing to confirm this diagnosis. I’d like to order an electrocardiogram (ECG), a stress test, and some blood work, including a lipid panel and cardiac enzyme levels.
[anxious] PATIENT: A stress test? What does that involve?
[reassuring] DOCTOR: The stress test will involve walking on a treadmill while we monitor your heart’s electrical activity. This will help us see how your heart responds to exertion and identify any potential blockages in your arteries. The ECG is a simple test to record the electrical activity of your heart at rest.
[professional] DOCTOR: After we get the results of these tests, we’ll discuss the findings and potential treatment options. These may include lifestyle modifications like diet and exercise, medication to manage your blood pressure and cholesterol, or, in more severe cases, procedures such as angioplasty or bypass surgery.
[reassuring] DOCTOR: For now, I’d recommend that you avoid strenuous activity and try to manage your stress levels. Please follow a heart-healthy diet, low in saturated fat, cholesterol, and sodium. Furthermore, it’s important to quit smoking if you are a smoker.
[professional] DOCTOR: I will have my nurse schedule you for the ECG, stress test, and blood work. We should have the results back within a week. Then, we can schedule a follow-up appointment to discuss them.
[anxious] PATIENT: Okay, thank you, Doctor. I appreciate your help.
[reassuring] DOCTOR: You’re welcome, Mr. Davies. Try not to worry too much. We’ll get to the bottom of this and develop a plan to manage your symptoms and protect your heart health. Please call the office if your symptoms worsen or if you have any questions before your follow-up appointment.
[neutral] DOCTOR: Take care, and I will see you soon.
[neutral] PATIENT: Thank you, Doctor. Goodbye.
Appendix A.3. Language Model Prompts for Medical Report Generation
- Generate a concise clinician-style report in the source language;
- Preserve the provided section headings and maintain SOAP-consistent organization;
- Include diagnoses, medications, and management details only when explicitly supported by the transcript;
- Avoid inventing information, recommendations, or clinical findings not present in the dialogue; and
- In Spanish, include all requested sections and explicitly mark missing content when needed.
Appendix A.4. Automated Evaluation Methodology with Gemini-2.0-Flash
Appendix A.4.1. Evaluation Model Configuration
- Model: Gemini-2.0-flash.
- Temperature: 0.1 (low temperature for consistent evaluation).
- Response format: Structured JSON output.
- Response schema: Pydantic BaseModel with enforced numerical outputs.
Appendix A.4.2. Clinical Quality Evaluation Prompt
Analyze and compare the original doctor-patient dialogue with the generated medical report. Provide numerical scores between 0.0 and 1.0 for each metric:
Original Dialogue:
{reference_text}
Generated Medical Report:
{medical_report}
Evaluate for: 1. Factual Accuracy - The extent to which all statements in the generated report accurately reflect the information in the original dialogue. 2. Information Coverage - The degree to which all key medical information points from the original dialogue are captured in the generated report. 3. Clinical Relevance - The degree to which the generated report focuses on medically important details and avoids irrelevant information. 4. Hallucination Score - A score representing the absence of added or fabricated information in the generated report. A score of 1.0 indicates no hallucinations. 5. Terminology Accuracy - The correctness and appropriateness of the medical terms used in the generated report.
Return as JSON with the score for each metric.
Appendix A.4.3. Structured Output Schema
Appendix A.5. Economic Analysis and Pricing Structure
Appendix A.5.1. Cost Components
- 1.
- Transcription Prices: Provider price schedules applied per minute of audio processed;
- 2.
- Processing Prices: LLM inference prices applied per token for input and output;
- 3.
- Infrastructure Costs: For locally deployed models (MedGemma), amortized deployment costs.
Appendix A.5.2. Cost Statistics
| Cost Metric | Value (USD) |
|---|---|
| Mean cost per consultation | $0.0620 |
| Median cost per consultation | $0.0297 |
| Minimum cost per consultation | $0.0095 |
| Maximum cost per consultation | $0.3830 |
| Standard deviation | $0.0891 |
Appendix A.6. Statistical Analysis Methodology
Appendix A.6.1. Statistical Tests Employed
| Contrast | W | Exact p | Bonferroni p | r | |
|---|---|---|---|---|---|
| Voxtral vs. Whisper | 20 | 0.0 | 0.877 | ||
| Voxtral vs. GPT-4o Transcribe | 20 | 0.0 | 0.877 | ||
| Voxtral vs. GCP | 10 | 0.0 | 0.001953 | 0.005859 | 0.886 |
Appendix A.6.2. Illustrative Pilot-Stage Reference Thresholds
- WER Improvement: 0.05 absolute improvement used as a practical reference point.
- Medical WER: 0.03 absolute improvement used as a practical reference point for medical terminology.
- Clinical Quality Metrics: 0.02 improvement on a 0–1 scale used as a practical reference point.
- Cost-Effectiveness: 10% improvement in the study-specific efficiency ratio used as a practical reference point.
Appendix A.7. Reproducibility and Data Availability
Appendix A.7.1. Experimental Reproducibility
Appendix A.7.2. Code and Configuration Availability
Appendix A.8. Clinical Safety Metrics: Hallucination Detection and Terminology Accuracy
Supplementary Pairwise Comparisons
| Metric | Contrast | Exact p | Bonferroni p | Result | |
|---|---|---|---|---|---|
| Hallucination score | GPT-4o vs. MedGemma | 0.0039 | 0.0117 | 0.0167 | Significant |
| Hallucination score | Gemini 1.5 Pro vs. GPT-4o | 0.0543 | 0.1629 | 0.0167 | Not significant |
| Hallucination score | Gemini 1.5 Pro vs. MedGemma | 0.3501 | 1.0000 | 0.0167 | Not significant |
| Terminology accuracy | GPT-4o vs. MedGemma | 0.0741 | 0.2223 | 0.0167 | Not significant |
| Terminology accuracy | Gemini 1.5 Pro vs. GPT-4o | 0.2580 | 0.7740 | 0.0167 | Not significant |
| Terminology accuracy | Gemini 1.5 Pro vs. MedGemma | 0.4347 | 1.0000 | 0.0167 | Not significant |
References
- Wenger, N.; Doyle, B.J. For Whom the Note Scrolls: A Brief History of the Medical Record’s Transition from a Tool for Physicians to a Bill. Ann. Intern. Med. 2024, 177, 566–569. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Asch, D.A.; Asch, R.M.; Asch, J.; Asch, S.M.; Asch, S.M. Clinical Documentation in the 21st Century: Executive Summary of a Policy Position Paper from the American College of Physicians. Ann. Intern. Med. 2015, 162, 797–798. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Strong, P.; Attal-Singer, J.; Singer, D.R.; Marcus, A.; Carew, A.; Althoff, A.; Singh, H. Enhancing clinical documentation with ambient artificial intelligence: A quality improvement survey assessing clinician perspectives on work burden, burnout, and job satisfaction. JAMIA Open 2025, 8, ooaf013. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Saadat, S.; Khalilizad, M.; Qorbani, M.; Hemmat, A.; Hariri, S. Enhancing Clinical Documentation with AI: Reducing Errors, Improving Interoperability, and Supporting Real-Time Note-Taking. InfoSci. Trends 2025, 2, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Lee, C.; Britto, S.; Diwan, K. Evaluating the Impact of Artificial Intelligence (AI) on Clinical Documentation Efficiency and Accuracy Across Clinical Settings: A Scoping Review. Cureus 2024, 16, e73994. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Leong, H.; Gao, Y.; Ji, S.; Zhang, Y.; Pamuksuz, U. Efficient Fine-Tuning of Large Language Models for Automated Medical Documentation. In Proceedings of the 2024 4th International Conference on Digital Society and Intelligent Systems (DSInS), Sydney, Australia, 20–22 November 2024. [Google Scholar]
- van Buchem, M.M.; Kant, I.M.J.; King, L.; Kazmaier, J.; Steyerberg, E.W.; Bauer, M.P. Impact of a Digital Scribe System on Clinical Documentation Time and Quality: Usability Study. JMIR AI 2024, 3, e60020. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Alsentzer, E.; Bergman, A.; McDermott, M.B.A.; Henry, P.; Ghassemi, M. Enhancing Clinical Documentation with Synthetic Data: Leveraging Generative Models for Improved Accuracy. arXiv 2024, arXiv:2406.06569. [Google Scholar]
- Radford, A.; Kim, J.W.; Xu, T.; Brockman, G.; McLeavey, C.; Sutskever, I. Robust Speech Recognition via Large-Scale Weak Supervision. arXiv 2023, arXiv:2212.04356. [Google Scholar]
- Adedeji, A.; Sanni, M.; Ayodele, E.; Joshi, S.; Olatunji, T. The Multicultural Medical Assistant: Can LLMs Improve Medical ASR Errors Across Borders? arXiv 2025, arXiv:2501.15310. [Google Scholar]
- Yu, S.; Lee, S.S.; Hwang, H. The ethics of using artificial intelligence in medical research. Kosin Med. J. 2024, 39, 229–237. [Google Scholar] [CrossRef] [Scilit]
- Ahuja, A.S.; Chen, E.P.; Hong, R.D.; Langlotz, C.P.; Lungren, M.P. Human-in-the-loop AI for clinical decision support: A review. J. Am. Med. Inform. Assoc. 2023, 30, 597–607. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gottlieb, M.; Lin, S.; Strong, P.; Lee, J.; Singer, D.R.; Marcus, A.; Singh, H. Assessing Patient-Reported Satisfaction with Care and Documentation Time in Primary Care Through AI-Driven Automatic Clinical Note Generation: Protocol for a Proof-of-Concept Study. JMIR Res. Protoc. 2025, 14, e66232. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Strong, P.; Lin, S.; Gottlieb, M.; Singh, H. Benchmarking and Datasets for Ambient Clinical Documentation: A Scoping Review of Existing Frameworks and Metrics for AI-Assisted Medical Note Generation. medRxiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Yim, W.w.; Fu, Y.; Ben Abacha, A.; Snider, N.; Lin, T.; Yetisgen, M. Aci-bench: A novel ambient clinical intelligence dataset for benchmarking automatic visit note generation. Sci. Data 2023, 10, 558. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ben Abacha, A.; Yim, W.w.; Michalopoulos, G.; Lin, T. An Investigation of Evaluation Methods in Automatic Medical Note Generation. In Proceedings of the Findings of the Association for Computational Linguistics; ACL: Toronto, ON, Canada, 2023; pp. 2575–2588. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Yang, R.; Alwakeel, M.; Kayastha, A.; Chowdhury, A.; Biro, J.M.; Sorrentino, A.D.; Handley, J.L.; Hantzmon, S.; Bessias, S.; et al. An evaluation framework for ambient digital scribing tools in clinical applications. npj Digit. Med. 2025, 8, 358. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sasseville, M.; Yousefi, F.; Ouellet, S.; Naye, F.; Stefan, T.; Carnovale, V.; Bergeron, F.; Ling, L.; Gheorghiu, B.; Hagens, S.; et al. The Impact of AI Scribes on Streamlining Clinical Documentation: A Systematic Review. Healthcare 2025, 13, 1447. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bracken, A.; Reilly, C.; Feeley, A.; Sheehan, E.; Merghani, K.; Feeley, I. Artificial Intelligence (AI)–Powered Documentation Systems in Healthcare: A Systematic Review. J. Med. Syst. 2025, 49, 28. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ng, J.J.W.; Wang, E.; Zhou, X.; Zhou, K.X.; Goh, C.X.L.; Sim, G.Z.N.; Tan, H.K.; Goh, S.S.N.; Ng, Q.X. Evaluating the Performance of Artificial Intelligence-Based Speech Recognition for Clinical Documentation: A Systematic Review. BMC Med. Inform. Decis. Mak. 2025, 25, 236. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shon, S.; Kim, H.K.; Kim, H.K.; Lee, H. Voice Activity Detection for Enhancing the Performance of Automatic Speech Recognition Systems. Appl. Sci. 2021, 11, 8146. [Google Scholar] [CrossRef] [Scilit]
- Ferrer, L.; Watanabe, S.; Gaur, Y.; Wiesner, M.; Delcroix, M.; Ogawa, A.; Nakatani, T.; Khudanpur, S. Studying the Effect of Silence on ASR Performance and How Its Usage Can Be Optimized in the ASR System. In Proceedings of the 23rd Annual Conference of the International Speech Communication Association (INTERSPEECH 2022), Incheon, Republic of Korea, 18–22 September 2022; pp. 1021–1025. [Google Scholar] [CrossRef] [Scilit]
- Lee, J.R.; Lungren, M.P.; Lee, P.H. Generating SOAP Notes from Doctor-Patient Conversations Using Large Language Models. arXiv 2023, arXiv:2303.06488. [Google Scholar]
- Agrawal, M.; Adams, M.; Brophy, E.; Chien, J.; Du, M.; Elibol, O.; Lin, C.C.; Logghe, P.; Maki, B.; Mehta, S.; et al. Medprompt: Cross-Modal Few-Shot Prompting for Ophthalmic Report Generation. In Proceedings of the Conference on Health, Inference, and Learning (CHIL), New York, NY, USA, 27–28 June 2024; pp. 1–25. [Google Scholar]
- Dash, S.; Shakyawar, S.K.; Sharma, M.; Kaushik, S. Big data in healthcare: Management, analysis and future prospects. J. Big Data 2019, 6, 54. [Google Scholar] [CrossRef] [Scilit]
- PHTI. Adoption of AI in Healthcare Delivery Systems: Early Applications & Impacts; Technical Report; Peterson Health Technology Institute: New York, NY, USA, 2025. [Google Scholar]
- Delaunay, J.; Girbes, D.; Cusido, J. Evaluating the Effectiveness of Large Language Models in Converting Clinical Data to FHIR Format. Appl. Sci. 2025, 15, 3379. [Google Scholar] [CrossRef] [Scilit]
- Nawab, K. Artificial intelligence scribe: A new era in medical documentation. Artif. Intell. Health 2024, 1, 12–15. [Google Scholar] [CrossRef] [Scilit]
- Measure and Improve Speech Accuracy. Google Cloud Speech-to-Text Documentation. Available online: https://cloud.google.com/speech-to-text/docs/speech-accuracy (accessed on 23 April 2025).
- Quiroz, J.C.; Laranjo, L.; Kocaballi, A.B.; Berkovsky, S.; Rezazadegan, D.; Coiera, E. Challenges of developing a digital scribe to reduce clinical documentation burden. npj Digit. Med. 2019, 2, 114. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Banerjee, S.; Agarwal, A.; Ghosh, P. High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR. arXiv 2024, arXiv:2412.00055. [Google Scholar]
- Zhou, L.; Blackley, S.V.; Kowalski, L.; Doan, R.; Acker, W.W.; Landman, A.B.; Kontrient, E.; Mack, D.; Meteer, M.; Bates, D.W.; et al. Analysis of Errors in Dictated Clinical Documents Assisted by Speech Recognition Software and Professional Transcriptionists. JAMA Netw. Open 2018, 1, e183458. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, W.; Hu, J.; Lv, F.; Tang, Z. A New Method for Long-Term Temperature Compensation of Structural Health Monitoring by Ultrasonic Guided Wave. Measurement 2025, 252, 117310. [Google Scholar] [CrossRef] [Scilit]
- Li, H.; Liu, W. Bone Damage Detection Using Ultrasonic Guided Waves: Multi-Feature Integration and Machine Learning Approaches. Nondestruct. Test. Eval. 2026, 1–34. [Google Scholar] [CrossRef] [Scilit]
- OpenAI. GPT-4 Technical Report. arXiv 2023, arXiv:2303.08774. [Google Scholar] [CrossRef] [Scilit]
- Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A.M.; Hauth, A.; Millican, K.; et al. Gemini: A Family of Highly Capable Multimodal Models. arXiv 2023, arXiv:2312.11805. [Google Scholar] [CrossRef] [Scilit]
- Saab, K.; Tu, T.; Weng, W.H.; Tanno, R.; Stutz, D.; Wulczyn, E.; Zhang, F.; Strother, T.; Park, C.; Vedadi, E.; et al. Capabilities of Gemini Models in Medicine. arXiv 2024, arXiv:2404.18416. [Google Scholar]
- Alsentzer, E.; Murphy, J.R.; Boag, W.; Weng, W.H.; Jindi, D.; Naumann, T.; McDermott, M. Publicly Available Clinical BERT Embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop, Minneapolis, MN, USA, 7 June 2019; pp. 72–78. [Google Scholar] [CrossRef] [Scilit]
- Adedeji, A.; Joshi, S.; Doohan, B. The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models. arXiv 2024, arXiv:2402.07658. [Google Scholar]
- Sharma, A.; Sharma, R.; Sharma, M. Transforming Health Care with Artificial Intelligence: Redefining Medical Documentation. Cureus 2025, 17, e56965. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yadav, S. Embracing artificial intelligence: Revolutionizing nursing documentation for a better future. Cureus 2024, 16, e57725. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sharma, A.; Sharma, R.; Sharma, M. Inspired Spine Smart Universal Resource Identifier (SURI): An Adaptive AI Framework for Transforming Multilingual Speech Into Structured Medical Reports. Cureus 2025, 17, e64606. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hodgson, T.; Magrabi, F.; Coiera, E. Automatic speech recognition performance for digital scribes: A performance comparison between general-purpose and specialized models tuned for patient-clinician conversations. JAMIA Open 2023, 6, ooad020. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Draper, T.C.; Leake, J.; Lamb-Riddell, K.; Cox, T.; McCormick, J.; Trowell, S.; Kiely, J.; Luxton, R. The Impact of Acoustic and Informational Noise on AI-Generated Clinical Summaries. medRxiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Federation of State Medical Boards. Navigating the Responsible and Ethical Incorporation of Artificial Intelligence into Clinical Practice; Policy Document; FSMB: Euless, TX, USA, 2024. [Google Scholar]
- Mianroodi, A.R.; Rezaie, A.; Todorov, N.G.; Rakovski, C.; Rudzicz, F. MedSynth: Realistic, Synthetic Medical Dialogue-Note Pairs. arXiv 2025, arXiv:2508.01401. [Google Scholar]
- Qiu, P.; Wu, C.; Zhang, X.; Lin, W.; Wang, H.; Zhang, Y.; Wang, Y.; Xie, W. Towards building multilingual language model for medicine. Nat. Commun. 2024, 15, 8384. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rotenstein, L.S.; Holmgren, A.J.; Thombley, R.; Sriram, A.; Dbouk, R.H.; Jost, M.; Aizenberg, D.; MacDonald, S.; Kanaparthy, N.; Williams, B.; et al. Changes in Clinician Time Expenditure and Visit Quantity with Adoption of Artificial Intelligence–Powered Scribes: A Multisite Study. JAMA 2026, 335, 1408–1417. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lee, K.C.K.; Clifton, D.A.; Chen, T.J.H. Generating synthetic medical data with large language models: Promises and pitfalls. Lancet Digit. Health 2023, 5, e753–e754. [Google Scholar]
- Chen, H.; Shu, K.; Wang, S.; Wang, S.; Chen, T.; Skiena, S. Generating Synthetic Text Data for Natural Language Processing. arXiv 2021, arXiv:2103.02762. [Google Scholar]
- Liu, A.; Ehrenberg, A.; Lo, A.; Denoix, C.; Barreau, C.; Lample, G.; Delignon, J.M.; Chandu, K.; Platen, P.; Muddireddy, P.; et al. Voxtral. arXiv 2025, arXiv:2507.13264. [Google Scholar] [CrossRef] [Scilit]
- OpenAI. GPT-4o System Card. 2024. Available online: https://openai.com/index/gpt-4o-system-card/ (accessed on 7 April 2026).
- Sellergren, A.; Kazemzadeh, S.; Jaroensri, T.; Kiraly, A.; Traverse, M.; Kohlberger, T.; Xu, S.; Jamil, F.; Hughes, C.; Lau, C.; et al. MedGemma Technical Report. arXiv 2025, arXiv:2507.05201. [Google Scholar]
- Zheng, L.; Chiang, W.L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.P.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv 2023, arXiv:2306.05685. [Google Scholar] [CrossRef] [Scilit]
- Panickssery, A.; Bowman, S.R.; Feng, S. LLM Evaluators Recognize and Favor Their Own Generations. arXiv 2024, arXiv:2404.13076. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.Y.; Wang, H.; Zhang, X.; Hu, E.; Lin, Y. Beyond the Surface: Measuring Self-Preference in LLM Judgments. arXiv 2025, arXiv:2506.02592. [Google Scholar]
- Warner, B.; Chaffin, A.; Clavié, B.; Weller, O.; Hallström, O.; Taghadouini, S.; Gallagher, A.; Biswas, R.; Ladhak, F.; Aarsen, T.; et al. Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Vienna, Austria, 27 July–1 August 2025; Volume 1, pp. 2526–2547. [Google Scholar] [CrossRef] [Scilit]
- Hada, R.; Gumma, V.; Wynter, A.; Diddee, H.; Ahmed, M.; Choudhury, M.; Bali, K.; Sitaram, S. Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation? In Proceedings of the Findings of the Association for Computational Linguistics: EACL 2024, St. Julian’s, Malta, 17–22 March 2024; pp. 1051–1070. [Google Scholar]
- Gupta, G.K.; Singh, A.; Manikandan, S.V.; Ehtesham, A. Digital Diagnostics: The Potential of Large Language Models in Recognizing Symptoms of Common Illnesses. arXiv 2024, arXiv:2405.06712. [Google Scholar]
- Delaunay, J.; Cusido, J. Evaluating the Performance of Large Language Models in Predicting Diagnostics for Spanish Clinical Cases in Cardiology. Appl. Sci. 2025, 15, 61. [Google Scholar] [CrossRef] [Scilit]
- Rosner, B.; Glynn, R.J.; Lee, M.L.T. The Wilcoxon signed rank test for paired comparisons of clustered data. Biometrics 2006, 62, 185–192. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Al Meslamani, A.Z. Beyond implementation: The long-term economic impact of AI in healthcare. J. Med. Econ. 2023, 26, 1566–1569. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lin, S.; Strong, P.; Singh, H.; Gottlieb, M. Clinicians’ Experiences with EHR Documentation and Attitudes Toward AI-Assisted Documentation; Stanford Medicine White Paper; Stanford Medicine: Stanford, CA, USA; Google Health: Mountain View, CA, USA, 2024. [Google Scholar]
- Li, Y.; Wang, H.; Yerebakan, H.Z.; Shinagawa, Y.; Luo, Y. FHIR-GPT Enhances Health Interoperability with Large Language Models. NEJM AI 2024, 1, AIcs2300301. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schmiedmayer, P.; Rao, A.; Zagar, P.; Ravi, V.; Zahedivash, A.; Fereydooni, A.; Aalami, O. LLM on FHIR—Demystifying Health Records. arXiv 2024, arXiv:2402.01711. [Google Scholar]
- Hong, N.; Wen, A.; Shen, F.; Sohn, S.; Liu, S.; Liu, H.; Jiang, G. Integrating Structured and Unstructured EHR Data Using an FHIR-based Type System: A Case Study with Medication Data. AMIA Jt. Summits Transl. Sci. Proc. 2018, 2018, 74–83. [Google Scholar]
- Fagherazzi, G.; Fischer, A.; Ismael, M.; Despotovic, V. Voice for Health: The Use of Vocal Biomarkers from Research to Clinical Practice. Digit. Biomark. 2021, 5, 78–88. [Google Scholar] [CrossRef] [Scilit]
- De Silva, U.; Subramaniam, P.; Doan, T.N.; Babar, Z.U.D.; Nandasena, M.; Li, C.; Masek, M.; Lee, A. Clinical Decision Support Using Speech Signal Analysis: Systematic Scoping Review of Neurological Disorders. J. Med. Internet Res. 2025, 27, e63004. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vilenchik, D.; Cwikel, J.; Ezra, Y.; Hausdorff, T.; Lazarov, M.; Sergienko, R.; Abramovitz, R.; Schmidt, I.; Perez, A.S. Method Matters: Enhancing Voice-Based Depression Detection with a New Data Collection Framework. Depress. Anxiety 2025, 2025, 4839334. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Montenegro, L.; Gomes, L.M.; Machado, J.M. What We Know About the Role of Large Language Models for Medical Synthetic Dataset Generation. AI 2025, 6, 109. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Yerebakan, M.; Luo, Y.; Amaba, B.; Swope, W.; Hu, B. The Effect of Different Occupational Background Noises on Voice Recognition Accuracy. J. Comput. Inf. Sci. Eng. 2022, 22, 050905. [Google Scholar] [CrossRef] [Scilit]
- Chiu, E.K.Y.; Chung, T.W.H. Protocol for Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations. medRxiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Saadi, N.; Raha, T.; Christophe, C.; Pimentel, M.A.; Rajan, R.; Kanithi, P.K. Bridging Language Barriers in Healthcare: A Study on Arabic LLMs. arXiv 2025, arXiv:2501.09825. [Google Scholar]
- Wang, X.; Zhang, N.; He, H.; Nguyen, T.; Yu, K.H.; Deng, H.; Brandt, C.; Bitterman, D.; Pan, L.; Cheng, C.y.; et al. Safety challenges of AI in medicine in the era of large language models. arXiv 2024, arXiv:2409.18968. [Google Scholar] [CrossRef] [Scilit]









| Prior Work | What It Contributes | How Our Study Is Positioned |
|---|---|---|
| ACI-BENCH [15] | Public benchmark corpus for automatic visit-note generation from clinical dialogue. | We are not the first benchmark. We add modular ASR–LLM comparison, English–Spanish synthetic audio, and deployment-grounded usage traces. |
| Ben Abacha et al. [16] | Analysis of evaluation methods and metrics for automatic medical note generation. | We do not claim that automated metrics solve evaluation. We use them as technical indicators within one applied benchmark. |
| SCRIBE framework [17] | Broad ambient digital scribing evaluation framework combining simulation, computational metrics, human review, and LLM-based evaluation. | We do not present a new general framework. We provide an applied integrated evaluation of one modular documentation pipeline. |
| Digital-scribe and ambient-AI studies [3,7,48] | Measure documentation time, note quality, editing, or clinician-reported outcomes. | Our telemetry does not measure these outcomes. It only describes finalized-export usage patterns. |
| Systematic reviews [18,19,20] | Show that the field is promising but heterogeneous, with variable methods and outcomes. | Our study contributes one structured evaluation, but it cannot close the evidence gap alone. |
| Metric | Mean | Std. Dev. |
|---|---|---|
| Audio duration (s) | 226.28 | 38.13 |
| Words per discussion | 573.88 | 99.76 |
| Turns per discussion | 22.62 | 3.50 |
| Doctor words | 421.56 | 69.72 |
| Patient words | 152.31 | 38.87 |
| Doctor turns | 14.38 | 2.32 |
| Patient turns | 8.25 | 1.60 |
| Proportion of doctor words (%) | 73.46 | – |
| Proportion of patient words (%) | 26.54 | – |
| Proportion of doctor turns (%) | 63.54 | – |
| Proportion of patient turns (%) | 36.46 | – |
| LLM Model | |||||
|---|---|---|---|---|---|
|
Gemini 1.5 Pro | GPT-4o | MedGemma |
Cross-Specialty Mismatch Control | ||
| ASR Model | GPT-4o Transcribe | 20 | 20 | 20 | 0 |
| Voxtral | 20 | 20 | 20 | 0 | |
| Whisper | 20 | 20 | 20 | 0 | |
| GCP | 10 | 10 | 10 | 0 | |
| Cross-specialty | 0 | 0 | 0 | 20 | |
| mismatch control | |||||
| LLM Model | Hallucination Score | Terminology Accuracy |
|---|---|---|
| GPT-4o | 0.980 ± 0.076 | 0.994 ± 0.024 |
| Gemini 1.5 Pro | 0.972 ± 0.069 | 0.993 ± 0.022 |
| MedGemma | 0.968 ± 0.068 | 0.981 ± 0.051 |
| Cross-specialty mismatch control | 0.010 ± 0.030 | 0.276 ± 0.360 |
| Model | Role | Rate (USD) |
|---|---|---|
| Google Cloud STT | ASR | 0.078/min |
| Whisper | ASR | 0.006/min |
| GPT-4o Transcribe | ASR | 0.006/min |
| Voxtral Mini | ASR | 0.004/min |
| Model | Usage | Input Price ($) | Output Price ($) |
|---|---|---|---|
| GPT-4o | Post-process | 2.5 | 10.0 |
| Gemini 2.0 Flash | Post-process | 0.10 | 0.40 |
| Gemini 1.5 Pro | Evaluation | 0.075 | 0.30 |
| Region | Active Users (≥1 Report) | Reports Generated | Total Minutes Processed | Mean Minutes Per Report |
|---|---|---|---|---|
| LATAM | 205 | 2314 | 23,190 | 10.02 |
| Europe | 96 | 793 | 6345 | 8.00 |
| Total | 301 | 3107 | 29,535 | — |
| Specialty | Active Users | Reports Generated | Total Minutes Processed |
|---|---|---|---|
| Psychology | 14 | 120 | 3345 |
| Traumatology & Orthopedics | 15 | 161 | 2458 |
| Psychiatry | 7 | 70 | 1693 |
| Cardiology (Electrophysiology) * | 1 | 311 | 2482 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Delaunay, J.; Sarkis, M.; Solé-Casals, J.; Cusido, J. A Modular Evaluation of AI-Assisted Clinical Documentation. Appl. Sci. 2026, 16, 6961. https://doi.org/10.3390/app16146961
Delaunay J, Sarkis M, Solé-Casals J, Cusido J. A Modular Evaluation of AI-Assisted Clinical Documentation. Applied Sciences. 2026; 16(14):6961. https://doi.org/10.3390/app16146961
Chicago/Turabian StyleDelaunay, Julien, Maissaa Sarkis, Jordi Solé-Casals, and Jordi Cusido. 2026. "A Modular Evaluation of AI-Assisted Clinical Documentation" Applied Sciences 16, no. 14: 6961. https://doi.org/10.3390/app16146961
APA StyleDelaunay, J., Sarkis, M., Solé-Casals, J., & Cusido, J. (2026). A Modular Evaluation of AI-Assisted Clinical Documentation. Applied Sciences, 16(14), 6961. https://doi.org/10.3390/app16146961

