Automated Structuring and Analysis of Unstructured Equipment Maintenance Text Data in Manufacturing Using Generative AI Models: A Comparative Study of Pre-Trained Language Models
Abstract
1. Introduction
1.1. Research Background and Objectives
1.2. Research Scope and Methodology
2. Literature Survey
2.1. Overview of Natural Language Processing
2.2. Text Summarization
2.3. Transformer Models
2.4. BART Model
2.5. T5 Model
2.6. Qwen Model
3. System Framework
3.1. Characteristics and Analysis of Target Dataset
3.2. Pipeline Configuration
3.2.1. Stage 1: Training Dataset Construction and Expert Validation
3.2.2. Stage 2: Generative AI Model Development
3.2.3. Stage 3: Model Performance Evaluation
3.2.4. Stage 4: Primary Structured Database Construction
3.2.5. Stage 5: Secondary Structured DB Construction and Standardization
3.2.6. Stage 6: Web-Based Service Implementation
4. Dataset Construction and Model Training
4.1. Training Dataset Construction
4.2. Model Development
4.2.1. BART
4.2.2. T5
4.2.3. Qwen
5. Model Performance Evaluation and Results Analysis
5.1. Evaluation Process
5.2. Evaluation Results Analysis
5.3. Performance Improvement Directions
5.4. Qualitative Error Analysis and Discussion
6. Data Structuring and Analysis Service Implementation
6.1. Structured Database Construction
- For failed components and failure types, which show relatively structured terminology patterns, a regular expression-based rule application approach was utilized. Through this method, terms with identical concepts described in various expression formats were identified and unified into standardized forms, ensuring analytical consistency.
- In contrast, corrective actions are described in various ways using natural language and have high expression variability, making structuring through regular expressions alone limited. Accordingly, the application of topic-level grouping techniques is proposed to group sentences with similar meanings. This approach performs automatic grouping based on semantic similarity between texts and can be utilized as an effective methodology for supporting the structuring of corrective actions, rather than strict algorithmic clustering. The specific application and validation of topic-level grouping techniques were set as future research tasks.
6.2. AI Model Service Implementation
- Time Series Analysis Module supports analysis of changes in failure occurrence pattern trends over time.
- Correlation Analysis Module enables identification of causal relationships by analyzing statistical associations between failed components and failure types.
- Frequency Analysis Module enables identification of priorities based on the occurrence frequency of specific failure types or corrective actions.
- Anomaly Detection Module is utilized to identify abnormal signs or exceptional patterns that deviate from normal ranges.
7. Conclusions and Future Research
7.1. Research Results
7.2. Limitations and Improvement Directions
7.3. Future Research Directions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| BART | Bidirectional and Auto-Regressive Transformers |
| MES | Manufacturing Execution Systems |
| GPT | Generative Pre-trained Transformer |
| ROUGE | Recall-Oriented Understudy for Gisting Evaluation |
| DX | Digital Transformation |
| AX | AI Transformation |
| PLC | Programmable Logic Controller |
| SMEs | Small and Medium-sized manufacturing Enterprises |
| JIT | Just-In-Time |
| ERP | Enterprise Resource Planning |
| NLP | Natural Language Processing |
| LLMs | Large Language Models |
| ATS | Automatic Text Summarization |
| AI | Artificial Intelligence |
Appendix A. Exact Input/Output Schematics for BART and T5 Models
| Item | BART (KoBART) | T5 (pko-t5-Base) |
|---|---|---|
| Model Architecture | Encoder–Decoder Transformer (BART) | Encoder–Decoder Transformer (T5, text-to-text) |
| Task Formulation | Sequence-to-sequence generation for slot filling | Text-to-text sequence-to-sequence generation for slot filling |
| Raw Input Source | Korean equipment maintenance log text | Korean equipment maintenance log text |
| Model Input String | <maintenance log text> | categorize: <maintenance log text> |
| Tokenizer | PreTrainedTokenizerFast (gogamza/kobart-base-v1) | T5TokenizerFast (paust/pko-t5-base) |
| Training Target (Label) | Structured annotation string (abstractive column) | Structured annotation string (abstractive column) |
| Raw Generated Output | Korean natural-language key-value sentence separated by commas and using case particles | Same as BART |
| Output Schema (Semantic Fields) | Failed components, failure types, corrective actions | Failed components, failure types, corrective actions |
| Output Format Assumption | 고장부품은_, 불량유형은_, 조치내용은_ | 고장부품은_, 불량유형은_, 조치내용은_ |
| Post-processing | Remove sentence-ending particles → split by comma → split by case particles → map to fields | Identical post-processing pipeline |
| Parsing Failure Handling | Raw generated text retained | Raw generated text retained |
| Decoding Strategy | Beam search (num_beams = 5, max_length = 512) | Beam search (num_beams = 5, max_length = 512) |
Appendix A.1. Note on Language-Specific Output Representation
Appendix B. Prompt Templates for LLM-Based Data Preparation and Extraction
Appendix B.1. GPT-4 Prompt for Training Data Preparation
- Input Example
- Output Example
- failed_component
- failure_type
- corrective_action
Appendix B.2. Qwen Prompt for Model Training and Inference
- “failed_component”: <extracted component>
- “failure_type”: <extracted failure type>
- “corrective_action”: <extracted corrective action>
Appendix B.3. Prompt Consistency and Reproducibility
References
- Garcia, C.I.; DiBattista, M.A.; Letelier, T.A.; Halloran, H.D.; Camelio, J.A. Framework for LLM Applications in Manufacturing. Manuf. Lett. 2024, 41, 253–263. [Google Scholar] [CrossRef] [Scilit]
- Werheid, J.; Melnychuk, O.; Zhou, H.; Huber, M.; Rippe, C.; Joosten, D.; Schmitt, R.H. Designing an LLM-Based Copilot for Manufacturing Equipment Selection. Manuf. Lett. 2025, 46, 123–127. [Google Scholar] [CrossRef] [Scilit]
- Shidaganti, G.; Shetty, R.; Edara, T.; Srinivas, P.; Tammineni, S.C. Exploratory analysis on the natural language processing models for task specific purposes. Bull. Electr. Eng. Inform. 2024, 13, 1245–1255. [Google Scholar] [CrossRef] [Scilit]
- Gupta, A.; Chugh, D.; Anjum; Katarya, R. Automated News Summarization Using Transformers. In Sustainable Advanced Computing; Lecture Notes in Electrical Engineering; Springer: Berlin/Heidelberg, Germany, 2022; pp. 249–259. [Google Scholar]
- Saxena, P.; El-Haj, M. Exploring Abstractive Text Summarisation for Podcasts: A Comparative Study of BART and T5 Models. In Proceedings of the Conference Recent Advances in Natural Language Processing—Large Language Models for Natural Language Processings; INCOMA Ltd.: Shumen, Bulgaria, 2023; pp. 1023–1033. [Google Scholar]
- Venkataramana, A.; Srividya, K.; Cristin, R. Abstractive Text Summarization Using BART. In Proceedings of the 2022 IEEE 2nd Mysore Sub Section International Conference (MysuruCon), Mysuru, India, 16–17 October 2022; pp. 1–6. [Google Scholar]
- Maghfiroh, N.A.; Abdurrachman Bachtiar, F.; Muflikhah, L. Comparative Analysis of Summarization Methods for Skin Care Product Reviews: A Study on BERT, BART, and T5 Models. In Proceedings of the 2023 International Conference on Advanced Mechatronics, Intelligent Manufacture and Industrial Automation (ICAMIMIA), Lombok, Indonesia, 14–15 November 2023; pp. 593–598. [Google Scholar]
- Rehman, T.; Das, S.; Sanyal, D.K.; Chattopadhyay, S. An analysis of abstractive text summarization using pre-trained models. In Proceedings of International Conference on Computational Intelligence, Data Science and Cloud Computing: IEM-ICDC 2021; Springer: Berlin/Heidelberg, Germany, 2022; pp. 253–264. [Google Scholar]
- Deokar, V.; Shah, K. Automated text summarization of news articles. Int. Res. J. Eng. Technol. 2021, 8, 1908–1911. [Google Scholar]
- Glass, J.R.; Hazen, T.J.; Cyphers, D.S.; Malioutov, I.; Huynh, D.; Barzilay, R. Recent progress in the MIT spoken lecture processing project. In Proceedings of the Interspeech, Antwerp, Belgium, 27–31 August 2007; pp. 2553–2556. [Google Scholar]
- Boorugu, R.; Ramesh, G. A survey on NLP based text summarization for summarizing product reviews. In Proceedings of the 2020 Second International Conference on Inventive Research in Computing Applications (ICIRCA), Coimbatore, India, 15–17 July 2020; pp. 352–356. [Google Scholar]
- Awasthi, I.; Gupta, K.; Bhogal, P.S.; Anand, S.S.; Soni, P.K. Natural language processing (NLP) based text summarization-a survey. In Proceedings of the 2021 6th International Conference on Inventive Computation Technologies (ICICT), Coimbatore, India, 20–22 January 2021; pp. 1310–1317. [Google Scholar]
- Luhn, H.P. The automatic creation of literature abstracts. IBM J. Res. Dev. 1958, 2, 159–165. [Google Scholar] [CrossRef] [Scilit]
- Radev, D.; Hovy, E.; McKeown, K. Introduction to the special issue on summarization. Comput. Linguist. 2002, 28, 399–408. [Google Scholar] [CrossRef] [Scilit]
- Christian, H.; Agus, M.P.; Suhartono, D. Single document automatic text summarization using term frequency-inverse document frequency (TF-IDF). ComTech Comput. Math. Eng. Appl. 2016, 7, 285–294. [Google Scholar] [CrossRef] [Scilit]
- Nomoto, T. Bayesian learning in text summarization. In Proceedings of Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Kerrville, TX, USA, 2005; pp. 249–256. [Google Scholar]
- Jing, H.; McKeown, K. Cut and paste based text summarization. In Proceedings of the 1st Meeting of the North American Chapter of the Association for Computational Linguistics, Seattle, WA, USA, 29 April–4 May 2000. [Google Scholar]
- Knight, K.; Marcu, D. Summarization beyond sentence extraction: A probabilistic approach to sentence compression. Artif. Intell. 2002, 139, 91–107. [Google Scholar] [CrossRef] [Scilit]
- Mishra, R.; Bian, J.; Fiszman, M.; Weir, C.B.; Jonnalagadda, S.; Mostafa, J.; Del Fiol, G. Text Summarization in the Biomedical Domain: A Systematic Review of Current Techniques. J. Biomed. Inform. 2014, 52, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Genest, P.-E.; Lapalme, G. Fully abstractive approach to guided summarization. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers); Association for Computational Linguistics: Kerrville, TX, USA, 2012; pp. 354–358. [Google Scholar]
- Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit]
- Shi, T.; Keneshloo, Y.; Ramakrishnan, N.; Reddy, C.K. Neural abstractive text summarization with sequence-to-sequence models. ACM Trans. Data Sci. 2021, 2, 1–37. [Google Scholar] [CrossRef] [Scilit]
- Nallapati, R.; Xiang, B.; Zhou, B. Sequence-to-Sequence Rnns for Text Summarization; IBM Watson: Yorktown Heights, NY, USA, 2016. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6010. [Google Scholar]
- Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers); Association for Computational Linguistics: Kerrville, TX, USA, 2019; pp. 4171–4186. [Google Scholar]
- Zhang, J.; Zhao, Y.; Saleh, M.; Liu, P. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the International Conference on Machine Learning, virtually, 13–18 July 2020; pp. 11328–11339. [Google Scholar]
- Dong, L.; Yang, N.; Wang, W.; Wei, F.; Liu, X.; Wang, Y.; Gao, J.; Zhou, M.; Hon, H.-W. Unified language model pre-training for natural language understanding and generation. Adv. Neural Inf. Process. Syst. 2019, 32, 13063–13075. [Google Scholar]
- Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I. Improving Language Understanding by Generative Pre-Training; OpenAI: San Francisco, CA, USA, 2018. [Google Scholar]
- Rao, R.; Sharma, S.; Malik, N. Automatic text summarization using transformer-based language models. Int. J. Syst. Assur. Eng. Manag. 2024, 15, 2599–2605. [Google Scholar] [CrossRef] [Scilit]
- Shaik Vadla, M.K.; Suresh, M.A.; Viswanathan, V.K. Enhancing Product Design through AI-Driven Sentiment Analysis of Amazon Reviews Using BERT. Algorithms 2024, 17, 59. [Google Scholar] [CrossRef] [Scilit]
- Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; Zettlemoyer, L. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv 2019, arXiv:1910.13461. [Google Scholar]
- Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 2020, 33, 1877–1901. [Google Scholar]
- Xue, L.; Constant, N.; Roberts, A.; Kale, M.; Al-Rfou, R.; Siddhant, A.; Barua, A.; Raffel, C. mT5: A massively multilingual pre-trained text-to-text transformer. arXiv 2020, arXiv:2010.11934. [Google Scholar]
- Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P.J. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 2020, 21, 1–67. [Google Scholar]
- Zhuang, F.; Qi, Z.; Duan, K.; Xi, D.; Zhu, Y.; Zhu, H.; Xiong, H.; He, Q. A comprehensive survey on transfer learning. Proc. IEEE 2020, 109, 43–76. [Google Scholar] [CrossRef] [Scilit]
- Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F. Qwen technical report. arXiv 2023, arXiv:2309.16609. [Google Scholar] [CrossRef] [Scilit]
- Team, Q. Qwen2 technical report. arXiv 2024, arXiv:2407.10671. [Google Scholar]
- Lin, C.-Y. Rouge: A package for automatic evaluation of summaries. In Proceedings of the Text Summarization Branches Out; Association for Computational Linguistics: Kerrville, TX, USA, 2004; pp. 74–81. [Google Scholar]










| Papers | Dataset | Applied Models | High-Performance Model |
|---|---|---|---|
| Automated Text Summarization of News Articles [9] | Internet news articles | BART, T5 | BART |
| Abstractive Text Summarization Using BART [6] | Internet news articles | BERT, Roberta, T5, BART | BART |
| An Analysis of Abstractive Text Summarization Using Pre-trained Models [8] | Internet news articles | PEGASU, T5, BART | BART |
| Automated News Summarization Using Transformers [4] | Internet news articles | BART, T5, PEGASUS | T5 |
| Comparative Analysis of Summarization Methods for Skin Care Product Reviews: A Study on BERT, BART, and T5 Models [7] | Product review data | BERT, BART, T5 | BART |
| Automatic text summarization using transformer-based language models [29] | Internet news articles | BART, T5 | BART |
| Exploring Abstractive Text Summarization for Podcasts: A Comparative Study of BART and T5 Models [5] | Podcasts | BART, T5 | BART |
| Enhancing Product Design through AI-Driven Sentiment Analysis of Amazon Reviews Using BERT [30] | Product review data | BERT, T5 | BERT |
| Exploratory analysis on the natural language processing models for task specific purposes [3] | Internet news articles | BERT, BART, T5 | BART |
| Model | Applied Models | Model Size |
|---|---|---|
| BART | KoBART (SK Telecom) | 110 million |
| T5 | pko-t5-base (PAUST) | 250 million |
| Qwen2.5 | Qwen2.5-0.5B-Instruct (Alibaba Cloud) | 500 million |
| Parameters | Setting Values |
|---|---|
| batch_size | 256 |
| max_len | 32 |
| num_workers | 4 |
| lr | 3 × 10−5 |
| max_epochs | 40 |
| warmup_ratio | 0.1 |
| Parameters | Setting Values |
|---|---|
| batch_size | 64 |
| max_len | 32 |
| num_workers | 4 |
| lr | 3 × 10−5 |
| max_epochs | 30 |
| warmup_ratio | 0.1 |
| Parameters | Setting Values |
|---|---|
| num_train_epochs | 3 |
| per_device_train_batch_size | 2 |
| gradient_accumulation_steps | 2 |
| learning_rate | 1 × 10−4 |
| warmup_ratio | 0.03 |
| optimizer | adamw_torch_fused |
| Items | Models | Exact Match (EM) | Precision | Recall | F1-Score |
|---|---|---|---|---|---|
| Failed components | BART | 0.473 | 0.794 | 0.802 | 0.798 |
| T5 | 0.564 | 0.832 | 0.838 | 0.835 | |
| Qwen | 0.635 | 0.860 | 0.867 | 0.863 | |
| Failure types | BART | 0.395 | 0.521 | 0.519 | 0.520 |
| T5 | 0.468 | 0.670 | 0.545 | 0.601 | |
| Qwen | 0.486 | 0.629 | 0.644 | 0.636 | |
| Corrective actions | BART | 0.387 | 0.729 | 0.768 | 0.748 |
| T5 | 0.415 | 0.826 | 0.644 | 0.724 | |
| Qwen | 0.513 | 0.805 | 0.793 | 0.799 |
| Items | Models | Rouge Score | ||
|---|---|---|---|---|
| ROUGE-1 | ROUGE-2 | ROUGE-L | ||
| failed components | BART | 0.790 | 0.711 | 0.790 |
| T5 | 0.833 | 0.761 | 0.833 | |
| Qwen | 0.855 | 0.790 | 0.854 | |
| failure types | BART | 0.571 | 0.260 | 0.571 |
| T5 | 0.641 | 0.285 | 0.641 | |
| Qwen | 0.670 | 0.327 | 0.670 | |
| corrective actions | BART | 0.773 | 0.607 | 0.766 |
| T5 | 0.756 | 0.599 | 0.751 | |
| Qwen | 0.817 | 0.666 | 0.809 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Cho, Y. Automated Structuring and Analysis of Unstructured Equipment Maintenance Text Data in Manufacturing Using Generative AI Models: A Comparative Study of Pre-Trained Language Models. Appl. Sci. 2026, 16, 1969. https://doi.org/10.3390/app16041969
Cho Y. Automated Structuring and Analysis of Unstructured Equipment Maintenance Text Data in Manufacturing Using Generative AI Models: A Comparative Study of Pre-Trained Language Models. Applied Sciences. 2026; 16(4):1969. https://doi.org/10.3390/app16041969
Chicago/Turabian StyleCho, Yongju. 2026. "Automated Structuring and Analysis of Unstructured Equipment Maintenance Text Data in Manufacturing Using Generative AI Models: A Comparative Study of Pre-Trained Language Models" Applied Sciences 16, no. 4: 1969. https://doi.org/10.3390/app16041969
APA StyleCho, Y. (2026). Automated Structuring and Analysis of Unstructured Equipment Maintenance Text Data in Manufacturing Using Generative AI Models: A Comparative Study of Pre-Trained Language Models. Applied Sciences, 16(4), 1969. https://doi.org/10.3390/app16041969
