AI-Powered Analytics in Technical Communication: Advancing Information Management Toward Content Data Science †
Abstract
1. Introduction
2. Conceptual Framework and Methodology
2.1. Information Architectures in Technical Communication
2.2. AI-Driven Content Delivery with RAG: From Chunks to Topics
2.3. Analysis Methods and Metrics
2.4. Similarity Metrics for AI-Readiness
2.5. Distribution Functions and Analytic Compression of Ensemble Properties
3. Academic Research and Content Data Science Projects
Academic Setting and the PIAI!-Lab
- Problem Definition: Business understanding of content-related processes and research questions was introduced during Bachelor studies, providing the domain foundation for the project.
- Data Acquisition: Data collection exports were technically and conceptually prepared by the course instructor in advance and extended at the start of the project in collaboration with the industry partner.
- Data Cleaning and Preparation: Data cleaning of XML export formats from content management systems was performed to enable selective access to topic-level content. Preparatory analyses, such as module sizes and variant frequencies, provide structural characterization prior to vectorization and ensemble computation.
- Data Integration and Transformation: Data transformation through normalization and semantic preparation of topics and ensembles for similarity analysis forms the core phase of the project.
- Exploratory Data Analysis: Analysis of ensemble properties includes visualization and preparation of data for subsequent investigations.
- Modeling: Data modeling based on standard ensemble metrics is extended by students with independent research contributions, developing additional models, applying project-specific perspectives, and deriving conclusions for the ensemble under investigation.
- Deployment: Operationalization is not part of the academic project scope.
- Communication: Documentation and knowledge transfer using a Python-based Jupyter environment (“PIAI!-Lab”) and presentation of results to academic and industry partners.
4. Results and Discussion
4.1. Ensemble Properties for AI-Readiness Measurements
4.2. Projects: Validation of Ensemble Properties
- Reuse behavior measurements analogous to the earlier introduced REx analytics [15,16] can be used to characterize ensembles and content management environments. Higher reuse rates are generally associated with shorter topics; large topics exhibit low reuse, consistent with variant-specific content creation in manufacturing environments.
- Product-level sub-ensembles show higher median MSD than the global ensemble, confirming that metadata-based pre-filtering improves retrieval sharpness.
- RAG pre-tests demonstrate that for low-MSD sub-ensembles, the correct topic fails to appear in the top 5 results in the global ensemble, while sub-ensemble filtering restores correct retrieval—confirming that ensemble sizing and extrinsic pre-filtering are strongly recommended.
- At small topic sizes, embedding truncation is unlikely to affect most topics in well-modularized TC corpora, supporting the use of all-mpnet-base-v2 [13] for this content type.
- Lower MSD thresholds detected in the data are ensemble-relative: the observed range must be interpreted within the specific ensemble context, as domain-specific semantic overlap imposes inherent bounds.
- Topics containing tables systematically exhibit low MSD, likely due to normalization reducing tabular content to similar text patterns—identified as a dedicated preprocessing challenge.
- An LLM-based metadata classification process with human-in-the-loop validation is usable for legacy content reclassification, enabling retrospective coverage of PI-Class or other classification schemes.
- PI-Class-based new information architecture achieves higher median MSD than legacy architectures, confirming that modular, classification-driven content structuring improves AI-readiness.
- Shorter topics have the greatest potential for high distinctiveness, but short topics containing tables or reference content can also show low MSD.
- Embedding model comparison shows that all-mpnet-base-v2 (384 tokens) [13] produces larger similarity amplitudes than higher-token models, making it more suitable for ensemble analysis despite its smaller token limit.
- Truncation tests confirm that MSD distributions remain structurally stable across token lengths below 384. In order to overcome token dependency and limitations, short topics yield the most reliable vectorization results.
- RAG pre-tests validate the metric: topics with high MSD yield steep retrieval score curves, while low MSD topics produce flat distributions with ambiguous context selection.
- Decreasing ensemble sizing reduces the similarity width, demonstrating that metadata-based retrieval space reduction is highly effective.
- Three semantic levels were used: Level 1 (content + title), Level 2 (XML structural tags), Level 3 (intrinsic metadata scoring function)—only Level 3 produces a significant improvement in MSD and SimWidth.
- Semantic Level 3 applies a simple and discrete metadata scoring function combined with cosine similarity, effectively pushing mismatching topics toward zero similarity while maintaining high scores for matching topics.
4.3. Summary
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| CCMS | Component Content Management System |
| LLM | Large Language Model |
| MSD | Mean Similarity Difference |
| MSSD | Mean Square Similarity Difference |
| PI | Product—Information |
| Portable Document Format | |
| RAG | Retrieval-Augmented Generation |
| TC | Technical Communication |
| XML | Extensible Markup Language |
References
- Kees, V.M. Improving the Quality of AI-Driven Technical Content Delivery. Available online: https://www.tcworld.info/e-magazine/intelligent-information/improving-the-quality-of-ai-driven-technical-content-delivery (accessed on 30 April 2026).
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Vlacic, V.; Karpukhin, V.; Oğuz, B.; Röder, M.; Alon, U.; Levy, O.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems; Curran Associates Inc.: New York, NY, USA, 2020; Volume 33, pp. 9459–9474. [Google Scholar]
- Behabetz, J.; Gatzke, N.; Ziegler, W. Content Reuse Analytics and AI Readiness of Content in the Domain for Maritime Product Component Supplier. In Proceedings of the International Conference on Educational Technology, Language and Technical Communication (ETLTC), Aizu-Wakamatsu, Japan, 20–25 January 2026. [Google Scholar]
- Groß, M.; Heck, N.; Ziegler, W. Semantic Information Architectures and Topic Ensemble Properties in AI Delivery of Product Information in the Domain of Water Treatment. In Proceedings of the International Conference on Educational Technology, Language and Technical Communication (ETLTC), Aizu-Wakamatsu, Japan, 20–25 January 2026. [Google Scholar]
- Muschinski, J.; Reiling, J.; Westenhoff, R.; Ziegler, W. A Comparison of Information Architectures of Software Documentation in RAG-based Delivery Scenarios. In Proceedings of the International Conference on Educational Technology, Language and Technical Communication (ETLTC), Aizu-Wakamatsu, Japan, 20–25 January 2026. [Google Scholar]
- Nguyen, G.L.; Schardt, E.A.; Ziegler, W. Content and AI-Delivery Analytics of Semantically Enriched Content in Engine Manufacturing. In Proceedings of the International Conference on Educational Technology, Language and Technical Communication (ETLTC), Aizu-Wakamatsu, Japan, 20–25 January 2026. [Google Scholar]
- Ziegler, W. Drivers of Digital Information Services: Intelligent Information Architectures in Technical Communication. In Proceedings of the ACM Chapter Conference on Educational Technology, Language and Technical Communication, Aizu-Wakamatsu, Japan, 28 January–1 February 2019; pp. 48–52. [Google Scholar]
- tekom Europe e.V. iiRDS—Intelligent Information Request and Delivery Standard, Version 1.2. Available online: https://iirds.org/fileadmin/iiRDS_specification/20231110-1.2-release/index.html (accessed on 30 April 2026).
- Es, S.; James, J.; Espinosa Anke, L.; Schockaert, S. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, St. Julians, Malta, 17–22 March 2024; pp. 150–158. [Google Scholar]
- Chen, H.; Chen, G.; Blasch, E.P.; Douville, P.; Pham, K. Information Theoretic Measures for Performance Evaluation and Comparison. In Proceedings of the 12th International Conference on Information Fusion, Seattle, WA, USA, 6–9 July 2009; pp. 874–881. [Google Scholar]
- Lai, S.; Cheung, T.-H.; Fung, K.-C.; Xue, K.; Lin, K.-H.; Choi, Y.-M.; Ng, V.; Lam, K.-M. Enhancing Technical Documents Retrieval for RAG. In Proceedings of the 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Singapore, 22–24 October 2025. [Google Scholar]
- Bhardwaj, P. Your Chunks Failed Your RAG in Production. Towards Data Science. Available online: https://towardsdatascience.com/your-chunks-failed-your-rag-in-production/ (accessed on 19 April 2026).
- Reimers, N.; Gurevych, I. Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks. Available online: https://huggingface.co/sentence-transformers/all-mpnet-base-v2 (accessed on 30 April 2026).
- Belcic, I.; Stryker, C. What Is Agentic RAG? Available online: https://www.ibm.com/think/topics/agentic-rag (accessed on 22 February 2026).
- Ziegler, W. Metrische Untersuchung der Wiederverwendung im Content Management; Karlsruhe University of Applied Sciences: Karlsruhe, Germany, 2008; Available online: https://www.i4icm.de/wp-content/uploads/2025/03/CMS-Metrik_Ziegler.pdf (accessed on 1 July 2026). (In German)
- Oberle, C.; Ziegler, W. Content Intelligence for Content Management Systems. Available online: https://www.tcworld.info/e-magazine/technical-writing/content-intelligence-for-content-management-systems-355 (accessed on 30 April 2026).









Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ziegler, W. AI-Powered Analytics in Technical Communication: Advancing Information Management Toward Content Data Science. Eng. Proc. 2026, 143, 41. https://doi.org/10.3390/engproc2026143041
Ziegler W. AI-Powered Analytics in Technical Communication: Advancing Information Management Toward Content Data Science. Engineering Proceedings. 2026; 143(1):41. https://doi.org/10.3390/engproc2026143041
Chicago/Turabian StyleZiegler, Wolfgang. 2026. "AI-Powered Analytics in Technical Communication: Advancing Information Management Toward Content Data Science" Engineering Proceedings 143, no. 1: 41. https://doi.org/10.3390/engproc2026143041
APA StyleZiegler, W. (2026). AI-Powered Analytics in Technical Communication: Advancing Information Management Toward Content Data Science. Engineering Proceedings, 143(1), 41. https://doi.org/10.3390/engproc2026143041

