1. Introduction
The exponential growth of data across domains has fundamentally transformed how knowledge is extracted, interpreted, and applied. From intelligent manufacturing and legal informatics to autonomous systems, urban analytics, and scientific research synthesis, the need for advanced data analysis and data mining techniques has never been more pressing. This Special Issue brings together a diverse yet coherent collection of contributions that advance the theory, methodology, and application of knowledge discovery in this area, highlighting the interdisciplinary nature and societal relevance of the field. The scope of the Special Issue includes works on important topics, such as the following:
- Advanced Data Mining Algorithms: Explorations of new algorithms and enhancements to existing methods for effective data mining, as in Bonde et al. (2022) [1].
- Big Data Analytics: Techniques and tools for handling and analyzing large-scale datasets, including distributed computing and cloud-based solutions, as in Fadil et al. (2025) [2].
- Machine Learning and Artificial Intelligence: Integration of machine learning and AI in data analysis to improve predictive accuracy and decision-making, as in Salem et al. (2024) [3].
- Data Visualization: Innovative methods for visualizing complex data to enhance interpretability and insights, as in Cvetkoska et al. (2023) [4].
- Text and Web Mining: Approaches for extracting valuable information from unstructured text and web data, as in Rueger et al. (2023) [5].
2. Overview of Published Articles
The exponential growth of data across domains has fundamentally transformed how knowledge is extracted, interpreted, and applied. From intelligent manufacturing and legal informatics to autonomous systems, urban analytics, and scientific research synthesis, the need for advanced data analysis and data mining techniques has never been more pressing. This Special Issue brings together a diverse yet coherent collection of contributions that advance the theory, methodology, and application of knowledge discovery, highlighting the interdisciplinary nature and societal relevance of the field.
A central theme emerging from this Issue is the integration of domain knowledge with data-driven methodologies. Xu et al. (2026) (Contribution 1) address a longstanding challenge in intelligent manufacturing: the extraction of reliable process decision knowledge (PDK) from high-dimensional and experience-driven datasets. By embedding prior knowledge through an association-discriminant matrix and coupling it with the Water Wave Optimization algorithm, the authors demonstrate how hybridizing optimization techniques with knowledge constraints can significantly improve rule mining accuracy. This work exemplifies how domain-aware data mining can bridge the gap between raw data and actionable industrial intelligence.
The importance of natural language processing (NLP) for knowledge discovery in complex textual domains is further highlighted by two contributions. In Zhang et al. (2025) (Contribution 2), the authors propose HybridSumm, a lexicon-enhanced extractive–abstractive framework tailored to the challenges of long and terminology-dense legal texts. By combining pre-trained language models with domain-specific lexicons and advanced sequence modeling, the authors demonstrate how structured knowledge can be distilled from unstructured legal corpora, contributing to the advancement of judicial intelligence systems.
Complementing this, Gana et al. (2024) (Contribution 3) introduce a semi-automated framework for synthesis of the literature. By integrating large language models into the review pipeline from query formulation to thematic extraction, the proposed approach addresses the growing complexity and volume of scientific publications, not only showcasing the transformative potential of LLMs in meta-research but also underscoring the evolving role of AI in accelerating knowledge consolidation and dissemination.
Another key direction represented in this Issue is data-driven decision-making under uncertainty and sparsity. Kakimoto et al. (2025) (Contribution 4) investigate the challenging problem of extracting meaningful insights from sparse spatio-temporal data. By leveraging Gaussian mixture models for feature extraction, the authors demonstrate that even limited and noisy trajectory data can yield robust classification outcomes when appropriately modeled, a finding with broad implications for urban analytics, marketing, and public policy, where incomplete data is often the norm.
Similarly, Chen et al. (2024) (Contribution 5) highlight the power of machine learning in capturing complex relationships in real-world economic systems. Through integrating feature selection techniques such as particle swarm optimization and CatBoost alongside neural network models implemented in PyTorch, the authors demonstrate significant improvements in predictive performance. This work reinforces the importance of combining algorithmic innovation with careful feature engineering to enable accurate and scalable decision support systems.
This Issue also reflects the growing impact of intelligent systems in physical and industrial environments. In Ribeiro et al. (2025) (Contribution 6), the proposed MARS-SLAM framework introduces a novel strategy for efficient exploration and mapping in unknown environments. By leveraging virtual markers to guide navigation and ensure coverage completeness, the approach significantly reduces computational effort and operational redundancy, illustrating the synergy between data-driven optimization and robotics.
From an industrial perspective, Rojas et al. (2025) (Contribution 7) provide a comprehensive overview of the state of the art in predictive maintenance. The analysis highlights the increasing adoption of AI techniques, including deep learning, reinforcement learning, and digital twins, to improve fault detection and operational efficiency. Importantly, the study identifies key challenges such as data standardization and system interoperability, offering a roadmap for future research and industry adoption.
Finally, the role of data visualization and interpretability is addressed in Gorgol et al. (2025) (Contribution 8). By systematically exploring the Fruchterman–Reingold and ForceAtlas2 layouts, the authors emphasize the importance of effective visual representation in understanding complex data structures. Visualization remains a critical component of knowledge discovery, enabling both experts and non-experts to interpret and communicate insights derived from data.
3. Conclusions and Future Perspectives
Taken together, the contributions in this Special Issue illustrate several important trends in contemporary data analysis and data mining research:
- The increasing fusion of domain knowledge and machine learning to enhance model accuracy and interpretability;
- The growing role of deep learning and large language models in extracting knowledge from unstructured data;
- The need for robust methods to handle high-dimensional, sparse, and noisy datasets;
- The expansion of AI-driven solutions in industrial and real-world applications, from manufacturing to mining and real estate;
- The continued importance of visualization and human-centered approaches in making data insights accessible and actionable.
As data continues to grow in scale and complexity, the challenge we face is no longer merely to analyze it, but to transform it into meaningful, trustworthy, and actionable knowledge. The works presented in this Issue make significant strides toward this goal, offering both theoretical advancements and practical solutions.
We hope that this Special Issue will inspire further research at the intersection of data mining, machine learning, and knowledge discovery, fostering innovations that will shape the next generation of intelligent systems.
Author Contributions
Conceptualization, N.N. and L.d.M.M.; methodology, N.N.; writing—original draft preparation, N.N.; writing—review and editing, L.d.M.M.; funding acquisition, N.N. All authors have read and agreed to the published version of the manuscript.
Conflicts of Interest
The authors declare no conflicts of interest.
List of Contributions
- Xu, X.; Huang, Z.; Qiao, L.; Wan, Y.; Chen, C.; Li, Z. A Novel Approach for Mining Machining Process Decision Knowledge Based on Knowledge Constraint Combined with Water Wave Optimization Algorithm. Appl. Sci. 2026, 16, 2806. https://doi.org/10.3390/app16062806.
- Zhang, L.; Li, Y.; Zhang, H. Deep Learning-Based Automatic Summarization of Chinese Maritime Judgment Documents. Appl. Sci. 2025, 15, 5434. https://doi.org/10.3390/app15105434.
- Gana, B.; Leiva-Araos, A.; Allende-Cid, H.; García, J. Leveraging LLMs for Efficient Topic Reviews. Appl. Sci. 2024, 14, 7675. https://doi.org/10.3390/app14177675.
- Kakimoto, Y.; Omae, Y.; Takahashi, H. Analysis of Sparse Trajectory Features Based on Mobile Device Location for User Group Classification Using Gaussian Mixture Model. Appl. Sci. 2025, 15, 982. https://doi.org/10.3390/app15020982.
- Chen, W.; Farag, S.; Butt, U.; Al-Khateeb, H. Leveraging Machine Learning for Sophisticated Rental Value Predictions: A Case Study from Munich, Germany. Appl. Sci. 2024, 14, 9528. https://doi.org/10.3390/app14209528.
- Ribeiro, L.; Nedjah, N.; de Carvalho, P. Optimized Navigation for SLAM Using Marker-Assisted Region Scanning, Path Finding, and Mapping Completion Control. Appl. Sci. 2025, 15, 3433. https://doi.org/10.3390/app15073433.
- Rojas, L.; Peña, Á.; Garcia, J. AI-Driven Predictive Maintenance in Mining: A Systematic Literature Review on Fault Detection, Digital Twins, and Intelligent Asset Management. Appl. Sci. 2025, 15, 3337. https://doi.org/10.3390/app15063337.
- Gorgol, I.; Salwa, H. Detailed Examples of Figure Preparation in the Two Most Common Graph Layouts. Appl. Sci. 2025, 15, 2645. https://doi.org/10.3390/app15052645.
References
- Bonde, P.; Pinjarkar, L.; Cengiz, K.; Shukla, A.; Joel, M.S. New Algorithms and Technologies for Data Mining, Data Mining and Machine Learning Applications; John Wiley & Sons, Ltd.: Hoboken, NJ, USA, 2022. [Google Scholar] [CrossRef] [Scilit]
- Fadil, M.S.; Abiy, A.M.; Bealu, G.G.; Kidus, A.M.; Abenezer, G.; Rajat, K.B.; Kumod, K. Unlocking the power of machine learning in big data: A scoping survey. Data Sci. Manag. 2025, 8, 519–535. [Google Scholar] [CrossRef] [Scilit]
- Salem, A.M.; Eyupoglu, S.Z.; Maaitah, M.K. The Influence of Machine Learning on Enhancing Rational Decision-Making and Trust Levels in e-Government. Systems 2024, 12, 373. [Google Scholar] [CrossRef] [Scilit]
- Cvetkoska, V.; Eftimov, L.; Kitanovikj, B. Enchanting performance measurement and management with data envelopment analysis: Insights from bibliometric data visualization and analysis. Decis. Anal. J. 2023, 9, 100367. [Google Scholar] [CrossRef] [Scilit]
- Rueger, J.; Dolfsma, W.; Aalbers, R. Mining and analysing online social networks: Studying the dynamics of digital peer support. MethodsX 2023, 10, 102005. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.