1. Introduction
This special issue presents a notably broad view of contemporary data mining and its applications. The published work covers a wide range of topics, including online misinformation, geoscience prediction, lyrical text mining, sparse keyword analysis, antiviral peptide discovery, cybersecurity, science mapping and STEM education analytics, process mining in large-scale assessments, and multimodal clinical communication analysis. Together, these contributions show how data mining is no longer confined to a narrow set of methods or applications. Instead, they reflect the growing reach of the field into complex, diverse, and high-impact real-world domains where data-driven insights can support critical decisions. Across the collection, recurring concerns include predictive performance, interpretability, sparse and heterogeneous data, multimodal analysis, and the practical translation of analytics into domain-specific applications.
Data mining has evolved from a limited technical endeavor to a deeply interdisciplinary field. The papers in this special issue show that today’s progress is driven not only by algorithmic novelty, but also by the ability to work effectively with sparse, heterogeneous, multimodal, sequential, and domain-specific data. Whether the task is detecting emotion, estimating subsurface pressure, characterizing antiviral peptide landscapes, or understanding student response behavior, the central challenge now is to turn complex data into insights that are robust and actionable. Despite the diversity of applications, several shared methodological trends can be observed. First is the transition to architecture that combine multiple views on the same problem, for example, multi-stage social media analysis, ensemble learning, cross modal attention, or joint use of graph and metadata networks. The second is visible concern for reliability and interpretability in the form of Bayesian confidence estimation, rule based explaninability and structured modeling pipelines. The third characteristic of this collection is its clear focus on practical relevance and real-world impact. What makes these contributions especially compelling is that they are not simply benchmark-driven experiments, but investigations shaped by real and often complex environments, ranging from social media and geoscience to healthcare, education, and bioinformatics. The Special Issue was designed to explore “recent advances in data mining” from both methodological and application perspectives, inviting work on scalable algorithms, semantic and text analysis, graph mining, interpretability, and practical deployment. The published contributions clearly align with this evolving vision.
2. An Overview of Published Articles
To better understand the range of contributions in this special issue, the published work can be viewed as four broad themes: communication and text centered mining, predictive and explainable analytics in high stakes domains, structure aware discovery through networks and sequences, and learning analytics and science mapping, as shown in Table 1. Although each study explores its own unique problem, this way of organizing them help bring up common ideas, practical challenges and emerging future directions that collectively shows how the field of data mining continues to evolve.
Table 1.
An overview of contributions of the special issue.
2.1. Communication and Text Centered Mining
Contribution 1 by Al-Mutair and Berri address the fake news diffusion by specifically focusing on influential users. They integrate content, user profile and propagation cues in a staged framework, showing that hierarchical feature integration improves detection over simple baseline. Travanca, Cruz and Oliveira in contribution 2 turn text mining toward contemporary music, using lyrical analysis to reveal distinct emotional and thematic patterns in the lyrics of Ed Sheeran and Sia. Jun’s Bayesian pattern mining work in contribution 3 advances sparse keyword analysis by embedding association rule-mining in a probabilistic framework that reports expected interestingness and credible intervals. Contribution 4 by Mallarapu and colleagues extend communication-centered data mining into the clinical domain through a bidirectional framework that couples patient emotion recognition with provider behavior analysis. They applied ClinicalBERT for multi-label patient-side classification and DeBERTa/WavLM with cross-modal attention for provider-side multimodal analysis.
2.2. Predictive and Explainable Analytics
Contribution 5 by Amjad and colleagues show how ensemble-based data mining can improve pore-pressure prediction from well-log data in the Potwar Basin of Pakistan. Their hybrid-meta-ensemble combines deep learning and tree-based learning methods. They reported strong predictive performance suggesting that ensemble learning can improve predictive robustness over single models. De Bernardi and co-authors focus on interpretable anomaly detection in contribution 6. Their proposed rule-based explainable autoencoder for DNS tunnelling detection seeks performance comparable to black-box approaches. They provide analysts clearer insights into system decisions and facilitate human inspection and rule refinement when required. These studies reflect an important shift in applied data mining, where achieving high accuracy is increasingly expected to go hand in hand with transparency and meaningful use within specific domains.
2.3. Structure Aware Discovery
Contribution 7 by De Llano Garcia and co-authors map antiviral peptide chemical space using half space proximal networks and metadata networks. They identified chemically and biologically distinct communities, compact representative scaffold sets, and a set of candidate motifs that includes both previously reported and apparently novel patterns. Jerez and colleagues in contribution 8 use sequence analysis and clustering of PISA 2012 log-file data in the context of international large-scale educational assessments. They showed that “slow responders” should not be treated as a single homogeneous group. Instead, the action sequences can reveal more nuanced behavioral subtypes. Despite their differences in domain, both studies point to the same methodological lesson, i.e., meaningful patterns are often better captured when the inherent structure of the data is preserved, rather than compressed too quickly into summary features.
2.4. Learning Analytics and Science Mapping
Contribution 9 by Lopez-Meneses and co-authors contribute a science-mapping and tool-analysis study focused on quantum computing in data science and STEM education. They examined 281 Scopus indexed publications and reviewed major educational platforms including Qiskit, Quantum Inspire, QuTiP, and Amazon Braket. Their work positions data mining not only as a modeling toolkit but also as a meta-research instrument for understanding how new technical fields emerge, cluster, and translate into educational practice.
3. Emerging Research Direction
Several emerging future directions and challenges are reported by the contributions of this special issue. First challenge is the cross-domain validation. Papers have reported strong results on tightly defined datasets, but generalization across datasets, environments, and institutions remain an open challenge, whether in geoscience, cybersecurity, counseling, or educational assessments. Second challenge is uncertainty aware learning under difficult data conditions, especially for sparse, imbalanced, or long-tailed problems. The Bayesian and network-based approaches in this issue suggest that reliability estimation should become a standard design feature rather than an afterthought. A third emerging direction in interpretable multimodal analytics, where focus is not only on combining different data types but also on understanding how text, audio, structure work together to lead to decisions. A fourth area is longitudinal and sequential modeling because many real-world applications involve the evolving interactions over time rather than isolated snapshots. A fifth concern that becomes increasingly important in several application areas represented in this Special Issue is ethical deployment, specifically in social, educational, and clinical settings. These are the areas where usefulness of data mining methods relates to issues of privacy, fairness, and responsible human oversight. Finally, a sixth direction points towards human in the loop discovery. The domain experts play an active role in guiding scaffold exploration, refining interpretable rules, and contextualizing discovered patterns. Collectively, these directions suggest how the field of data mining is continuing to evolve based on the patterns that emerge across the studies in this collection.
Acknowledgments
The Guest Editor warmly thanks all the authors for their valuable contributions, the reviewers for their time and careful evaluations, and the editorial team of Computers for their continuous support during the Special Issue journey. The final collection stands as a collaborative effort that brings together a wide range of disciplines and real-world applications.
Conflicts of Interest
The author declares no conflict of interest.
List of Contributions
- Al-Mutair, H.; Berri, J. MSDSI-FND: Multi-Stage Detection Model of Influential Users’ Fake News in Online Social Networks. Computers 2025, 14, 517.
- Travanca, C.; Cruz, M.; Oliveira, A. Emotion in Words: The Role of Ed Sheeran and Sia’s Lyrics on the Musical Experience. Computers 2025, 14, 460.
- Jun, S. Sparse Keyword Data Analysis Using Bayesian Pattern Mining. Computers 2025, 14, 436.
- Mallarapu, S.; Liu, X.; Zargarian, P.; Mottaghian, S.F.; Suresha, R.; Jain, V.; Bayat, A. From Patient Emotion Recognition to Provider Understanding: A Multimodal Data Mining Framework for Emotion-Aware Clinical Counseling Systems. Computers 2026, 15, 161.
- Amjad, M.R.; Varghese, R.B.; Amjad, T. Machine Learning Models for Subsurface Pressure Prediction: A Data Mining Approach. Computers 2025, 14, 499.
- De Bernardi, G.; Gaggero, G.B.; Patrone, F.; Zappatore, S.; Marchese, M.; Mongelli, M. Rule-Based eXplainable Autoencoder for DNS Tunneling Detection. Computers 2025, 14, 375.
- García, D.d.L.; Marrero-Ponce, Y.; Agüero-Chapin, G.; Rodríguez, H.; Ferri, F.J.; Márquez, E.A.; Mora, J.R.; Martinez-Rios, F.; Pérez-Castillo, Y. Mapping the Chemical Space of Antiviral Peptides with Half-Space Proximal and Metadata Networks Through Interactive Data Mining. Computers 2025, 14, 423.
- Jerez, D.; Mazzullo, E.; Bulut, O. Exploring Slow Responses in International Large-Scale Assessments Using Sequential Process Analysis. Computers 2026, 15, 64.
- López-Meneses, E.; Cáceres-Tello, J.; Galán-Hernández, J.J.; López-Catalán, L. Quantum computing in data science and STEM education: mapping academic trends and analyzing practical tools. Computers 2025, 14, 235.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.