Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (978)

Search Parameters:
Keywords = word classification

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
38 pages, 1868 KB  
Article
Balancing Sentiment Analysis Datasets Through Representative-Word-Guided Synthetic Review Generation: A Case Study on Mexican Spanish Tourism Reviews
by Angel Díaz-Pacheco, Andrea Bethsabe García-Gutiérrez, Ansel Y. Rodríguez-González, Ramón Aranda and Miguel Á. Álvarez-Carmona
Appl. Sci. 2026, 16(15), 7398; https://doi.org/10.3390/app16157398 - 23 Jul 2026
Viewed by 219
Abstract
Class imbalance remains one of the most challenging problems in sentiment analysis, particularly in tourism review datasets where positive opinions substantially outnumber neutral and negative comments. This issue is especially critical because minority classes often contain the most valuable information regarding customer dissatisfaction, [...] Read more.
Class imbalance remains one of the most challenging problems in sentiment analysis, particularly in tourism review datasets where positive opinions substantially outnumber neutral and negative comments. This issue is especially critical because minority classes often contain the most valuable information regarding customer dissatisfaction, service failures, and opportunities for improvement. In this work, we propose a hybrid balancing methodology for sentiment analysis in Mexican Spanish tourism reviews that combines undersampling and Large Language Model (LLM)-based oversampling. The proposed framework first extracts representative words from each sentiment class using Mutual Information, then enriches them through dictionary-based or embedding-based lexical substitutions, and finally generates synthetic reviews using GPT-4o-mini guided by these representative terms. Experiments were conducted on a corpus that contains approximately 300,000 tourism reviews collected from TripAdvisor, exhibiting severe sentiment imbalance. Three undersampling strategies and multiple oversampling configurations were evaluated across six traditional machine learning classifiers and one Transformer-based model (BETO). Results show that random undersampling consistently outperformed centroid-based and K-means-based alternatives while also requiring the lowest computational cost. The best overall performance was obtained by BETO, achieving a Macro-F1 score of 0.57 compared to 0.51 on the original imbalanced dataset, representing an improvement of 11.8%. Significant gains were also observed for minority classes, with improvements exceeding 16% for the most underrepresented category. Furthermore, the proposed methodology consistently outperformed direct prompt-based generation using GPT-4o-mini, Gemini 2.5 Flash, and Llama 3.3 70B. These findings suggest that guiding synthetic review generation through representative words effectively preserves domain-specific lexical and semantic patterns of Mexican Spanish tourism reviews, resulting in more balanced datasets and improved sentiment classification performance. Full article
(This article belongs to the Special Issue Advances in Expert Systems for Natural Language Processing)
Show Figures

Figure 1

25 pages, 2321 KB  
Article
Activity Classification in E-Commerce Product Reviews Using Deep Learning and Transformer Models
by Tinashe Wamambo, Arooj Fatima, Bethwel Kiplagat, Mahdi Maktab Dar Oghaz and Cristina Luca
Informatics 2026, 13(8), 120; https://doi.org/10.3390/informatics13080120 - 23 Jul 2026
Viewed by 170
Abstract
Existing research on e-commerce product reviews has primarily focused on analysing consumers’ opinions, emotions, sentiments and associated star ratings. Whilst these approaches provide insights into consumers’ perceptions of products, they offer limited understanding of how products are used in real-world contexts. Therefore, they [...] Read more.
Existing research on e-commerce product reviews has primarily focused on analysing consumers’ opinions, emotions, sentiments and associated star ratings. Whilst these approaches provide insights into consumers’ perceptions of products, they offer limited understanding of how products are used in real-world contexts. Therefore, they do little to enhance the e-commerce experience by helping consumers make more informed purchasing decisions based on products’ intended uses without requiring them to read numerous reviews during the decision-making process. To address this problem, this paper investigates the feasibility of automatically identifying and classifying product usage activities from e-commerce reviews. A methodology combining natural language processing, manual activity-level annotation and deep learning-based text classification was developed and evaluated. An initial dataset of 60,000 Amazon product reviews was manually labelled according to six activity classes: run, walk, hike, swim, climb and unknown. Following quality inspection and data cleaning, a final dataset of 50,843 reviews was used for model training and evaluation. Multiple classification approaches were assessed, including CNN, LSTM, hybrid LSTM-CNN architectures and transformer-based models (DistilBERT and DistilBERT-CNN). Experimental evaluation was conducted using multiple random seeds to ensure robustness and reproducibility. The results indicate that activity classification from e-commerce reviews is a challenging task due to ambiguity and overlapping usage descriptions, with all evaluated models achieving comparable performance on the full dataset. Among the evaluated models, the hybrid LSTM-CNN-GloVe architecture achieved the highest performance on the keyword-filtered dataset, whilst the DistilBERT-CNN model also demonstrated strong results. The findings demonstrate the feasibility of extracting activity-oriented information from product reviews and highlight activity classification as a distinct and under-explored natural language processing task that complements traditional sentiment analysis. The proposed methodology provides a foundation for improving product discovery and supporting usage-oriented search and recommendation systems in e-commerce environments. Full article
(This article belongs to the Section Big Data Mining and Analytics)
Show Figures

Figure 1

22 pages, 4603 KB  
Article
A Phase-Coherent Four-Stage Pipeline for the Dereverberation of Quránic Recitation
by Osama Al Maaini, Khizar Hayat, Khalil Al Ruqeishi and Baptiste Magnier
Information 2026, 17(7), 714; https://doi.org/10.3390/info17070714 - 22 Jul 2026
Viewed by 198
Abstract
The accuracy of spectro-temporal features for Makhaarij al-Huroof and Sifaat distinguishes between the ten canonical Qiraát recitation styles of the Holy Quran. However, real-world room reverberations blur formant contours and corrupt inter-word energies, thus making Qiraat discrimination difficult. The current dereverberation methods were [...] Read more.
The accuracy of spectro-temporal features for Makhaarij al-Huroof and Sifaat distinguishes between the ten canonical Qiraát recitation styles of the Holy Quran. However, real-world room reverberations blur formant contours and corrupt inter-word energies, thus making Qiraat discrimination difficult. The current dereverberation methods were designed to work under ordinary speech conditions and are not capable of preserving phonetic qualities for domain-specific purposes. This paper introduces a four-step, phase-consistent signal-processing approach prioritizing phonetic preservation over direct reverberation suppression. The four steps are: (1) adaptive noise-floor attenuation; (2) soft-voice activity detection using power-law boundary decay; (3) application-specific spectral contour adjustment from clean Quranic reference audio; and (4) Griffin–Lim algorithm-based phase correction. A total of 48 real-world room recordings were utilized for the evaluation of this approach based on Energy Ratio (ER), Spectral Contrast (SC), and Spectral Contour Stability (SCS)—measures specific to the Quran audio domain—alongside conventional speech-quality metrics. The proposed approach yielded the highest scores in three of seven metrics, namely SC (+40.11), SCS (+822.94), and PESQ (+1.251), alongside the second-highest Energy Ratio (+19.58 dB), while being superior to Spectral Subtraction, Wiener Filtering, and WPE Dereverberation approaches. Moreover, the perceptual enhancement was verified in a synthetic controlled experiment where the proposed approach scored an improved PESQ metric (+2.495; SNR −1.874 dB). The results illustrate the fact that an optimization for general-purpose metrics does not necessarily ensure phonetic preservation required for specific classification. Full article
(This article belongs to the Section Information Applications)
Show Figures

Graphical abstract

30 pages, 7319 KB  
Article
Computational Modelling of Institutional Perceptions in a Civilisational Intergovernmental Organisation
by Fahim Sufi, A. K. M. Iftekharul Islam and Anowara Akter
Computation 2026, 14(7), 164; https://doi.org/10.3390/computation14070164 - 21 Jul 2026
Viewed by 214
Abstract
Civilisational intergovernmental organisations occupy a distinctive position in global governance because their legitimacy is shaped by institutional performance, symbolic identity, collective representation, and normative expectation. This study develops a sentiment-driven computational modelling framework for analysing institutional perceptions in a 57-member civilisational intergovernmental organisation. [...] Read more.
Civilisational intergovernmental organisations occupy a distinctive position in global governance because their legitimacy is shaped by institutional performance, symbolic identity, collective representation, and normative expectation. This study develops a sentiment-driven computational modelling framework for analysing institutional perceptions in a 57-member civilisational intergovernmental organisation. The empirical corpus comprises 30 semi-structured interviews, including 15 expert interviews and 15 general respondent interviews. After preprocessing, removal of administrative material, and exclusion of courtesy-only responses, the final corpus contained 239 answer segments and 42,940 respondent-generated words. The methodology integrates answer-level segmentation, thematic classification, lexical sentiment analysis, transformer-based sentiment modelling, latent semantic clustering, non-parametric statistical testing, bootstrap confidence estimation, and manual validation. The results show that institutional perception is broadly positive but thematically uneven. Experts recorded a mean polarity score of 0.0803, while general respondents recorded 0.1016. However, group-level differences were not statistically significant at segment level, Mann–Whitney U=6171.50, p=0.0704, or interview level, U=71.00, p=0.0888. In contrast, sentiment differed significantly across institutional themes, Kruskal–Wallis H=45.8768, p<0.001, and across latent semantic clusters, with Kruskal–Wallis H=35.0124 and p<0.001. The strongest positive sentiment appeared in reform and future orientation, mean polarity =0.1700, socioeconomic cooperation, 0.1330, and institutional effectiveness, 0.1325. The weakest sentiment concerned institutional weakness and constraints, 0.0546, and political and security role, 0.0693. Latent semantic clustering identified socioeconomic and economic cooperation as the most positively evaluated cluster, with mean polarity 0.1465. Overall, the findings reveal a measurable tension between symbolic-developmental legitimacy and operational scepticism. Full article
Show Figures

Figure 1

40 pages, 4767 KB  
Review
The Experience and Use of Power Mobility by Children with Complex Non-Ambulant Cerebral Palsy: A Scoping Review
by Roslyn W. Livingstone, Ginny S. Paleg, Benjamin W. Fullerton, Débora Claësson, Pragashnie Govender and Lisbeth Nilsson
Disabilities 2026, 6(4), 64; https://doi.org/10.3390/disabilities6040064 - 16 Jul 2026
Viewed by 403
Abstract
Background/Objectives: To map the literature and describe the meaning, use, and experience of power mobility for children with complex non-ambulant cerebral palsy (Gross Motor Classification System (GMFCS) levels IV–V and Manual Abilities Classification System (MACS) levels III–V). Methods: Included searches in [...] Read more.
Background/Objectives: To map the literature and describe the meaning, use, and experience of power mobility for children with complex non-ambulant cerebral palsy (Gross Motor Classification System (GMFCS) levels IV–V and Manual Abilities Classification System (MACS) levels III–V). Methods: Included searches in five electronic databases, grey literature, and hand searches with no restrictions on date, study type, or language, as well as independent duplicate screening and data extraction. Outcomes and experiences were mapped to the integrated F-words Interdependence Human Activity Assistive Technology (iHAAT) framework. Results: In total, 90 studies, from randomized trials to case reports and qualitative designs, included 916 children (10 months–18 years; 432 GMFCS IV; 262 GMFCS V; 222 GMFCS IV/V), with 351 parents, therapists, or educators. Only 32 studies reported MACS levels. Power wheelchairs were used by 724 children (68 used switches rather than joysticks). Other children used modified ride-on cars, specialty pediatric devices, or platform/smart training devices. Based on 22 studies where this information was provided, alternate access/control methods were primarily used by children classified at GMFCS/MACS V, but there was considerable variability. Introduction predominantly occurred in natural settings with limited training or support. Significant and meaningful improvements in power mobility use were reported for intensive play-based, child-led, and caregiver-supported approaches; for virtual training with joystick users; and for skills-training approaches with older children who already achieved functional power wheelchair use. Conclusions: Children classified at GMFCS IV and V may benefit from power mobility experience to promote fitness, functioning, friends, family, fun, and future outcomes. Their use and experience of power mobility may be interdependent with parents, therapists, and educators, changing attitudes and perceptions of child potential. Full article
Show Figures

Graphical abstract

19 pages, 1648 KB  
Article
A Secure Lightweight SMS Spam Detection Framework with Robustness to Text Obfuscation Attacks
by Baraa Tareq Hammad, Ismail Taha Ahmed, Mohamed A. Hafez and Betty Wan Niu Voon
Computers 2026, 15(7), 451; https://doi.org/10.3390/computers15070451 - 16 Jul 2026
Viewed by 206
Abstract
The proliferation of mobile communications has led to a significant increase in SMS spam, posing challenges related to security, privacy, and user experience. Although numerous machine-learning-based spam detection approaches have been proposed, developing systems that are simultaneously lightweight and resilient to adversarial manipulation [...] Read more.
The proliferation of mobile communications has led to a significant increase in SMS spam, posing challenges related to security, privacy, and user experience. Although numerous machine-learning-based spam detection approaches have been proposed, developing systems that are simultaneously lightweight and resilient to adversarial manipulation remains an open problem. This paper proposes an SMS spam detection framework that incorporates multiple feature extraction methods, including bag-of-words (BoW), Term Frequency–Inverse Document Frequency (TF-IDF), and N-gram models with dimensionality reduction using principal component analysis (PCA), followed by classification using decision tree (DT) and Logistic Regression (LogReg) models. Experimental evaluations on the UCI SMS Spam Collection dataset demonstrate that the TF-IDF-PCA-DT pipeline achieves a detection accuracy of 99% while reducing model size by 77% and inference time by 75%. Robustness evaluation under adversarial text perturbations indicates minimal performance degradation, maintaining an accuracy of 96.5%. These findings demonstrate the practicality of the proposed framework for real-world deployment in resource-constrained environments. Full article
Show Figures

Figure 1

18 pages, 926 KB  
Article
Construction of Customized Personas for Decision-Making Cognition Regarding Oral Microbiota Transplantation in Head and Neck Cancer Patients Undergoing Radiotherapy: A Qualitative Study
by Xue Liu, Hang Wang, Xinyao Yang, Yufei Li, Like Zhang, Lei Cui, Hao Li and Lili Hou
Healthcare 2026, 14(14), 2073; https://doi.org/10.3390/healthcare14142073 - 10 Jul 2026
Viewed by 254
Abstract
Background: Patients with head and neck cancer who are undergoing radiotherapy frequently suffer from oral mucositis and oral microecological disorders, which severely impair their quality of life. Oral microbiota transplantation is an emerging oral microecological intervention that offers a novel approach for [...] Read more.
Background: Patients with head and neck cancer who are undergoing radiotherapy frequently suffer from oral mucositis and oral microecological disorders, which severely impair their quality of life. Oral microbiota transplantation is an emerging oral microecological intervention that offers a novel approach for reconstructing oral microecological balance and relieving mucositis. However, regarding this innovative therapy, there is a paucity of in-depth research into patients’ decision-making cognition, and existing evidence is insufficient to support individualized clinical decision-making guidance. Methods: A descriptive qualitative research design was employed. From July to December 2025, patients diagnosed with head and neck cancer undergoing radiotherapy were recruited from a tertiary hospital in Shanghai via purposive sampling. The data were collected through semi-structured interviews and analyzed using Colaizzi’s seven-step analysis method. The user label system was refined and summarized to construct user portraits. These portraits were visualized in the form of WordArt word clouds and character labels. Results: A total of 21 eligible patients with head and neck cancer undergoing radiotherapy participated in the study. The construct of decision-making cognition encompasses five dimensions: treatment prioritization, information needs, health literacy, psychological status, and decision quality. The patients were categorized into four types: proactive participation, passive dependence, weigh carefully, and symptom-driven. These classifications reflect the cognitive characteristics and group differences regarding the Oral Microbiota Transplantation decision-making process among different patients. Conclusions: Patients exhibit considerable variability in their decision-making cognition regarding the innovative OMT therapy. This phenomenon can be categorized into four distinct persona types, which, respectively, reflect unique information processing styles, risk assessments, and behavioral coping strategies when patients encounter novel therapeutic interventions. This typology provides a theoretical foundation for individualized clinical decision support, delineates targets for the formulation of targeted communication strategies, and ultimately enhances patient decision quality and treatment adherence. Full article
Show Figures

Figure 1

23 pages, 1417 KB  
Article
EPECT: An Eigenvalue-Guided Positional Encoding Classification Transformer for Cross-Subject EEG-fNIRS Decoding
by Chayut Bunterngchit, Laith H. Baniata and Sangwoo Kang
Mathematics 2026, 14(13), 2416; https://doi.org/10.3390/math14132416 - 6 Jul 2026
Viewed by 296
Abstract
Decoding mental states from non-invasive neural recordings is central to brain-computer interface research. Multimodal acquisition that combines electroencephalography (EEG) and functional near-infrared spectroscopy (fNIRS) couples the high temporal resolution of EEG with the spatial specificity of fNIRS, compensating for the individual limitations of [...] Read more.
Decoding mental states from non-invasive neural recordings is central to brain-computer interface research. Multimodal acquisition that combines electroencephalography (EEG) and functional near-infrared spectroscopy (fNIRS) couples the high temporal resolution of EEG with the spatial specificity of fNIRS, compensating for the individual limitations of each modality. While such hybrid systems achieve strong intra-subject performance, cross-subject generalization remains constrained by inter-individual variability in neural responses. This study introduces the Eigenvalue-Guided Positional Encoding Classification Transformer (EPECT), an architecture that integrates eigenvalue-aware multi-head self-attention with sinusoidal positional encoding to capture both the spectral structure of the learned feature representations and the temporal ordering of multimodal sequences. Stacked one-dimensional convolutions extract local patterns prior to transformer encoding, and global average pooling aggregates the final representation for classification. EPECT was evaluated on two publicly available EEG-fNIRS datasets covering motor imagery (MI), n-back, discrimination/selection response (DSR), and word generation (WG) paradigms under a cross-subject protocol. The model achieved classification accuracies of 97.3%, 96.3%, 98.1%, and 97.9% on the MI, n-back, DSR, and WG tasks, respectively. Ablation studies quantified the contribution of each architectural component, and integrated gradients analysis revealed structured modality-specific attribution patterns aligned with task-relevant cortical regions. Additional experiments with synthetic cortical perturbations demonstrate the sensitivity of EPECT to subtle activity changes, indicating potential utility for tracking neurorehabilitation outcomes in future clinical applications. Full article
Show Figures

Figure 1

22 pages, 700 KB  
Article
Cross-Layer Resource Optimization for Ultra-Low-Power TinyML Inference on ARM Cortex-M Microcontrollers
by Abdulaziz G. Alanazi, Haifa A. Alanazi and Nasser S. Albalawi
Electronics 2026, 15(13), 2918; https://doi.org/10.3390/electronics15132918 - 3 Jul 2026
Viewed by 374
Abstract
Running neural networks on battery-powered Internet of Things (IoT) sensor nodes is difficult because flash memory, SRAM, latency, and energy per inference are limited at the same time. Existing TinyML co-design methods usually improve model size or memory use, but runtime voltage–frequency control [...] Read more.
Running neural networks on battery-powered Internet of Things (IoT) sensor nodes is difficult because flash memory, SRAM, latency, and energy per inference are limited at the same time. Existing TinyML co-design methods usually improve model size or memory use, but runtime voltage–frequency control is often handled as a separate step. This separation limits energy saving because the power policy does not use the layer-wise compute profile of the final compressed model. We propose the Cross-Layer Resource Optimizer (CLRO), a three-stage resource optimization pipeline for TinyML inference on an ARM Cortex-M7 target. The first stage, Mixed-Precision Aware Pruning and Distillation (MPAD), assigns per-layer bit widths and pruning ratios using calibration-set sensitivity scores. The second stage, consisting of the Activation Lifetime-Aware Tensor Scheduler (ALTS), uses the compressed graph to find an execution order that reduces peak live static random-access memory (SRAM). The third stage, Reinforcement Learning-Based Dynamic Voltage and Frequency Scaling (DVFS-RL), trains a tabular Q-learning policy from the multiply–accumulate (MAC) utilization profile of the compressed and scheduled model. The learned voltage–frequency policy is stored as a small flash lookup table, so it adds no runtime decision cost during inference. We evaluate the CLRO on all four MLPerf Tiny tasks using an STM32H743ZI microcontroller with 512 kB SRAM and 2 MB flash. The CLRO reaches 91.7% image classification accuracy, 95.4% keyword-spotting accuracy, 89.6% visual wake words accuracy, and 0.913 anomaly detection AUC. The final deployment uses 198 kB flash and 174 kB peak SRAM, with 387 μJ energy per inference and 38 ms latency. Compared with the MCUNet baseline, the CLRO reduces energy by 58.1% and peak SRAM by 39% while keeping the same accuracy level. Full article
Show Figures

Figure 1

20 pages, 6116 KB  
Article
SlideRing: Robust Dual-IMU Thumb-to-Finger Text Input for Virtual Reality
by Tao Sun, Nuo Jia and Dawei Jiao
Sensors 2026, 26(13), 4210; https://doi.org/10.3390/s26134210 - 3 Jul 2026
Viewed by 250
Abstract
Text entry remains a bottleneck for productivity-oriented Virtual Reality (VR), especially in scenarios where optical hand tracking is unstable because of self-occlusion, poor lighting, or out-of-view interaction. We present SlideRing, a dual-thumb wearable text-entry method that senses thumb-to-finger micro-gestures with two miniature Inertial [...] Read more.
Text entry remains a bottleneck for productivity-oriented Virtual Reality (VR), especially in scenarios where optical hand tracking is unstable because of self-occlusion, poor lighting, or out-of-view interaction. We present SlideRing, a dual-thumb wearable text-entry method that senses thumb-to-finger micro-gestures with two miniature Inertial Measurement Units (IMUs). SlideRing defines a 30-command interaction space from two hands, three target fingers, and five gesture types, then maps these commands to a full alphabetic keyboard through two complementary strategies: an ergonomic layout optimized for low movement cost and a QWERTY-compatible layout optimized for learnability. To decode subtle inertial signals, we design a dual-stream recognition model with a Statistical Feature Encoder, a Temporal Feature Encoder, and a context-aware gating module for joint finger–action classification. In offline evaluation, the model reaches 96.5% target-finger accuracy and 94.2% action-type accuracy. In a five-day text-entry study, the ergonomic layout improves from 7.43 to 15.75 words per minute (WPM), while the QWERTY-compatible layout improves from 10.55 to 15.25 WPM. The ergonomic layout reduces physical demand, whereas the QWERTY-compatible layout lowers initial mental load. These results suggest that IMU-based thumb-to-finger input has the potential to provide robust, low-visual-demand text entry for constrained VR environments. Full article
(This article belongs to the Section Wearables)
Show Figures

Figure 1

25 pages, 2077 KB  
Article
From API to Action: A Multi-Model Comparison of OpenAI, Anthropic, Google, and Meta LLMs for Clinical Trial Data Extraction
by Richard J. Young, Jorge Fonseca and Brach Poston
Bioengineering 2026, 13(7), 773; https://doi.org/10.3390/bioengineering13070773 - 2 Jul 2026
Viewed by 763
Abstract
(1) Background: Clinical trial data extraction from registries such as ClinicalTrials.gov remains labor-intensive and error-prone, often missing critical details hidden in unstructured protocol descriptions. Large Language Models (LLMs) offer potential to automate this process, yet systematic multi-model comparisons on real clinical trial data [...] Read more.
(1) Background: Clinical trial data extraction from registries such as ClinicalTrials.gov remains labor-intensive and error-prone, often missing critical details hidden in unstructured protocol descriptions. Large Language Models (LLMs) offer potential to automate this process, yet systematic multi-model comparisons on real clinical trial data remain scarce. (2) Methods: Four LLMs (OpenAI o4-mini-high, Anthropic Claude-Sonnet-4, Google Gemini 2.5-Pro, and Meta Llama-4-Maverick) extracted brain stimulation parameters from 67 transcranial direct current stimulation (tDCS) trials in Parkinson’s disease via a structured JSON schema. Pairwise inter-model agreement was quantified with Cohen’s Kappa and percentage agreement across binary, categorical, and multi-component task tiers. (3) Results: Under exact-string matching, agreement was near-perfect for binary classifications (non-invasive classification: 100%; brain stimulation presence: 99.3%, κ = 0.50) and substantial for categorical extractions (primary stimulation type: 96.4%, κ = 0.70), but fell to 48.6% (κ = 0.43) for complex anatomical targets. Numeric parameters revealed model-specific strengths: o4-mini-high and Claude-Sonnet-4 achieved perfect duration agreement (r = 1.000, n = 19) while Llama-4-Maverick diverged substantially (r < 0.12). Validation against an expert gold standard (100% inter-annotator agreement on a 20-trial overlap) confirmed high extraction accuracy across all features (mean 93.7–98.9%). Crucially, the low agreement on anatomical targets proved to be an artifact of exact-string scoring: under the same semantic matching used to measure accuracy, inter-model agreement rose to 97.0%, coinciding with the 95.5% expert accuracy. Inter-model agreement therefore tracks accuracy once both are measured on a common basis. (4) Conclusions: Exact-string inter-model agreement decreases with task complexity, but this decline largely reflects interchangeable free-text wording rather than reduced accuracy. Evaluated semantically, agreement and expert accuracy are both high and closely aligned. A residual risk is not low accuracy but the rare error shared across all models, which agreement cannot detect, and which overall accuracy can itself mask when one class dominates. These findings inform hybrid human–AI systematic review pipelines in which targeted expert oversight focuses on shared-error and minority-class detection. Full article
(This article belongs to the Special Issue Biomedical Data Mining: Emerging Methods and Applications)
Show Figures

Graphical abstract

44 pages, 3647 KB  
Article
Forensic-BERT: Explainable Transformer-Based Detection of Concealed Evidence in Cross-Platform Volatile Memory
by Yousef Sanjalawe, Salam Al-E’mari and Sharif Naser Makhadmeh
Computers 2026, 15(7), 420; https://doi.org/10.3390/computers15070420 - 29 Jun 2026
Viewed by 369
Abstract
Advanced cyber threats increasingly exploit volatile memory to execute malicious payloads without touching persistent storage, rendering traditional disk-centric forensic tools insufficient for comprehensive digital investigations. This paper presents Forensic-BERT, an AI-driven forensic framework that automatically extracts and classifies potentially relevant artifacts from unstructured [...] Read more.
Advanced cyber threats increasingly exploit volatile memory to execute malicious payloads without touching persistent storage, rendering traditional disk-centric forensic tools insufficient for comprehensive digital investigations. This paper presents Forensic-BERT, an AI-driven forensic framework that automatically extracts and classifies potentially relevant artifacts from unstructured memory dumps across heterogeneous operating environments. The framework combines byte-boundary-preserving Hex-to-ASCII conversion, sliding-window Shannon entropy filtering (H>7.2 bits per byte, 256-byte windows) to isolate high-probability artifact regions, and a binary-aware WordPiece tokenizer extended with 2048 domain-specific tokens covering hexadecimal byte patterns, Windows API names, and Linux system-call sequences. These components feed a transformer-based classifier fine-tuned from bert-base-uncased (110 M parameters) on memory-derived text, with sliding-window inference and majority-vote aggregation for large images. A SHAP DeepExplainer module and averaged 12-head attention heatmaps provide transparent, analyst-accessible explanations for classification decisions. We evaluate the framework on a multi-source corpus of 735 labeled memory segments drawn from 197 distinct images across four independent collections, MemLabs, the DARPA Transparent Computing program, Digital Corpora, and live sandbox execution traces from Any.run and Joe Sandbox, spanning Windows XP through Windows 11, Ubuntu Linux 16.04/18.04, and FreeBSD. Source-stratified five-fold cross-validation yields an overall F1-score of 0.92±0.02 and AUC-ROC of 0.95±0.01 (95% CI). Forensic-BERT outperforms all six baselines, Volatility with YARA rules (F1 =0.71), Random Forest (F1 =0.82), BiLSTM with GloVe embeddings (F1 =0.85), MRm-DLDet (F1 =0.87), SPECTRE (F1 =0.89), and SecBERT (F1 =0.90), with every pairwise difference statistically significant under the McNemar test with Bonferroni correction. Explainability quality is independently confirmed by a Spearman rank correlation of ρ=0.81 between model SHAP token rankings and expert forensic-indicator rankings and by a System Usability Scale score of 73.2 among certified examiners. The complete pipeline processes 512 MB memory images in 7.5–10.2 s (GPU) or 38–52 s (CPU-only), scaling to 4 GB images with near-linear throughput. These results indicate that, on the corpus evaluated here, combining domain-adapted NLP preprocessing, transformer-based sequence modeling, and quantified explainability can improve the effectiveness and usability of analyst decision support and investigative triage for volatile memory analysis. Full article
(This article belongs to the Section AI-Driven Innovations)
Show Figures

Figure 1

18 pages, 21844 KB  
Article
Evaluating Cultural Ecosystem Services of Nature-Based Solutions in Urban Renewal Using Social Media Data
by Xin Cheng, Peisi Xu and Sylvie Van Damme
Forests 2026, 17(7), 749; https://doi.org/10.3390/f17070749 - 27 Jun 2026
Viewed by 334
Abstract
Urban renewal increasingly adopts Nature-Based Solutions (NBSs) to address environmental challenges and enhance social well-being. However, it remains unclear whether and to what extent NBSs contribute to cultural ecosystem services (CESs), which reflect people’s perceptions, values, and experiences of urban nature. This study [...] Read more.
Urban renewal increasingly adopts Nature-Based Solutions (NBSs) to address environmental challenges and enhance social well-being. However, it remains unclear whether and to what extent NBSs contribute to cultural ecosystem services (CESs), which reflect people’s perceptions, values, and experiences of urban nature. This study develops an integrated framework combining text and image mining of social media data to evaluate the CES outcomes of NBS in regenerated urban districts in Chengdu, China. The comment data were analyzed for CES using Jieba word segmentation and dictionary matching, while images were categorized into NBS types by manual classification. By integrating these multimodal data, the framework effectively clarifies the relationship between NBSs and CESs from the perspective of public perception. Results indicate that recreation and leisure, inspiration, and spiritual values are the most prominent aspects of public perception, with linear green infrastructure and pocket parks being the most frequently identified NBS types. Correspondence analysis further reveals significant associations between specific NBS interventions and CES categories. By integrating textual and visual data, this study offers a practical and real-time approach for capturing public perceptions of CESs and provides actionable insights for the design and management of NBS-driven urban regeneration. Full article
Show Figures

Figure 1

32 pages, 3409 KB  
Article
xServeNet: An Explainable Deep Neural Network for Web Services Classification
by Yilong Yang, Muhammad Ali Khan, Zhaotian Li and Weiru Wang
Electronics 2026, 15(12), 2711; https://doi.org/10.3390/electronics15122711 - 18 Jun 2026
Viewed by 306
Abstract
Web service classification plays an important role in software reuse, service discovery, and automatic metadata organization. Although recent deep learning approaches have improved classification performance by using service names and natural-language descriptions, most existing methods still operate as black-box models and offer limited [...] Read more.
Web service classification plays an important role in software reuse, service discovery, and automatic metadata organization. Although recent deep learning approaches have improved classification performance by using service names and natural-language descriptions, most existing methods still operate as black-box models and offer limited insight into how different metadata sources influence classification decisions. This lack of transparency reduces their practical usefulness for developers who need to verify predicted categories, analyze incorrect classifications, and improve service metadata quality. A well-trained interpretable model can not only help developers choose more appropriate and reliable categories for each web service, but also help write a more reasonable service name and description. In this paper, we present xServeNet, an explainability-oriented extension of ServeNet for transparent web service classification. xServeNet preserves the BERT-based representation and CNN–BiLSTM feature extractor of ServeNet and introduces (i) an instance-wise dynamic source-fusion mechanism that adaptively combines service-name and service-description features according to their semantic contribution, and (ii) model-internal importance indicators at both the source and word levels that support inspection of classification decisions without introducing additional trainable parameters. We benchmark xServeNet against eleven machine learning baselines on two real-world ProgrammableWeb datasets of 10,943 and 14,086 services covering 50 categories. xServeNet reaches 71.08% Top-1/91.35% Top-5 accuracy on the original dataset and 74.10% Top-1/92.95% Top-5 accuracy on the updated dataset, consistently improving Top-1 accuracy over ServeNet while remaining competitive on Top-5, and achieving the lowest per-category Top-5 standard deviation among all twelve compared methods. In practice, the importance indicators support three concrete activities at the service registry: helping developers verify predicted categories at registration time, iterating on description wording when the predicted category looks wrong, and supporting registry curators in flagging likely mislabelled services for review. Full article
(This article belongs to the Special Issue New Trends in Machine Learning, System and Digital Twins)
Show Figures

Figure 1

24 pages, 7402 KB  
Article
Public Value Perception and Conservation Strategies for Urban Industrial Heritage: Evidence from UGC
by Ziyang Wang, Qixuan Zhou, Yi Tai, Rong Zhu and Kexin Wei
Buildings 2026, 16(12), 2391; https://doi.org/10.3390/buildings16122391 - 16 Jun 2026
Viewed by 331
Abstract
Urban industrial heritage is increasingly embedded in urban regeneration, public space provision, and community governance, yet existing studies have insufficiently examined how heterogeneous publics perceive its value through everyday digital discourse. Taking the Guangzhou Iron and Steel Plant industrial heritage site (hereafter, the [...] Read more.
Urban industrial heritage is increasingly embedded in urban regeneration, public space provision, and community governance, yet existing studies have insufficiently examined how heterogeneous publics perceive its value through everyday digital discourse. Taking the Guangzhou Iron and Steel Plant industrial heritage site (hereafter, the Guanggang industrial heritage site) as a case study, this study used user-generated content from Rednote posts and local WeChat public-account comments to identify platform-mediated expressions of public value perception. A corpus of 745 valid samples comprising 51,459 Chinese characters was constructed after data collection, screening, and text preprocessing. Word-frequency analysis, semantic network analysis, and sentiment analysis were conducted using ROST CM 6.0. The results show that the two retrieved platform-contextual corpora foregrounded different concerns. Rednote discourse foregrounded ruin landscapes, industrial aesthetics, photography-based check-ins, and exploratory experiences, whereas WeChat comments emphasized park construction, public facilities, governance responsiveness, safety, and the residential environment. At the corpus level, lexicon-based sentiment classification indicated that Rednote texts were dominated by positive and neutral categories, while WeChat comments contained a higher proportion of texts classified as negative. This study conceptualizes dual foregrounding as a bounded selection process through which platform affordances, user self-selection, and users’ relationships with the site influence which concerns become visible in each corpus; it does not treat the observed differences as a causal platform effect. It argues that industrial heritage regeneration must translate historical, technological, and aesthetic values into public values that are interpretable, accessible, usable, and trusted by local communities. Full article
Show Figures

Figure 1

Back to TopTop