Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (29)

Search Parameters:
Keywords = variable recognition ambiguity

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
35 pages, 4214 KB  
Article
Biomedical Text Mining and Information Extraction Using Prompt-Enhanced and LoRA-Adapted Large Language Models
by Feng Yan, Dequan Zheng, Feng Yu and Jing Kang
Appl. Sci. 2026, 16(16), 7956; https://doi.org/10.3390/app16167956 - 10 Aug 2026
Viewed by 343
Abstract
Biomedical named entity recognition (NER) and relation extraction (RE) remain challenging because biomedical texts contain ambiguous abbreviations, complex entity boundaries, domain-specific terminology, and implicit relations. This study proposes a prompt-enhanced and QLoRA-adapted large language model framework for biomedical information extraction. For NER, abbreviation-aware [...] Read more.
Biomedical named entity recognition (NER) and relation extraction (RE) remain challenging because biomedical texts contain ambiguous abbreviations, complex entity boundaries, domain-specific terminology, and implicit relations. This study proposes a prompt-enhanced and QLoRA-adapted large language model framework for biomedical information extraction. For NER, abbreviation-aware prompting supports candidate detection, contextual interpretation, boundary-aware generation, and schema-constrained outputs. For RE, entity markers identify a predefined target pair, while filtered UMLS and MeSH concepts provide concise evidence. DeepSeek-R1-Distill-Qwen-7B is adapted using LoRA over a 4-bit quantized frozen backbone. Experiments cover three NER and three RE datasets. Across three training seeds, the dataset-level macro-average F1 values are 0.909 ± 0.001 for NER and 0.787 ± 0.001 for RE. Seed-balanced paired bootstrap resampling with 10,000 aligned instance-level resamples confirms significant improvements over a matched deterministic simple-prompt baseline on all six datasets after Holm–Bonferroni correction (adjusted p < 0.001), with absolute F1 gains from +0.091 to +0.131. Repeated-run ablations show low variability and complementary contributions from task-structured prompting, knowledge filtering, deterministic validation, and parameter-efficient adaptation. Full article
Show Figures

Figure 1

25 pages, 4200 KB  
Article
Challenges in Emotion Recognition Across Modalities: A Comparative Analysis
by Rafał Gasz
Appl. Sci. 2026, 16(14), 7239; https://doi.org/10.3390/app16147239 - 20 Jul 2026
Viewed by 387
Abstract
Emotion recognition remains a challenging task despite substantial progress in machine learning and affective computing. This study examines challenges in emotion recognition through a comparative analysis of two widely used modalities: facial images and speech signals. The analysis was conducted using FER-2013 for [...] Read more.
Emotion recognition remains a challenging task despite substantial progress in machine learning and affective computing. This study examines challenges in emotion recognition through a comparative analysis of two widely used modalities: facial images and speech signals. The analysis was conducted using FER-2013 for facial emotion recognition and the TESS and RAVDESS datasets for speech emotion recognition. A MobileNetV2-based approach was applied to visual data, while speech analysis employed MFCC-based representations and both classical and deep learning models. The study combines quantitative performance evaluation with qualitative analysis of classification behavior, focusing on emotion-specific recognition difficulties and recurring error patterns across modalities. Model performance was assessed using accuracy, precision, recall, F1-score, and confusion matrices. Across the analysed datasets, overall classification accuracy ranged from approximately 73% to 96%, while class-level F1-scores ranged from 0.48 to 0.89 depending on the emotion and modality. Happiness and surprise consistently achieved the highest recognition performance, whereas neutral emotion, fear, and disgust exhibited the lowest class-level F1-scores and generated the highest numbers of misclassifications. The experimental results confirmed that happiness and surprise achieved the highest classification performance across modalities, while neutral emotion, fear, and disgust showed reduced recognition accuracy due to weak expressive cues and overlapping feature representations. These difficulties are associated with weak or ambiguous expressive signals, overlap between emotional categories, and variability in emotional expression. The comparative findings suggest that recognition challenges arise from both modality-specific limitations and the inherent properties of emotional expression. The results highlight the importance of multimodal approaches and more flexible representations for improving emotion recognition systems. Full article
(This article belongs to the Special Issue Computational Models and Machine Learning for Biomedical Applications)
Show Figures

Figure 1

34 pages, 32804 KB  
Article
Emotion-Aware Contextual Modelling for Robust Driver Fatigue Detection
by Sebastian Budzan and Roman Wyżgolik
Sensors 2026, 26(13), 4120; https://doi.org/10.3390/s26134120 - 30 Jun 2026
Viewed by 496
Abstract
Vision-based driver fatigue detection remains challenging because facial signals associated with fatigue are often ambiguous, while geometric indicators such as Eye Aspect Ratio (EAR) and Percentage of Eye Closure (PERCLOS) are prone to false positives caused by normal facial activity, including smiling or [...] Read more.
Vision-based driver fatigue detection remains challenging because facial signals associated with fatigue are often ambiguous, while geometric indicators such as Eye Aspect Ratio (EAR) and Percentage of Eye Closure (PERCLOS) are prone to false positives caused by normal facial activity, including smiling or speaking. This paper proposes a context-aware framework that integrates behavioural, geometric, and emotional information for robust fatigue assessment. Facial landmarks are extracted using MediaPipe Face Mesh, while adaptive eye-closure detection is performed through multi-stage validation combining EAR trajectories, mouth activity, head-pose analysis, and event-level filtering. Emotion recognition is achieved using an EfficientNet-B0 convolutional neural network trained on the AffectNet dataset, enabling frame-level estimation of facial expression probabilities. These predictions are aggregated into descriptors representing emotional variability and fatigue-related emotional relevance over time. Behavioural information obtained from blinking, yawning, head nodding, and validated PERCLOS is fused with emotional context to construct a multi-level fatigue assessment model. The final Driver Fatigue Risk Index combines physiological eye-closure information with contextual behavioural–emotional analysis, providing an interpretable estimation of driver state rather than a binary classification alone. Experimental evaluation on the NTHU-DDD dataset achieved 94% accuracy and demonstrated improved robustness under non-frontal head poses and expressive facial behaviour. Full article
(This article belongs to the Section Optical Sensors)
Show Figures

Figure 1

23 pages, 2606 KB  
Article
Adaptive Confidence-Gated Hybrid Ensemble Framework for Speech Emotion Recognition
by Salem Titouni, Nadhir Djeffal, Abdallah Hedir, Massinissa Belazzoug, Boualem Hammache and Idris Messaoudene
Electronics 2026, 15(9), 1931; https://doi.org/10.3390/electronics15091931 - 2 May 2026
Viewed by 508
Abstract
Speech Emotion Recognition (SER) is a key enabling technology for advanced human–computer interaction and affective computing. This paper presents an adaptive hybrid SER framework that combines a deep neural feature extraction module with a heterogeneous ensemble of machine learning classifiers, including XGBoost, Support [...] Read more.
Speech Emotion Recognition (SER) is a key enabling technology for advanced human–computer interaction and affective computing. This paper presents an adaptive hybrid SER framework that combines a deep neural feature extraction module with a heterogeneous ensemble of machine learning classifiers, including XGBoost, Support Vector Machines (SVMs), and Random Forest. To overcome the limitations of static fusion strategies, a confidence-gated meta-classification mechanism is introduced to dynamically weight the contribution of each base classifier according to its instance-level reliability. The proposed approach is evaluated on two widely adopted benchmark datasets, EmoDB and SAVEE, achieving competitive accuracies of 98.88% and 91.92%, respectively. Experimental results demonstrate that the proposed fusion strategy significantly improves robustness against inter-speaker variability and emotional ambiguity, while maintaining low computational complexity suitable for real-time implementation. These findings highlight the effectiveness of the proposed framework as a robust and efficient solution for speech emotion recognition. While the model is evaluated on benchmark datasets, it is intended as a foundational component for future emotion-aware systems, including applications in human–computer interaction. Full article
(This article belongs to the Section Bioelectronics)
Show Figures

Figure 1

14 pages, 579 KB  
Article
Wearable Sensor-Free Adult Physical Activity Monitoring Using Smartphone IMU Signals: Cross-Subject Deep Learning with Window-Length and Sensor Modality Studies
by Mussa Turdalyuly, Ay Zholdassova, Tolganay Turdalykyzy and Aydin Doshybekov
Information 2026, 17(4), 368; https://doi.org/10.3390/info17040368 - 14 Apr 2026
Viewed by 1093
Abstract
Human activity recognition (HAR) using inertial sensors is essential for health monitoring and wellness applications, yet robust classification in real-world adult scenarios remains challenging due to subject variability and activity transitions in smartphone sensing environments. This study investigated smartphone-based physical activity recognition using [...] Read more.
Human activity recognition (HAR) using inertial sensors is essential for health monitoring and wellness applications, yet robust classification in real-world adult scenarios remains challenging due to subject variability and activity transitions in smartphone sensing environments. This study investigated smartphone-based physical activity recognition using accelerometer and gyroscope signals under a cross-subject evaluation protocol. To reduce label ambiguity and improve generalization, the original activity set was grouped into a reduced 6-class taxonomy. We evaluated lightweight deep learning models, including a smartphone-only convolutional neural network (CNN) and a multimodal fusion model combining smartphone and smartwatch signals. Using GroupKFold cross-subject validation, the smartphone-only CNN achieved competitive performance with Macro-F1 ≈ 0.46, while multimodal fusion did not provide consistent improvements. We also examined temporal segmentation and showed that shorter windows (2.0 s) yield better results than longer windows. Sensor ablation confirmed the importance of gyroscope information, and per-class analysis indicated that dynamic activities could be recognized reliably, whereas stairs and static categories remained difficult. Overall, the results demonstrate the practicality of smartphone-based activity recognition using built-in smartphone sensors without external wearable devices for adult activity monitoring and provide recommendations for window length and sensor selection in cross-subject HAR. Full article
Show Figures

Figure 1

18 pages, 1642 KB  
Article
Foundation Protein Language Models for Influenza A Virus T-Cell Epitope Prediction: A Transformer-Based Viroinformatics Framework
by Syed Nisar Hussain Bukhari and Kingsley A. Ogudo
Viruses 2026, 18(3), 380; https://doi.org/10.3390/v18030380 - 18 Mar 2026
Viewed by 1347
Abstract
Influenza A virus remains a major cause of respiratory disease worldwide and poses a persistent challenge to vaccine development due to its rapid genetic evolution and antigenic variability. T-cell-based immunity has therefore gained increasing importance, as it can provide broader and more durable [...] Read more.
Influenza A virus remains a major cause of respiratory disease worldwide and poses a persistent challenge to vaccine development due to its rapid genetic evolution and antigenic variability. T-cell-based immunity has therefore gained increasing importance, as it can provide broader and more durable protection by targeting conserved viral regions. Accurate identification of T-cell epitopes (TCEs) is a fundamental requirement for epitope-based vaccine design and immunological research. Although numerous computational methods have been proposed, many existing approaches rely on handcrafted physicochemical features, which offer limited ability to capture contextual sequence dependencies. In this study, a transformer-based viroinformatics framework is proposed for the binary prediction of TCEs from Influenza A virus peptide sequences. The framework employs a pretrained Evolutionary Scale Modeling-2 (ESM-2) protein language model (PLM) to generate rich, contextualized embeddings directly from raw amino acid sequences, eliminating the need for manual feature engineering. These embeddings are processed using a lightweight attention-based transformer classifier to learn epitope-specific sequence patterns. The model achieves strong and stable predictive performance, attaining an accuracy of approximately 97% and an AUC close to 0.99 under stratified cross-validation. Ablation analysis further confirms that protein language model representations and self-attention contribute substantially to performance gains over classical machine learning baselines. To enhance practical reliability, Monte Carlo dropout is incorporated during inference to provide uncertainty-aware predictions, enabling differentiation between high-confidence and ambiguous peptide candidates. In addition, attention-based interpretability is used to identify residue-level contributions to model decisions, offering biologically meaningful insights into epitope recognition. Overall, this study demonstrates that PLMs combined with Transformer architectures provide an effective, interpretable, and a promising computational framework for Influenza A TCE discovery and vaccine research. Full article
(This article belongs to the Special Issue Viroinformatics and Viral Diseases)
Show Figures

Figure 1

21 pages, 3006 KB  
Article
Emotion Recognition from Facial Expressions Considering Individual Differences in Emotional Intelligence
by Yubin Kim, Ayoung Cho, Hyunwoo Lee and Mincheol Whang
Biomimetics 2026, 11(3), 174; https://doi.org/10.3390/biomimetics11030174 - 2 Mar 2026
Viewed by 771
Abstract
Facial expression recognition (FER) in naturalistic settings is constrained by label ambiguity and variability in stimulus–response alignment. Adopting a data-centric perspective, this study examined whether emotional intelligence (EI)-stratified training data influence FER performance by treating EI as a qualitative factor associated with affective [...] Read more.
Facial expression recognition (FER) in naturalistic settings is constrained by label ambiguity and variability in stimulus–response alignment. Adopting a data-centric perspective, this study examined whether emotional intelligence (EI)-stratified training data influence FER performance by treating EI as a qualitative factor associated with affective data consistency. Naturally elicited facial expressions were collected in a controlled emotion induction experiment with subjective arousal and valence ratings. Using response-driven labeling, neutral ratings were retained as indicators of ambiguity. Participants were grouped into High and Low EI based on the alignment between subjective evaluations and outputs from a pretrained affect estimator. Identical binary classifiers for arousal and valence recognition were trained while varying only the training data composition and evaluated across baseline, unambiguous, and ambiguous test sets using independent training repetitions with repetition-level statistical aggregation. EI-stratified training was associated with statistically detectable, context-dependent performance differences: group effects were observed primarily under baseline conditions and, to a lesser extent, under ambiguous conditions, whereas no reliable differences emerged under unambiguous conditions. Pooled discrimination differences were modest, but item-level analyses identified significant differences in classification correctness in specific task–condition combinations. Comparable patterns were observed across alternative backbone architectures. These findings indicate that FER performance in naturalistic contexts is influenced not only by model architecture but also by the statistical structure and internal coherence of the training data, supporting EI-informed data selection in ambiguity-prone scenarios. Full article
Show Figures

Figure 1

22 pages, 979 KB  
Article
Case Series and Literature Narrative Review of Immune-Mediated Thrombotic Thrombocytopenic Purpura in Children
by Letiția-Elena Radu, Andreea Nicoleta Șerbănică, Andra Daniela Marcu, Ana-Maria Bică, Cristina Georgiana Jercan, Radu Obrișcă, Georgiana Gherghe, Gabriela Droc, Dana Tomescu and Anca Coliță
Children 2026, 13(3), 350; https://doi.org/10.3390/children13030350 - 28 Feb 2026
Viewed by 1224
Abstract
Background/Objectives: Immune-mediated thrombotic thrombocytopenic purpura (iTTP) is a rare but life-threatening thrombotic microangiopathy in children. Secondary forms, occurring in association with immune dysregulation, autoimmune disease, or other triggers, are particularly challenging to diagnose and manage, and pediatric-specific data remain limited. This study [...] Read more.
Background/Objectives: Immune-mediated thrombotic thrombocytopenic purpura (iTTP) is a rare but life-threatening thrombotic microangiopathy in children. Secondary forms, occurring in association with immune dysregulation, autoimmune disease, or other triggers, are particularly challenging to diagnose and manage, and pediatric-specific data remain limited. This study aimed to describe the clinical characteristics, diagnostic pathways, and management of pediatric iTTP and to contextualize these findings within the recent literature. Methods: We conducted a retrospective case series of pediatric patients diagnosed with iTTP at a tertiary referral center, between November 2021 and January 2026. Clinical presentation, laboratory findings, including ADAMTS13 activity and ADAMTS13 inhibitors, associated conditions, treatment strategies, and outcomes were reviewed. In parallel, a narrative literature review was performed focusing on pediatric immune-mediated secondary TTP published over the past five years. Results: Four pediatric patients (three females, one male; median age 14 years) met inclusion criteria. All presented with severe thrombocytopenia and microangiopathic hemolytic anemia, accompanied by prominent neurologic manifestations in three cases. Severe ADAMTS13 activity deficiency (≤10%) with positive inhibitors was documented in all patients. Secondary iTTP occurred in association with evolving systemic autoimmunity, systemic lupus erythematosus, common variable immunodeficiency, or without an identifiable trigger at presentation. High clinical probability scores facilitated early diagnosis. Management required plasma exchange, corticosteroids, and targeted and immunomodulatory therapy. Conclusions: Pediatric secondary iTTP is a heterogeneous condition that frequently presents with diagnostic ambiguity and severe neurologic involvement. Early recognition, prompt initiation of TTP-directed therapy, and comprehensive immunologic evaluation are critical for favorable outcomes. Case series combined with narrative reviews remain valuable for advancing understanding and optimizing individualized care in this rare pediatric disorder. Full article
(This article belongs to the Special Issue Advances in the Epidemiology of Hemostasis Disorders in Children)
Show Figures

Figure 1

23 pages, 498 KB  
Review
Recognition and Management of Cognitive Impairment in Chronic Obstructive Pulmonary Disease (COPD): Implications of Clinical Confidence
by Rayan A. Siraj
Medicina 2026, 62(3), 438; https://doi.org/10.3390/medicina62030438 - 26 Feb 2026
Viewed by 930
Abstract
Cognitive impairment is a serious comorbidity in chronic obstructive pulmonary disease (COPD), consistently associated with adverse clinical outcomes, including impaired self-management, poor treatment adherence, reduced participation in pulmonary rehabilitation, and increased risk of mortality. Despite this, it remains inconsistently recognised and insufficiently addressed [...] Read more.
Cognitive impairment is a serious comorbidity in chronic obstructive pulmonary disease (COPD), consistently associated with adverse clinical outcomes, including impaired self-management, poor treatment adherence, reduced participation in pulmonary rehabilitation, and increased risk of mortality. Despite this, it remains inconsistently recognised and insufficiently addressed during routine COPD assessment. This narrative review synthesises current evidence on the recognition and management of cognitive impairment in COPD, with a particular focus on understanding why it continues to be under-recognised and inadequately managed in clinical practice. Across care settings, cognitive concerns are commonly identified informally, assessed selectively, or deferred altogether, even when clinicians acknowledge their relevance to respiratory assessment, treatment implementation, and patient engagement. This persistent evidence–practice gap suggests the influence of factors extending beyond disease- or patient-related explanations alone. Emerging evidence indicates that clinician-level determinants, particularly clinical confidence, play a central role in shaping cognitive care practices. Limited clinical confidence appears to mediate the translation of existing knowledge and competence into clinical action, influencing decisions to initiate assessment, communicate cognitive concerns, assume clinical ownership, and pursue follow-up or referral. These confidence-related barriers are further reinforced by educational limitations, time constraints, diagnostic ambiguity, particularly in the early cognitive impairment stage, and the absence of clear operational guidance within COPD-specific frameworks. Conceptualising cognitive care through the lens of clinical confidence provides a coherent explanation for the underrecognition of cognitive impairment in COPD. It also helps account for observed variability in clinical decision-making, highlighting clinical confidence as a modifiable intermediary between knowledge, competence, and practice and a potential target for strengthening integrated, patient-centred COPD care. Full article
(This article belongs to the Special Issue New Trends in Chronic Obstructive Pulmonary Disease (COPD))
Show Figures

Figure 1

27 pages, 11057 KB  
Article
A Variable-Speed and Multi-Condition Bearing Fault Diagnosis Method Based on Adaptive Signal Decomposition and Deep Feature Fusion
by Ting Li, Mingyang Yu, Tianyi Ma, Yanping Du and Shuihai Dou
Algorithms 2025, 18(12), 753; https://doi.org/10.3390/a18120753 - 28 Nov 2025
Cited by 4 | Viewed by 1124
Abstract
To address the challenges in identifying effective fault features and achieving sufficient diagnostic accuracy and robustness in variable-speed printing press bearings, where complex mixed-condition vibration signals exhibit non-stationarity, strong nonlinearity, ambiguous time-frequency characteristics, and overlapping fault features across multiple operating conditions, this paper [...] Read more.
To address the challenges in identifying effective fault features and achieving sufficient diagnostic accuracy and robustness in variable-speed printing press bearings, where complex mixed-condition vibration signals exhibit non-stationarity, strong nonlinearity, ambiguous time-frequency characteristics, and overlapping fault features across multiple operating conditions, this paper proposes an adaptive optimization signal decomposition method combined with dual-modal time-series and image deep feature fusion for variable-speed multi-condition bearing fault diagnosis. First, to overcome the strong parameter dependency and significant noise interference of traditional adaptive decomposition algorithms, the Crested Porcupine Optimization Algorithm is introduced to adaptively search for the optimal noise amplitude and integration count of ICEEMDAN for effective signal decomposition. IMF components are then screened and reorganized based on correlation coefficients and variance contribution rates to enhance fault-sensitive information. Second, multidimensional time-domain features are extracted in parallel to construct time-frequency images, forming time-sequence-image bimodal inputs that enhance fault representation across different dimensions. Finally, a dual-branch deep learning model is developed: the time-sequence branch employs gated recurrent units to capture feature evolution trends, while the image branch utilizes SE-ResNet18 with embedded channel attention mechanisms to extract deep spatial features. Multimodal feature fusion enables classification recognition. Validation using a bearing self-diagnosis dataset from variable-speed hybrid operation and the publicly available Ottawa variable-speed bearing dataset demonstrates that this method achieves high-accuracy fault identification and strong generalization capabilities across diverse variable-speed hybrid operating conditions. Full article
(This article belongs to the Special Issue Machine Learning Algorithms for Signal Processing)
Show Figures

Figure 1

18 pages, 1126 KB  
Article
The Ambiguous Morpheme Processing in Chinese Compound Word Recognition in Deaf Readers
by Yang Liu, Mengfang Zhang and Yan Wu
Behav. Sci. 2025, 15(12), 1625; https://doi.org/10.3390/bs15121625 - 25 Nov 2025
Viewed by 852
Abstract
This study used event-related potentials (ERPs) to examine how deaf individuals process ambiguous morphemes during Chinese compound word recognition in a masked priming lexical decision paradigm. Ambiguous morphemes were classified as balanced or biased, and two experiments employed a 3 × 2 within-subject [...] Read more.
This study used event-related potentials (ERPs) to examine how deaf individuals process ambiguous morphemes during Chinese compound word recognition in a masked priming lexical decision paradigm. Ambiguous morphemes were classified as balanced or biased, and two experiments employed a 3 × 2 within-subject design. Each morpheme’s two meanings served as both primes and targets. The independent variables were prime type (meaning1 vs. meaning2 vs. unrelated) and target type (meaning1 vs. meaning2), with meaning1 being the dominant meaning and meaning2 being the subordinate meaning for biased morphemes. In the N250 (sublexical processing), balanced morphemes showed a main effect of prime type: any orthographically similar prime elicited priming. In the N400 (semantic processing), an interaction of prime and target type emerged, with only contextually congruent meanings activated. For biased morphemes, interactions were observed across N250 and N400 stages. The dominant meaning was consistently activated: when the target was dominant, both meanings showed priming; when the target was subordinate, only the subordinate meaning produced priming. These results reveal a dissociation in how deaf readers process ambiguous morphemes: balanced morphemes rely on contextual information, whereas biased morphemes are influenced by meaning frequency. The findings provide novel insights into the temporal dynamics of morpheme-based lexical access in deaf Chinese readers, with implications for reading and vocabulary instruction. Full article
(This article belongs to the Section Cognition)
Show Figures

Figure 1

22 pages, 7307 KB  
Article
Unified Spatiotemporal Detection for Isolated Sign Language Recognition Using YOLO-Act
by Nada Alzahrani, Ouiem Bchir and Mohamed Maher Ben Ismail
Electronics 2025, 14(23), 4589; https://doi.org/10.3390/electronics14234589 - 23 Nov 2025
Cited by 1 | Viewed by 1874
Abstract
Isolated Sign Language Recognition (ISLR), which focuses on identifying individual signs from sign language videos, presents substantial challenges due to small and ambiguous hand regions, high visual similarity among signs, and large intra-class variability. This study investigates the adaptability of YOLO-Act, a unified [...] Read more.
Isolated Sign Language Recognition (ISLR), which focuses on identifying individual signs from sign language videos, presents substantial challenges due to small and ambiguous hand regions, high visual similarity among signs, and large intra-class variability. This study investigates the adaptability of YOLO-Act, a unified spatiotemporal detection framework originally developed for generic action recognition in videos, when applied to large-scale sign language benchmarks. YOLO-Act jointly performs signer localization (identifying the person signing within a video) and action classification (determining which sign is performed) directly from RGB sequences, eliminating the need for pose estimation or handcrafted temporal cues. We evaluate the model on the WLASL2000 and MSASL1000 datasets for American Sign Language recognition, achieving Top-1 accuracies of 67.07% and 81.41%, respectively. The latter represents a 3.55% absolute improvement over the best-performing baseline without pose supervision. These results demonstrate the strong cross-domain generalization and robustness of YOLO-Act in complex multi-class recognition scenarios. Full article
Show Figures

Figure 1

18 pages, 1694 KB  
Article
FAIR-Net: A Fuzzy Autoencoder and Interpretable Rule-Based Network for Ancient Chinese Character Recognition
by Yanling Ge, Yunmeng Zhang and Seok-Beom Roh
Sensors 2025, 25(18), 5928; https://doi.org/10.3390/s25185928 - 22 Sep 2025
Cited by 1 | Viewed by 1274
Abstract
Ancient Chinese scripts—including oracle bone carvings, bronze inscriptions, stone steles, Dunhuang scrolls, and bamboo slips—are rich in historical value but often degraded due to centuries of erosion, damage, and stylistic variability. These issues severely hinder manual transcription and render conventional OCR techniques inadequate, [...] Read more.
Ancient Chinese scripts—including oracle bone carvings, bronze inscriptions, stone steles, Dunhuang scrolls, and bamboo slips—are rich in historical value but often degraded due to centuries of erosion, damage, and stylistic variability. These issues severely hinder manual transcription and render conventional OCR techniques inadequate, as they are typically trained on modern printed or handwritten text and lack interpretability. To tackle these challenges, we propose FAIR-Net, a hybrid architecture that combines the unsupervised feature learning capacity of a deep autoencoder with the semantic transparency of a fuzzy rule-based classifier. In FAIR-Net, the deep autoencoder first compresses high-resolution character images into low-dimensional, noise-robust embeddings. These embeddings are then passed into a Fuzzy Neural Network (FNN), whose hidden layer leverages Fuzzy C-Means (FCM) clustering to model soft membership degrees and generate human-readable fuzzy rules. The output layer uses Iteratively Reweighted Least Squares Estimation (IRLSE) combined with a Softmax function to produce probabilistic predictions, with all weights constrained as linear mappings to maintain model transparency. We evaluate FAIR-Net on CASIA-HWDB1.0, HWDB1.1, and ICDAR 2013 CompetitionDB, where it achieves a recognition accuracy of 97.91%, significantly outperforming baseline CNNs (p < 0.01, Cohen’s d > 0.8) while maintaining the tightest confidence interval (96.88–98.94%) and lowest standard deviation (±1.03%). Additionally, FAIR-Net reduces inference time to 25 s, improving processing efficiency by 41.9% over AlexNet and up to 98.9% over CNN-Fujitsu, while preserving >97.5% accuracy across evaluations. To further assess generalization to historical scripts, FAIR-Net was tested on the Ancient Chinese Character Dataset (9233 classes; 979,907 images), achieving 83.25% accuracy—slightly higher than ResNet101 but 2.49% lower than SwinT-v2-small—while reducing training time by over 5.5× compared to transformer-based baselines. Fuzzy rule visualization confirms enhanced robustness to glyph ambiguities and erosion. Overall, FAIR-Net provides a practical, interpretable, and highly efficient solution for the digitization and preservation of ancient Chinese character corpora. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

24 pages, 2115 KB  
Article
MHD-Protonet: Margin-Aware Hard Example Mining for SAR Few-Shot Learning via Dual-Loss Optimization
by Marii Zayani, Abdelmalek Toumi and Ali Khalfallah
Algorithms 2025, 18(8), 519; https://doi.org/10.3390/a18080519 - 16 Aug 2025
Cited by 1 | Viewed by 2009
Abstract
Synthetic aperture radar (SAR) image classification under limited data conditions faces two major challenges: inter-class similarity, where distinct radar targets (e.g., tanks and armored trucks) have nearly identical scattering characteristics, and intra-class variability, caused by speckle noise, pose changes, and differences in depression [...] Read more.
Synthetic aperture radar (SAR) image classification under limited data conditions faces two major challenges: inter-class similarity, where distinct radar targets (e.g., tanks and armored trucks) have nearly identical scattering characteristics, and intra-class variability, caused by speckle noise, pose changes, and differences in depression angle. To address these challenges, we propose MHD-ProtoNet, a meta-learning framework that extends prototypical networks with two key innovations: margin-aware hard example mining to better separate confusable classes by enforcing prototype distance margins, and dual-loss optimization to refine embeddings and improve robustness to noise-induced variations. Evaluated on the MSTAR dataset in a five-way one-shot task, MHD-ProtoNet achieves 76.80% accuracy, outperforming the Hybrid Inference Network (HIN) (74.70%), as well as standard few-shot methods such as prototypical networks (69.38%), ST-PN (72.54%), and graph-based models like ADMM-GCN (61.79%) and DGP-NET (68.60%). By explicitly mitigating inter-class ambiguity and intra-class noise, the proposed model enables robust SAR target recognition with minimal labeled data. Full article
Show Figures

Figure 1

18 pages, 4979 KB  
Systematic Review
Discordant High-Gradient Aortic Stenosis: A Systematic Review
by Nadera N. Bismee, Mohammed Tiseer Abbas, Hesham Sheashaa, Fatmaelzahraa E. Abdelfattah, Juan M. Farina, Kamal Awad, Isabel G. Scalia, Milagros Pereyra Pietri, Nima Baba Ali, Sogol Attaripour Esfahani, Omar H. Ibrahim, Steven J. Lester, Said Alsidawi, Chadi Ayoub and Reza Arsanjani
J. Cardiovasc. Dev. Dis. 2025, 12(7), 255; https://doi.org/10.3390/jcdd12070255 - 3 Jul 2025
Cited by 2 | Viewed by 3169
Abstract
Aortic stenosis (AS), the most common valvular heart disease, is traditionally graded based on several echocardiographic quantitative parameters, such as aortic valve area (AVA), mean pressure gradient (MPG), and peak jet velocity (Vmax). This systematic review evaluates the clinical significance and prognostic implications [...] Read more.
Aortic stenosis (AS), the most common valvular heart disease, is traditionally graded based on several echocardiographic quantitative parameters, such as aortic valve area (AVA), mean pressure gradient (MPG), and peak jet velocity (Vmax). This systematic review evaluates the clinical significance and prognostic implications of discordant high-gradient AS (DHG-AS), a distinct hemodynamic phenotype characterized by elevated MPG despite a preserved AVA (>1.0 cm2). Although often overlooked, DHG-AS presents unique diagnostic and therapeutic challenges, as high gradients remain a strong predictor of adverse outcomes despite moderately reduced AVA. Sixty-three studies were included following rigorous selection and quality assessment of the key studies. Prognostic outcomes across five key studies were discrepant: some showed better survival in DHG-AS compared to concordant high-gradient AS (CHG-AS), while others reported similar or worse outcomes. For instance, a retrospective observational study including 3209 patients with AS found higher mortality in CHG-AS (unadjusted HR: 1.4; 95% CI: 1.1 to 1.7), whereas another retrospective multicenter study including 2724 patients with AS observed worse outcomes in DHG-AS (adjusted HR: 1.59; 95% CI: 1.04 to 2.56). These discrepancies may stem from delays in intervention or heterogeneity in study populations. Despite the diagnostic ambiguity, the presence of high gradients warrants careful evaluation, aggressive risk stratification, and timely management. Current guidelines recommend a multimodal approach combining echocardiography, computed tomography (CT) calcium scoring, transesophageal echocardiography (TEE) planimetry, and, when needed, catheterization. Anatomic AVA assessment by TEE, CT, and cardiac magnetic resonance imaging (CMR) can improve diagnostic accuracy by directly visualizing valve morphology and planimetry-based AVA, helping to clarify the true severity in discordant cases. However, these modalities are limited by factors such as image quality (especially with TEE), radiation exposure and contrast use (in CT), and availability or contraindications (in CMR). Management remains largely based on CHG-AS protocols, with intervention primarily guided by transvalvular gradient and symptom burden. The variability among the different guidelines in defining severity and therapeutic thresholds highlights the need for tailored approaches in DHG-AS. DHG-AS is clinically relevant and associated with substantial prognostic uncertainty. Timely recognition and individualized treatment could improve outcomes in this complex subgroup. Full article
(This article belongs to the Special Issue Cardiovascular Imaging in Heart Failure and in Valvular Heart Disease)
Show Figures

Figure 1

Back to TopTop