Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (589)

Search Parameters:
Keywords = categorical data clustering

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
23 pages, 696 KB  
Article
A Domain-Guided Feature-Fusion Framework for Ship Equipment Based on Multi-Type Features
by Ruoyi Yin, Yali Zhai, Zhengxuan Gu and Songshi Shao
Algorithms 2026, 19(9), 798; https://doi.org/10.3390/a19090798 - 17 Sep 2026
Abstract
Reliability assessment of ship equipment is often constrained by insufficient or unavailable failure data, particularly for highly reliable components with extremely low failure frequencies. To support the use of reference information from similar equipment in subsequent reliability analysis, this study proposes a domain-guided [...] Read more.
Reliability assessment of ship equipment is often constrained by insufficient or unavailable failure data, particularly for highly reliable components with extremely low failure frequencies. To support the use of reference information from similar equipment in subsequent reliability analysis, this study proposes a domain-guided multi-type feature-fusion framework for ship equipment clustering. The proposed framework addresses the heterogeneous nature of ship equipment records by integrating textual, categorical, and numerical attributes into a unified representation. Specifically, equipment names and specification/model information are represented using character-level TF-IDF features; categorical attributes are encoded through One-Hot representation, and numerical attributes are processed using logarithmic transformation and standardization. Six engineering attributes, including equipment name, specification/model information, technical category, measurement unit, number of installations per platform, and reference unit price, are incorporated into the fused feature space. In addition, a domain-knowledge-driven feature-group weighting mechanism is introduced to emphasize attributes that directly reflect functional and technical similarities, especially equipment names and specification/model information. K-Means clustering is then performed in the weighted fused feature space, and the number of clusters is determined by jointly considering the Silhouette Coefficient, Calinski–Harabasz index, Davies–Bouldin index, and cluster-size distribution. Experiments on 1345 practical ship equipment samples show that K = 36 provides a reasonable balance among clustering structure, candidate-set availability, engineering consistency, and stability under random initialization. Compared with the equal-weight scheme, Attribute-Weighted K-Means, K-Prototypes, and Gower-based hierarchical clustering, the proposed framework achieves higher equipment-name and specification/model similarities while maintaining meaningful technical-category consistency. These results indicate that the proposed framework can effectively identify latent functional and technical similarity relationships among ship equipment and provide candidate reference sets for reliability assessment under sparse failure data conditions. Full article
Show Figures

Figure 1

23 pages, 3294 KB  
Review
Physical AI: A Data-Driven Survey of Foundations, Technologies, and Applications
by Johannes Stübinger and Fabio Metz
Technologies 2026, 14(9), 588; https://doi.org/10.3390/technologies14090588 - 17 Sep 2026
Abstract
This paper presents a systematic, data-driven literature review of research on Physical Artificial Intelligence (AI) based on the top 100 Google Scholar publications related to the search terms “Physical Artificial Intelligence” and “Physical AI”. The rapid advancement of Physical AI, driven by the [...] Read more.
This paper presents a systematic, data-driven literature review of research on Physical Artificial Intelligence (AI) based on the top 100 Google Scholar publications related to the search terms “Physical Artificial Intelligence” and “Physical AI”. The rapid advancement of Physical AI, driven by the convergence of advanced sensor technologies and foundation world models, has resulted in a diverse and fragmented research landscape that lacks comprehensive quantitative overviews. To address this gap, we implement and apply an AI-assisted computational analysis pipeline to this domain. The collected publications are processed using a Large Language Model accessed via a Python-based Application Programming Interface (API), enabling a structured computational analysis of the literature to assist thematic categorization. Based on this approach, the publications are grouped into five data-driven thematic clusters reflecting primary research perspectives within the analyzed sample. Specifically, the identified clusters comprise “Sensor Infrastructure and Architectures”, “Core Learning and Modeling Methodologies”, “Sim-to-Real and Digital Twins”, “Applications”, and “Safety, Governance, and Ethics”. By synthesizing the literature in a structured manner, this work provides a consolidated overview of central research patterns, identifies key operational challenges, and highlights fragmentation across Physical AI research, establishing a solid foundation for future trustworthy autonomous systems. Full article
Show Figures

Figure 1

26 pages, 116305 KB  
Article
Annual Gridded Anthropogenic CH4 Emissions Estimation in China (2019–2025) Integrating Multisource Data: SHAP-Based Driver Attribution and Spatio-Temporal Patterns
by Chaokang He, Qinjun Wang and Wenyue Xie
Remote Sens. 2026, 18(18), 3168; https://doi.org/10.3390/rs18183168 - 15 Sep 2026
Abstract
Accurately quantifying the spatiotemporal dynamics and driving mechanisms of anthropogenic methane (CH4) emissions (MEs) is of great significance for achieving regional “dual-carbon” goals and global climate collaborative governance. However, existing ME inventories and macro-inversion models generally face bottlenecks such as coarse [...] Read more.
Accurately quantifying the spatiotemporal dynamics and driving mechanisms of anthropogenic methane (CH4) emissions (MEs) is of great significance for achieving regional “dual-carbon” goals and global climate collaborative governance. However, existing ME inventories and macro-inversion models generally face bottlenecks such as coarse spatial resolution, lack of data update timeliness, and the inability of traditional static emission factors to capture non-linear responses. To address these issues, this study proposes an annual ME inventory enhancement framework integrating multi-source geographic and remote sensing data. This framework evaluates four advanced machine learning (ML) algorithms, including Random Forest (RF), Categorical Boosting (CB), Extreme Gradient Boosting (XGB), and Light Gradient Boosting Machine (LGBM), to construct a 0.1° high-resolution spatial grid of anthropogenic ME in China from 2019 to 2025. Furthermore, it introduces the SHapley Additive exPlanations (SHAP) framework and multi-scale spatial autocorrelation analysis to parse the driving mechanisms and clustering patterns. The results show the following: (1) LGBM exhibits the optimal comprehensive estimation accuracy (R2 = 0.938, RMSE = 3.707 Kt) and robust capability in capturing extreme ME sources (RTop2 = 0.929). (2) SHAP attribution reveals that coal mining and nighttime light (NTL) represent the primary contributing features to ME predictions (with a cumulative contribution of 65.60%), followed by agricultural and pastoral activities (24.98%), and all factors exhibit significant non-linear threshold and step-response characteristics. (3) Regarding spatiotemporal evolution, China’s total anthropogenic ME shows a trend of initial slow increase followed by high-level stabilization; spatially, it presents a “hot in the north, cold in the south” pattern, with extreme high values highly clustered in the Shanxi–Shaanxi–Inner Mongolia energy triangle and its peripheral expansion nodes. This study provides scientific references for formulating tailored, multi-scale, and refined CH4 mitigation strategies. Full article
(This article belongs to the Special Issue Satellite Remote Sensing of Quantifying Greenhouse Gases Emissions)
Show Figures

Figure 1

32 pages, 44743 KB  
Article
Exploiting Projection Trajectories Discrepancy for Multipath Suppression and Detailed Feature Extraction of Buildings in SAR Adjacent Sub-Aperture Images
by Yi Zhang, Daoxiang An, Di Wang, Jinxing Li and Leping Chen
Remote Sens. 2026, 18(18), 3164; https://doi.org/10.3390/rs18183164 - 15 Sep 2026
Abstract
In synthetic aperture radar (SAR) imagery of built-up areas, multipath effects generate false targets that closely resemble genuine structural features, severely hindering refined interpretation of building structures. Existing methods based on interferometric SAR, tomographic SAR, or full-angle circular SAR (CSAR), while effective in [...] Read more.
In synthetic aperture radar (SAR) imagery of built-up areas, multipath effects generate false targets that closely resemble genuine structural features, severely hindering refined interpretation of building structures. Existing methods based on interferometric SAR, tomographic SAR, or full-angle circular SAR (CSAR), while effective in 3D information extraction, impose stringent requirements on radar systems, data acquisition conditions, and prior information, rendering them less applicable to time-critical scenarios with limited observation constraints. To address this issue, this paper proposes a multipath suppression and detailed feature extraction method for buildings based on projection offset discrepancies across adjacent sub-aperture images. First, a projection offset model for elevated target points and a multipath effect model between elevated targets are established, theoretically revealing that the projections of elevated targets and multipath ghosts are offset to opposite sides of the target in successive sub-aperture images. Building upon this theoretical foundation, a complete image-domain processing pipeline is developed: an improved iterative watershed algorithm for robust building region segmentation, non-edge Hough transform combined with Thresholded Connected Component Analysis clustering for wall line extraction, multi-dimensional feature-based Hungarian algorithm for wall matching and tracking across sub-apertures, and normalized cross-correlation (NCC) for pixel-level offset estimation. Based on the distinct offset characteristics, building structures are categorized into three classes—stationary walls, elevated structures, and multipath ghosts—enabling simultaneous multipath suppression and structural extraction. Experimental results on Ku-band UAV-borne circular SAR data demonstrate that the proposed method requires only a small number of sub-aperture images with narrow angular spans to effectively distinguish different scattering structures, suppress multipath ghosts, and extract major structural details, providing a viable solution for building interpretation in SAR imagery under observation-constrained scenarios. Full article
(This article belongs to the Special Issue Physics-Informed Information Exploitation in Radar Remote Sensing)
Show Figures

Figure 1

38 pages, 572 KB  
Article
Statistical Methods for Assessing Diagnostic Agreement
by Maximilian Pilz
Appl. Sci. 2026, 16(17), 8808; https://doi.org/10.3390/app16178808 - 4 Sep 2026
Viewed by 201
Abstract
With the rise of artificial intelligence (AI), an increasing number of AI-based diagnostic tools are being developed. Before clinical implementation, these tools must be validated against existing gold standards. This requires trials that quantify the agreement between AI predictions and reference measurements. However, [...] Read more.
With the rise of artificial intelligence (AI), an increasing number of AI-based diagnostic tools are being developed. Before clinical implementation, these tools must be validated against existing gold standards. This requires trials that quantify the agreement between AI predictions and reference measurements. However, designing such agreement studies poses methodological challenges that differ substantially from classical superiority trials. This paper aims to provide statistical methods for assessing diagnostic agreement. Methods were categorized according to the measurement scale of the data (nominal, ordinal, continuous)—with a separate group for methods that apply across several scales—and according to the number of raters or measurements involved. A decision tree is provided as a simplified educational framework for method selection rather than as a general method-selection algorithm: design features such as repeated measurements, clustering, spectrum effects, dependence between raters, and an imperfect reference method are not encoded in it and are discussed separately, together with the circularity and confounding issues specific to the validation of AI-based tools. For each method, we summarized assumptions, appropriate use cases, interpretation of results, and available open-source software for sample size calculation and analysis, and we illustrate the sample size calculations in three fully worked examples covering binary, ordinal, and continuous outcomes. We further distinguish conditional inference about one fixed, frozen model version from the broader generalization to a class of algorithms or to future model versions, which require additional sources of algorithmic and dataset variability to be represented in the design and analysis. We conclude by outlining open methodological questions—including Bayesian approaches to agreement estimation, methods for complex AI outputs, agreement models for clustered and repeated-measures designs, and the limited software support for Gwet’s AC1/AC2 sample size planning—that warrant further work as diagnostic technologies and statistical methodology continue to evolve. Full article
(This article belongs to the Special Issue Statistics in Data Science: Latest Methods and Applications)
Show Figures

Figure 1

21 pages, 290 KB  
Article
Future Drivers of Electronic Auditing Under Electronic Governance: A Delphi Study from Iraq
by Ahmed Sameer Abdulhussein Dakheel, Alireza Rahrovi Dastjerdi and Amin Rostami
J. Risk Financ. Manag. 2026, 19(9), 655; https://doi.org/10.3390/jrfm19090655 - 1 Sep 2026
Viewed by 191
Abstract
The rapid digitalization of public administration has transformed auditing environments, especially in emerging economies expanding their electronic governance (e-governance) frameworks. This study identifies and prioritizes the key drivers influencing electronic auditing (e-auditing) development in Iraq over the next decade. Using a mixed-methods design, [...] Read more.
The rapid digitalization of public administration has transformed auditing environments, especially in emerging economies expanding their electronic governance (e-governance) frameworks. This study identifies and prioritizes the key drivers influencing electronic auditing (e-auditing) development in Iraq over the next decade. Using a mixed-methods design, the research first identifies potential drivers through qualitative interviews and open-ended questionnaires. These drivers were then evaluated and prioritized via a two-round Delphi survey involving a purposive panel of 20 experts, including senior auditors, accounting academics, and IT-audit specialists with over 15 years of professional experience. The analysis identified 19 significant drivers categorized into four clusters: (1) emerging audit technologies, (2) information security and data quality, (3) e-governance and transparency, and (4) professional capabilities. Results highlight that technological innovations, specifically real-time monitoring and machine learning, are the most influential drivers. Furthermore, cybersecurity and transparent governance mechanisms are identified as essential pillars for digital auditing in the Iraqi context. By providing a foresight perspective in a post-conflict, emerging economy, this study offers a unique conceptual framework that integrates e-governance maturity with auditing evolution. The findings provide actionable insights for policymakers and regulatory bodies to modernize auditing practices in high-uncertainty environments. Full article
(This article belongs to the Section Business and Entrepreneurship)
30 pages, 19375 KB  
Article
Assessing the Spatial Heterogeneity of Village Homestead Improvement Potential: Village Environment and Multidestination Urban Housing Purchases
by Cheng-Xiang Wang, Chi Chen, Sai-Zu Wang and Wei-Ling Hsu
Buildings 2026, 16(17), 3381; https://doi.org/10.3390/buildings16173381 - 25 Aug 2026
Viewed by 303
Abstract
Understanding the interactions between rural household decision-making behavior and the environment is critical to promoting sustainable rural development. Analyzing the interrelationship between multidestination urban housing purchases, homestead improvement potential, and village environments can provide a theoretical basis for formulating rural revitalization policies. This [...] Read more.
Understanding the interactions between rural household decision-making behavior and the environment is critical to promoting sustainable rural development. Analyzing the interrelationship between multidestination urban housing purchases, homestead improvement potential, and village environments can provide a theoretical basis for formulating rural revitalization policies. This study uses full-sample household survey data, applying multiscale geographically weighted regression to assess spatial variability and K-means clustering to categorize effects. The findings reveal significant spatial heterogeneity in village homestead improvement potential, most strongly associated with locational attributes, water network density, residential quality, and social factors. The association between multidestination urban housing purchases and improvement potential varies by destination: township purchases are positively associated with it, whereas purchases in county centers and beyond show negative associations. The spatial variability of different factors is significant, allowing villages to be classified into five zones based on dominant factors. Tailored, zone-specific policy directions—developing a county-level dual-core urban system, promoting rural tourism, and encouraging concentrated local habitation—are proposed as testable hypotheses for supporting in situ urbanization and homestead land improvement. Full article
(This article belongs to the Special Issue Research on Health, Wellbeing, and Urban Design—2nd Edition)
Show Figures

Figure 1

21 pages, 3154 KB  
Article
Tinnitus Phenotypes and Their Response to the Enriched Acoustic Environment (EAE) Treatment
by Charles Bouzou, Marta Fernández-Ledesma, María Cuesta and Pedro Cobo
Brain Sci. 2026, 16(8), 894; https://doi.org/10.3390/brainsci16080894 - 21 Aug 2026
Viewed by 449
Abstract
Background/Objectives: The aim of this study was to apply data-driven clustering techniques for the subtyping of tinnitus severity to a retrospective cross-sectional cohort of 564 subjects. Additionally, the response of the resulting phenotypes to an enriched acoustic environment (EAE) treatment was investigated. [...] Read more.
Background/Objectives: The aim of this study was to apply data-driven clustering techniques for the subtyping of tinnitus severity to a retrospective cross-sectional cohort of 564 subjects. Additionally, the response of the resulting phenotypes to an enriched acoustic environment (EAE) treatment was investigated. Methods: A hierarchical decision cascade was applied to classify the measured audiograms of participants into six hearing loss (HL) patterns. K-clustering techniques were applied to demographic (gender), hearing (average audiometric threshold, HL pattern), tinnitus (tinnitus onset age, duration, severity, lateralisation, type of sound, aetiology) and emotional variables. Subjects received customised EAE stimulus to be listened to for one hour daily over four months. Results: Five phenotypes were found by applying K-clustering techniques to numerical and categorical variables. EAE provided an average tinnitus severity reduction of 24.3 points, which was clinically relevant and statistically significant, in 78% of subjects who completed the four-month treatment. Conclusions: Analysis of the performance of EAE treatment in the different phenotypes afforded clear differences in the composition of excluded, dropped out, and not improved subjects. Absolute mean improvement of both severity and emotional scores was larger in phenotypes PH3 and PH4, as they had higher tinnitus distress and emotional symptoms at baseline. Nevertheless, relative change of severity score was notably smaller in PH4. Full article
Show Figures

Figure 1

22 pages, 1985 KB  
Article
A Semantic Clustering Framework for Discovering Latent Offense Patterns: A Case Study of Thai Police Records
by Krittakom Srijiranon, Tanatorn Tanantong, Nattanon Keeratiwattapong, Nawarerk Chalarak and Usanut Sangtongdee
Digital 2026, 6(3), 69; https://doi.org/10.3390/digital6030069 - 18 Aug 2026
Viewed by 351
Abstract
Crime offense descriptions are often recorded as unstructured text, making large-scale analysis and categorization difficult. This study proposes a semantic clustering framework for Thai crime offense descriptions using sentence embeddings, dimensionality reduction, and unsupervised clustering. Two datasets were obtained from Thonglor Metropolitan Police [...] Read more.
Crime offense descriptions are often recorded as unstructured text, making large-scale analysis and categorization difficult. This study proposes a semantic clustering framework for Thai crime offense descriptions using sentence embeddings, dimensionality reduction, and unsupervised clustering. Two datasets were obtained from Thonglor Metropolitan Police Station and Mueang Nonthaburi Police Station, Thailand. After preprocessing, the datasets contained 962 and 902 unique offense descriptions, respectively. Each description was transformed into a 768-dimensional embedding using SimCSE-PhayaThaiBERT. The embeddings were represented in Principal Component Analysis (PCA) Space and Uniform Manifold Approximation and Projection (UMAP) Space and clustered using K-Means, DBSCAN, HDBSCAN, and OPTICS. The results showed that UMAP Space generally provided more useful clustering results than PCA Space. Although DBSCAN achieved the highest internal clustering scores, it classified most records as noise. In contrast, HDBSCAN provided a more balanced result by maintaining strong clustering quality while retaining more records for interpretation. Qualitative analysis showed that the discovered clusters corresponded to meaningful offense categories. The proposed framework can support exploratory analysis of Thai crime records without requiring manually labeled data. Full article
Show Figures

Figure 1

20 pages, 713 KB  
Article
How Consumer Engagement Shapes Corporate Technology for Good: The Mediation of Knowledge Co-Creation
by Mengmeng Meng, Qing Li, Yan Huang, Qiaohua Li and Jiasu Lei
Systems 2026, 14(8), 982; https://doi.org/10.3390/systems14080982 - 13 Aug 2026
Viewed by 304
Abstract
How does consumer engagement shape a firm’s technology adoption strategy? Based on the knowledge-based perspective, this paper explores the impact mechanism of consumer engagement on corporate technology for good using survey data from medium–high R&D intensity manufacturing firms within China’s industrial clusters. The [...] Read more.
How does consumer engagement shape a firm’s technology adoption strategy? Based on the knowledge-based perspective, this paper explores the impact mechanism of consumer engagement on corporate technology for good using survey data from medium–high R&D intensity manufacturing firms within China’s industrial clusters. The research sample covers five industries with medium-to-high R&D intensity, categorized according to the OECD classification. The findings suggest that consumer engagement positively affects knowledge co-creation and corporate technology for good, and knowledge co-creation plays a mediating role in the relationship between consumer engagement and corporate technology for good. Further analysis reveals that knowledge absorption ability positively moderates the relationship between consumer engagement and knowledge co-creation, and the mediating effect of knowledge co-creation on the relationship between consumer engagement and corporate technology for good is positively moderated by knowledge absorption ability. The study expands the research on the driving factors of corporate technology for good from the stakeholder theory perspective, providing insights for firms to facilitate consumer value co-creation. Full article
(This article belongs to the Section Systems Practice in Social Science)
Show Figures

Figure 1

17 pages, 2493 KB  
Article
Establishment of Standard Models Using Copula-Based Data Augmentation and Genetic Algorithms for Improving the Energy Performance of Small-Scale Aging Buildings
by Shin Kim, Joung-Joo Choi, Yong-Joon Jun and Kyung-Soon Park
Buildings 2026, 16(15), 3030; https://doi.org/10.3390/buildings16153030 - 30 Jul 2026
Viewed by 296
Abstract
Simulation-dependent energy analysis has long dominated building retrofit research, yet this paradigm presents substantial barriers for non-expert building owners who lack technical software proficiency and detailed building documentation-a challenge compounded by the “curse of dimensionality” when multivariate analysis requires thousands of samples beyond [...] Read more.
Simulation-dependent energy analysis has long dominated building retrofit research, yet this paradigm presents substantial barriers for non-expert building owners who lack technical software proficiency and detailed building documentation-a challenge compounded by the “curse of dimensionality” when multivariate analysis requires thousands of samples beyond available empirical records. Leveraging retrofit data accumulated through Korea’s Green Remodeling programs since 2017, this study proposes a Copula-Genetic Algorithm (Copula-GA) integrated framework that enables rational retrofit decision-making with minimal user inputs (construction year, floor area, structural type). From 178 documented retrofit cases, Gaussian copula-based multivariate sampling generated 10,000 synthetic records while preserving inter-variable dependency structures. Building physics constraints addressing vintage-thermal performance and capacity-efficiency relationships filtered implausible combinations, yielding 9898 valid cases with correlation matrix fidelity confirmed by a Frobenius norm deviation of 0.043. Evolutionary clustering employing a composite fitness function of Silhouette coefficient (0.68) and Davies-Bouldin Index (0.52) identified K = 16 as the optimal partition, categorizing outcomes into four reference model archetypes: Lightweight Structure (Type A, 27.0% reduction, 15.7-year payback), Masonry Structure (Type B, 29.0%, 14.8 years), RC Structure (Type C, 30.7%, 13.4 years), and Mixed Structure (Type D, 30.9%, 13.1 years). The proposed Copula-GA framework bridges the gap between advanced energy optimization methodologies and practical accessibility for non-expert building owners. By transforming limited empirical samples into reliable reference models, this research supports building-sector decarbonization. Using three minimal inputs, a building can be matched to one of the 16 standard models to obtain its expected saving rate, payback period, and recommended measures without detailed simulation. Full article
(This article belongs to the Section Building Energy, Physics, Environment, and Systems)
Show Figures

Figure 1

19 pages, 15974 KB  
Article
Classification Evolution and Epitope Prediction of the Porcine Epidemic Diarrhea Virus (PEDV) Spike Protein in Thailand (2008–2024): Updated Insights for Preventive Strategies
by Christopher James Stott, Tanakamol Mahawan, Pablo Piñeyro, Hongyao Lin, Angkana Tantituvanont and Dachrit Nilubol
Animals 2026, 16(15), 2314; https://doi.org/10.3390/ani16152314 - 27 Jul 2026
Viewed by 745
Abstract
This study analyzed Porcine epidemic diarrhea virus (PEDV) spike protein sequences and structures in Thailand from 2008 to 2024 to provide predicted structural templates that could inform regional vaccine selection and planned exposure frameworks. Using an in-silico approach, the researchers reduced sequence redundancy [...] Read more.
This study analyzed Porcine epidemic diarrhea virus (PEDV) spike protein sequences and structures in Thailand from 2008 to 2024 to provide predicted structural templates that could inform regional vaccine selection and planned exposure frameworks. Using an in-silico approach, the researchers reduced sequence redundancy via CD-HIT (v4.8.1), established evolutionary lineages with BEAST (v1.10.4), and reconstructed protein structures using SWISS-MODEL. Structural comparisons and clustering were performed using DALI Z-scores and DBSCAN (v1.2.2), while Discotope 3 (v3.0) and ElliPro mapped B-cell epitope landscapes against a G1 reference strain. The results revealed a major lineage shift from G2a to G2b strains around 2017, with the spike proteins categorized into 14 subtypes and 6 eigenvalue clusters. Notably, minor amino acid substitutions altered properties such as hydrophobicity without disrupting the core structure, and certain deletions caused minimal structural deviations, indicating that sequence data or predicted structures alone do not fully dictate viral virulence or immunogenicity. Furthermore, primitive TH2 strains shared evolutionary links with G1 or US-InDel strains despite their G2 classification, identifying Cluster 1 as a potential ancestral structural type. In conclusion, this updated analysis provides crucial baseline data to optimize regional PEDV preventative measures, though further rigorous structural investigations are needed to definitively link specific spike alterations to virulence and host immune response. Full article
Show Figures

Graphical abstract

16 pages, 1466 KB  
Article
Associations Among Piglet Characteristics, Farrowing Kinetics, and Stillbirth Risk in Hyperprolific Sows
by Chananchida Mueansree, Phubet Satsook, Napatsorn Kamnaray, Nitikan Nikhomjit, Thanakorn Jongprasert, Niratchaporn Faksongsakul, Suppachok Taveekaikun and Nitipong Homwong
Animals 2026, 16(14), 2223; https://doi.org/10.3390/ani16142223 - 17 Jul 2026
Viewed by 387
Abstract
Stillbirth remains a major cause of piglet loss in hyperprolific sow herds. This observational study investigated the effects of farrowing kinetics, including birth weight (BW), birth order (BO), and expulsion interval (EI), on stillbirth risk. Data from 117 crossbred sows between Landrace sires [...] Read more.
Stillbirth remains a major cause of piglet loss in hyperprolific sow herds. This observational study investigated the effects of farrowing kinetics, including birth weight (BW), birth order (BO), and expulsion interval (EI), on stillbirth risk. Data from 117 crossbred sows between Landrace sires and Yorkshire dams and 1726 piglets were analyzed using linear mixed models (LMM) and generalized linear mixed models (GLMM), with sow included as a random effect to account for litter clustering. BW, BO, time of farrowing (TF), and EI were categorized into three (BWG, low, medium, high), four (BOQ, Q1–Q4), four (H1–H4), and seven (≤5, 6–10, 11–15, 16–20, 21–30, 31–45, and >45 min) groups, respectively. BWG were not associated with BO (p = 0.22) nor with EI (p = 0.49). Univariate GLMM analysis showed that stillbirth risk increased across BOQ, with Q4 piglets exhibiting a higher stillbirth rate than Q1 piglets (9.1% vs. 1.6%; p < 0.01), whereas TF and EI groups were not associated with stillbirth risk (p > 0.05). In the multivariable GLMM, lower BW (OR = 0.49 per 0.25 kg increase; p < 0.001), longer cumulative expulsion interval (CEI) (OR = 1.71 per 60 min increase; p < 0.001), and greater within-litter birth weight variation (BWCV; OR = 1.51 per 10% increase; p = 0.077) were associated with an increased risk of stillbirth. In contrast, BOQ, EI group, litter size (LS), and TF were not retained in the final model. Reverse Kaplan–Meier survival analysis demonstrated a progressive ascend in the probability of being stillborn throughout the farrowing process (p < 0.01), with a median stillborn time of 450 min. These findings indicate that prolonged cumulative farrowing duration, rather than individual birth intervals or birth sequence alone, is the primary farrowing-related determinant of stillbirth risk. These findings suggest that reducing CEI, providing timely assistance to late-born piglets and lowering BWCV may help decrease stillbirth losses in hyperprolific sow herds. Full article
(This article belongs to the Special Issue Strategies to Improve Piglet Survival and Sow Longevity)
Show Figures

Figure 1

81 pages, 989 KB  
Review
Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods
by Behnam Yousefimehr, Mehdi Ghatee, Javad Fazli, Shervin Ghaffari, Zahra Rafei, Mohammad Amin Seifi, Sajed Tavakoli, Abolfazl Nikahd, Mahdi Razi Gandomani, Alireza Orouji, Ramtin Mahmoudi Kashani, Sarina Heshmati and Negin Sadat Mousavi
Mach. Learn. Knowl. Extr. 2026, 8(7), 211; https://doi.org/10.3390/make8070211 - 16 Jul 2026
Cited by 1 | Viewed by 1115
Abstract
Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing predictions toward the majority class and degrading classifier performance. This paper provides a comprehensive, systematic review of data balancing methods, extending beyond foundational oversampling techniques such [...] Read more.
Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing predictions toward the majority class and degrading classifier performance. This paper provides a comprehensive, systematic review of data balancing methods, extending beyond foundational oversampling techniques such as the Synthetic Minority Oversampling Technique (SMOTE) and its variants (e.g., Borderline SMOTE, K-Means SMOTE, and Safe-Level SMOTE) to encompass advanced adaptive methods (MWMOTE, AMDO), deep generative models (generative adversarial networks, variational autoencoders, and diffusion models), undersampling techniques (NearMiss, Tomek Links), combination/hybrid methods (SMOTE-ENN, SMOTE-Tomek, and SMOTE+OCSVM), ensemble strategies (SMOTEBoost, RUSBoost, Balanced Random Forest, and One-Sided Selection), and specialized approaches for multi-label and clustered data. Beyond descriptive categorization, this review critically examines each method’s underlying assumptions, operational mechanisms, and suitability for diverse data characteristics, including high dimensionality, mixed feature types, class overlap, and noise. Key findings demonstrate that no single method universally outperforms others; optimal selection depends critically on dataset characteristics, classifier choice, and evaluation metrics. The paper concludes by identifying emerging research directions, including self-supervised learning for imbalance, diffusion-based generative oversampling, distribution-preserving resampling, knowledge distillation for imbalanced deployment, and the adaptation of foundation models to skewed distributions, offering practical guidelines for practitioners and a roadmap for future methodological development. Full article
Show Figures

Figure 1

43 pages, 8097 KB  
Article
Toward Reliable Diabetic Retinopathy Screening
by Hendrio Bragança, Ítalo P. Caliari, Wington L. Vital, Antonio Fontenele, Sergio Cavalcante and Glaucio Messias
Sensors 2026, 26(14), 4515; https://doi.org/10.3390/s26144515 - 16 Jul 2026
Viewed by 687
Abstract
Diabetic retinopathy (DR) grading requires reliable five-grade severity assessment under substantial acquisition variability and cross-dataset distribution shift. We propose PRISM-DR, a multi-objective five-grade DR grading framework trained under a gradient-partitioned strategy. The architecture is organized as a feedforward pipeline: a data-driven preprocessing stage [...] Read more.
Diabetic retinopathy (DR) grading requires reliable five-grade severity assessment under substantial acquisition variability and cross-dataset distribution shift. We propose PRISM-DR, a multi-objective five-grade DR grading framework trained under a gradient-partitioned strategy. The architecture is organized as a feedforward pipeline: a data-driven preprocessing stage followed by a ConvNeXtV2-Base backbone, a Recurrent BiFPN neck for multi-scale feature fusion, a Frequency-Aware Fusion module, a lightweight multi-scale reasoning transformer, dual classification heads with gradient-isolated pathways (categorical and ordinal), and a prototype memory module for embedding regularization. The CORAL ordinal head operates through a dedicated projection layer and is gradient-isolated from the backbone; the backbone is shaped by the cross-entropy, prototype contrastive, and view-consistency objectives, which carry indirect ordinal signal through severity-weighted class penalties and grade-indexed cluster regularization. The model is trained in a multi-crop setting with a phased loss curriculum designed for severely imbalanced DR datasets. Evaluated across six datasets under Fixed-Source, Multi-Target (FSMT) protocols, PRISM-DR trained on EyePACS + DDR achieves QWK of 0.835 on IDRiD, 0.865 on APTOS2019, and 0.720 on Messidor-2, with in-domain QWK = 0.920 and AUC-PR = 0.941 on EyePACS, outperforming RETFound, RETFound-Green, and MedGemma-4B in AUC-PR across all evaluated datasets. Quantitative interpretability evaluation against 755 expert-annotated lesion images yields 8.0× Energy Ratio Enrichment and a FAF gate retention ratio of 4.4× inside lesion regions, confirming that anatomically plausible spatial priors emerge from grade-level supervision alone, without pixel-level annotation. PRISM-DR establishes a superior accuracy–robustness–capacity trade-off for scalable, automated DR screening. Full article
(This article belongs to the Section Biomedical Sensors)
Show Figures

Figure 1

Back to TopTop