Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (2,391)

Search Parameters:
Keywords = sample-selection framework

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
21 pages, 3081 KB  
Article
Risk of Erythritol-Associated Ischemic Stroke: Integrated Genetic, Transcriptomic, and Mfachine Learning Evidence
by Ao Zhong, Fangyang Yu, Chuyue Xia, Xiang Ma, Qiucheng Zhu, Peilin Du, Ruonan Wang and Si Jin
Int. J. Mol. Sci. 2026, 27(17), 7811; https://doi.org/10.3390/ijms27177811 (registering DOI) - 31 Aug 2026
Abstract
Erythritol is a widely used low-calorie sugar substitute, but its relationship with cerebrovascular risk remains uncertain. We investigated the association between genetically predicted erythritol levels and ischemic stroke and explored stroke-related molecular features using an integrative bioinformatics framework. Two-sample and multivariable Mendelian randomization [...] Read more.
Erythritol is a widely used low-calorie sugar substitute, but its relationship with cerebrovascular risk remains uncertain. We investigated the association between genetically predicted erythritol levels and ischemic stroke and explored stroke-related molecular features using an integrative bioinformatics framework. Two-sample and multivariable Mendelian randomization were performed across cardiovascular–kidney–metabolic outcomes, followed by target prediction, enrichment and protein–protein interaction analyses, transcriptomic profiling, machine-learning feature selection, SHAP interpretation, immune-cell analysis, gene set variation analysis, and exploratory molecular docking. Genetically predicted erythritol showed the strongest association with stroke among the tested sweetener-related traits (OR = 1.246, 95% CI: 1.101–1.410, p < 0.001) and remained significant after adjustment for selected hemodynamic, glycometabolic, and lipid-related traits. Downstream analyses highlighted inflammatory, oxidative-stress, hypoxic, and vascular-injury pathways and prioritized MMP9, TLR4, and HIF1A as a reproducible stroke-related three-gene signature. These downstream bioinformatic findings are exploratory and do not establish erythritol-specific molecular regulation. Overall, the results support an association between genetically predicted erythritol levels and ischemic stroke and identify candidate pathways and genes for further investigation; the MR exposure should not be interpreted as direct evidence that dietary erythritol intake causes stroke. Full article
36 pages, 30239 KB  
Article
Framework for Cross-Disaster Building Damage Assessment Using Cost-Sensitive Learning
by Omer Aviv, Armin Shmilovici and Ofer Hadar
Remote Sens. 2026, 18(17), 2920; https://doi.org/10.3390/rs18172920 (registering DOI) - 31 Aug 2026
Abstract
Rapid and reliable assessment of structural damage following disasters is critical for prioritizing rescue operations. In this study, we present a unified deep learning framework for building damage assessment from satellite imagery. The proposed approach integrates segmentation-driven feature construction, morphological processing, cost-sensitive learning, [...] Read more.
Rapid and reliable assessment of structural damage following disasters is critical for prioritizing rescue operations. In this study, we present a unified deep learning framework for building damage assessment from satellite imagery. The proposed approach integrates segmentation-driven feature construction, morphological processing, cost-sensitive learning, and cross-disaster evaluation to support robust performance under limited and imbalanced data conditions. The framework combines an adapted U-Net for building localization with a hybrid convolutional neural network (CNN)-deep neural network (DNN) classifier for damage-level prediction and evaluates transferability across disaster events, geographic regions, and sensing conditions. The proposed method is evaluated on selected events from the xView2 Building Damage Assessment (xBD) and BRIGHT datasets, using optical imagery from xBD and pre-disaster optical and post-disaster Synthetic Aperture Radar (SAR) imagery from BRIGHT. Despite the limited and highly imbalanced event-specific samples, the framework achieves a mean cross-validation macro-F1 score of 70% and a maximum fold-level score of 77% on the Mexico earthquake subset of xBD and up to 98% on earthquake-related events in BRIGHT. Cross-validation characterizes performance variability across source-image-grouped data partitions, while cross-disaster evaluation reveals event-dependent transferability and provides a preliminary indication that structural domain similarity may be related to transfer performance. Although the evaluation is constrained by data availability, the results indicate that lightweight, cost-sensitive deep learning frameworks may support auxiliary post-disaster screening and decision support in resource-constrained scenarios. This study highlights both the potential and the remaining challenges of deploying artificial intelligence (AI) for rapid post-disaster assessment. Full article
(This article belongs to the Section AI Remote Sensing)
Show Figures

Figure 1

26 pages, 757 KB  
Review
Selenium and Iodine as Susceptibility Modifiers of Thyroid Disruption Associated with Metals: An Overview of Human and Experimental Research Data
by Maria-Nefeli Georgaki, Despoina Ioannou, Kanellos Skourtsidis, Georgios Kiosis, Theodora Papamitsou and Dimosthenis Sarigiannis
J. Xenobiotics 2026, 16(5), 164; https://doi.org/10.3390/jox16050164 (registering DOI) - 31 Aug 2026
Abstract
Background: Thyroid hormone synthesis, deiodination, and redox regulation depend on iodine availability and selenium-dependent proteins. Therefore, these nutrients may alter sensitivity to thyroid disturbance brought on by metals and metalloids; nevertheless, direct human evidence has not been compiled independently from rescue experiments. Goal [...] Read more.
Background: Thyroid hormone synthesis, deiodination, and redox regulation depend on iodine availability and selenium-dependent proteins. Therefore, these nutrients may alter sensitivity to thyroid disturbance brought on by metals and metalloids; nevertheless, direct human evidence has not been compiled independently from rescue experiments. Goal: To determine the mechanistic, biomarker, and study-design needs for interpretable human research, as well as to critically assess whether iodine or selenium alters metal-associated thyroid effects. Methods: Terms for metals, metalloids, iodine, selenium, and thyroid endpoints were used to search PubMed/MEDLINE until 14 July 2026. The database search was improved by selective forward citation searching, backward citation searching, and exact-title and DOI retrieval. A thyroid-specific outcome, a measurable or experimentally manipulated iodine or selenium variable, and a metal or metalloid exposure were all necessary for studies to be eligible. An author-developed framework that distinguished between formal interaction, stratification, joint-exposure modeling, contextual co-measurement, factorial nutritional-status experiments, physiologically interpretable supplementation, and pharmacological or nanoparticle rescue was used to categorize experimental evidence from humans and mammals. Results: Seven human studies and nine mammalian experimental studies made up the core evidence set. One additional human study was retained as contextual evidence. The results of the three human studies that directly assessed modification were mixed. One showed no clear interactions with iodine or selenium, one discovered an isolated strontium-by-iodine interaction, and one reported a suggestive mercury-by-iodine-supplement interaction. Most experimental trials employed high-dose, combination, parenteral, or nanoparticle rescue methods, but they more consistently demonstrated mitigation of thyroid damage by selenium-containing treatments. Conclusions: Although iodine- and selenium-dependent sensitivity is biologically feasible, there is currently little human data to support a consistent protective or detrimental modifying impact. Rather than supporting population-level prevention, experimental rescue promotes mechanistic modifiability. Repeated iodine testing, functional selenium biomarkers, metal speciation, vulnerable-window sampling, thyroid-specific outcomes, and predetermined interaction analyses are all necessary for future research. Full article
19 pages, 4469 KB  
Article
Construction of an Extended-Reach Well Torque Prediction Model Based on Multi-Source Drilling Data Fusion and Analysis of Dominant Factors
by Junrui Ge, Pengbo Li, Wei Liu, Bin Cai, Yanfei Li, Xuyue Chen, Xiujin Yuan and Penglei Tang
Processes 2026, 14(17), 2801; https://doi.org/10.3390/pr14172801 (registering DOI) - 31 Aug 2026
Abstract
Extended-reach wells contain long, highly inclined intervals in which drill-string/wellbore contact and cumulative friction evolve continuously with depth, making torque prediction strongly nonstationary. This study develops a multi-source torque-prediction framework using 5977 field samples from a deepwater extended-reach well in the East China [...] Read more.
Extended-reach wells contain long, highly inclined intervals in which drill-string/wellbore contact and cumulative friction evolve continuously with depth, making torque prediction strongly nonstationary. This study develops a multi-source torque-prediction framework using 5977 field samples from a deepwater extended-reach well in the East China Sea. To eliminate the optimistic bias caused by randomly mixing adjacent depth samples, all records were ordered by measured depth; preprocessing and feature screening were fitted only on historical data; hyperparameters were selected with forward-chaining TimeSeriesSplit; and model performance was assessed using both a fixed deep-depth stress test and expanding-window rolling-origin prediction. Training-only correlation screening and variance inflation factor analysis reduced multicollinearity before comparing ExtraTrees, XGBoost, LightGBM, BPNN, and time-aware ensemble regression. Static extrapolation to the deepest 15% of the well revealed severe distribution shift and poor long-range generalization. In contrast, 50 m rolling-forward updating produced reliable prediction: ExtraTrees achieved an MAE of 1.121, RMSE of 1.643, MAPE of 2.750%, R2 of 0.907, and Pearson correlation of 0.964, while LightGBM achieved the lowest MAE of 1.090. ExtraTrees R2 decreased from 0.927 at a 25 m horizon to 0.711 at 200 m, demonstrating progressive concept drift with prediction distance. SHAP analysis identified measured depth as the dominant predictive proxy, whereas WOB, flow rate, and RPM were more relevant to operational intervention. The revised framework therefore emphasizes leakage-free validation, adaptive updating, and separation of predictive from controllable factors rather than random-split interpolation accuracy. Full article
(This article belongs to the Special Issue Modeling, Simulation and Intelligentization in Offshore Drilling)
Show Figures

Figure 1

27 pages, 47370 KB  
Article
Geometry-Constrained Reference Sample Construction from Forest Inventory Compartments for Dominant Tree Species Mapping
by Pengfei Zheng, Wendou Liu, Xin Huang, Dongyang Han, Yibing Li and Shaozhi Chen
Remote Sens. 2026, 18(17), 2915; https://doi.org/10.3390/rs18172915 (registering DOI) - 31 Aug 2026
Abstract
Forest inventory compartments provide extensive and management-relevant reference information for satellite-based tree species mapping, but their dominant-species attributes are defined at the stand level rather than for individual image pixels. Existing applications commonly derive training samples from compartment centres or assign polygon labels [...] Read more.
Forest inventory compartments provide extensive and management-relevant reference information for satellite-based tree species mapping, but their dominant-species attributes are defined at the stand level rather than for individual image pixels. Existing applications commonly derive training samples from compartment centres or assign polygon labels to enclosed pixels, which may introduce boundary effects, uneven class representation, and disproportionate contributions from individual compartments. However, the intermediate step of converting inventory polygons into spatially controlled pixel-level reference samples has received comparatively limited attention. Here, we developed a geometry-constrained reference sample construction framework that integrates interior-position screening, class balancing, source compartment contribution control, and spatial-spacing constraints. The framework was evaluated for mapping Korean pine, larch, white birch, and spruce in a temperate mixed forest in northeastern China using Sentinel-1/2 time series and ancillary predictors. Predictor–classifier combinations were selected using compartment-grouped out-of-fold evaluation, sampling workflows were compared on 60 independently withheld compartments, and the final map was further assessed using 306 independent reference points. Relative to centroid sampling, the geometry-constrained workflow increased compartment-level macro-F1 from 0.612 to 0.709. XGBoost with optical time series and ancillary predictors achieved the best development-set performance, while inclusion of the complete Sentinel-1 time series provided no further gain. The final model achieved an overall accuracy of 0.827 and a macro-F1 of 0.820 on the independent reference points. Aggregation of 10 m predictions further enabled compartment-level characterization of mapped dominant species, dominance strength, and mixing intensity. These results demonstrate that reference sample construction is a consequential step in tree species mapping from polygon-based forest inventories and provide a practical approach for linking pixel-level remote sensing classification with forest management units. Full article
Show Figures

Figure 1

33 pages, 30121 KB  
Article
Geometry-Enhanced Point–Voxel Fusion with Active Learning for Label-Efficient Point Cloud Semantic Segmentation
by Cheng Zhang, Fei Meng, Yichang Qiu, Haofei Zhao, Yefei Liu and Jianfeng Huang
Appl. Sci. 2026, 16(17), 8650; https://doi.org/10.3390/app16178650 (registering DOI) - 31 Aug 2026
Abstract
Point cloud semantic segmentation is fundamental for 3D scene understanding and has been widely used in autonomous driving and infrastructure inspection applications. However, its performance is often limited by insufficient representation of local geometric structures and the high cost of point-wise annotation. To [...] Read more.
Point cloud semantic segmentation is fundamental for 3D scene understanding and has been widely used in autonomous driving and infrastructure inspection applications. However, its performance is often limited by insufficient representation of local geometric structures and the high cost of point-wise annotation. To address these issues, this paper proposes GeoFuse-AL, a label-efficient segmentation framework that integrates a geometry-guided point–voxel network with multi-cue active learning. The base model, GeoFuseNet, builds on a hybrid point–voxel backbone and incorporates a Local Geometry Prototype Attention module to enhance object boundaries, fine-grained structures, and local geometric patterns. An Adaptive Channel Fusion module is further designed to improve feature interaction between point-level details and voxel-level context. To reduce annotation dependence, a Multi-Cue Diversity Active Sampling strategy combines prediction uncertainty, color-gradient variation, geometric curvature, and feature-space clustering to select informative and diverse samples. Experiments on S3DIS and SemanticKITTI demonstrate that the proposed model achieves mIoU scores of 63.3% and 61.7%, respectively, outperforming several representative methods. Under limited annotation settings, the proposed strategy reaches 99.2% of fully supervised performance with only 15% labeled data on S3DIS and 97.6% with only 5% labeled data on SemanticKITTI. These results demonstrate that GeoFuse-AL improves segmentation accuracy while substantially reducing annotation requirements. Full article
Show Figures

Figure 1

22 pages, 5136 KB  
Article
Uneven Fisheries Observability Across the Western Pacific: Implications of Nighttime Light–AIS Integration for Sustainable Fisheries Monitoring
by Lei Chen, Chun Wang, Guangyuan Liu, Zhenxiang Ling, Zheng Cao, Qifei Zhang, Zhifeng Wu, Zhicheng Yang and Zihao Zheng
Sustainability 2026, 18(17), 8881; https://doi.org/10.3390/su18178881 (registering DOI) - 30 Aug 2026
Abstract
Background: Public AIS data may represent light-attracted fisheries unevenly across coastal regions, whereas VIIRS–DNB provides a complementary optical observation of bright offshore activity. Aims: We develop a regional framework to compare VIIRS-derived fishing-light patterns with AIS-derived vessel activity and to compare [...] Read more.
Background: Public AIS data may represent light-attracted fisheries unevenly across coastal regions, whereas VIIRS–DNB provides a complementary optical observation of bright offshore activity. Aims: We develop a regional framework to compare VIIRS-derived fishing-light patterns with AIS-derived vessel activity and to compare model-inferred environmental associations between East Asian and Southeast Asian coastal waters. Methods: Monthly VIIRS–DNB composites, GFW AIS-derived fishing effort and gear categories, and monthly CMEMS sea-surface temperature (SST), sea-surface height (SSH), sea-surface salinity (SSS), surface-current velocities, dissolved oxygen (DO), mixed-layer depth (MLD), and chlorophyll-a (Chl-a) were harmonized for 2017–2020. VIIRS detections were screened for coastal and fixed-light contamination, and spatial density, temporal correlation, spatial-error model (SEM), Geodetector, and MaxEnt analyses were applied. Results: East Asia showed more continuous high-intensity fishing-light belts with summer–autumn peaks, whereas Southeast Asia showed more fragmented activity with a stronger spring peak. In Southeast Asia, squid-jigging activity had the closest spatial correspondence among AIS categories and strong temporal synchronization with fishing-light intensity (Spearman’s rs = 0.890, p < 0.001), but its spatial explanatory power remained limited (q = 7.8%, p < 0.001). Within the selected predictor set, SST had the highest model-derived contribution in East Asia, whereas SSS had the highest model-derived contribution in Southeast Asia; these results represent model-inferred environmental associations rather than causal effects. Conclusions and recommendations: Regional contrasts describe VIIRS-derived fishing-light patterns, whereas differences in AIS correspondence indicate potential observation gaps but may also reflect fleet composition, vessel-light intensity, and genuine differences in fishing activity. We recommend combining optical and tracking observations and reporting uncertainty from AIS coverage, VIIRS detection limits, spatial sampling, and environmental-model assumptions. Full article
Show Figures

Figure 1

29 pages, 948 KB  
Article
Leakage-Free Multimodal Depression Screening: Controlled Evaluation of Text, Facial Behavior, and Prosodic Fusion
by Souaad Hamza-Cherif and Nesma Settouti
Bioengineering 2026, 13(9), 1009; https://doi.org/10.3390/bioengineering13091009 - 30 Aug 2026
Abstract
Multimodal behavioral sensing may support depression screening, but evaluation on small clinical-interview datasets is particularly vulnerable to data leakage and model-selection bias. We present a leakage-audited trimodal framework evaluated on DAIC-WOZ (n=180, PHQ-8 10), combining SBERT text [...] Read more.
Multimodal behavioral sensing may support depression screening, but evaluation on small clinical-interview datasets is particularly vulnerable to data leakage and model-selection bias. We present a leakage-audited trimodal framework evaluated on DAIC-WOZ (n=180, PHQ-8 10), combining SBERT text embeddings, OpenFace facial-behavior descriptors, and COVAREP prosodic features. Participant-level partitioning is performed before augmentation, while decision thresholds and neural-model checkpoints are selected exclusively from internal validation data. A controlled five-seed experiment showed that a deliberately leaky full-pool MixUp construction, in which a retained development sample could include a held-out participant as its second parent, was associated with a 27–33 percentage-point increase in Macro-F1 across four fusion configurations. Under the participant-level leakage-free 5-fold protocol, trimodal late fusion achieved a Macro-F1 of 0.532±0.041, compared with 0.446±0.021 for Text+Imaging late fusion. Paired participant-level correctness outcomes also favored trimodal fusion (McNemar χ2=6.618, p=0.010), consistent with improved paired classification when the audio modality was included under the leakage-free protocol. On 86 participant-disjoint E-DAIC sessions, using a consistent PHQ-8-based outcome definition (PHQ-8 10), trimodal late fusion achieved Macro-F1 = 0.621 and AUC = 0.676; this experiment is interpreted as within-family generalization rather than independent cross-corpus validation. Overall, the results show that leakage control can substantially alter both absolute performance and comparative conclusions, and that multimodal gains should be established through participant-level evaluation, modality-specific analysis, and reproducible model-selection procedures. Full article
Show Figures

Graphical abstract

23 pages, 1263 KB  
Article
Evidence-Based Decision-Making for Intraoral Scanner Selection in Clinical Dental Practice: Development of the EBIOS Pilot Framework
by Socratis Thomaidis, Georgios Chrisochoou, Eleni-Ioanna Tzaferi, Aikaterini Petropoulou and Maria Antoniadou
Prosthesis 2026, 8(9), 90; https://doi.org/10.3390/prosthesis8090090 (registering DOI) - 29 Aug 2026
Abstract
Background/ Objectives: Selecting an intraoral scanner (IOS) has evolved into a complex clinical decision that extends beyond technical performance alone. This study aimed to investigate the clinical, technical, educational, and organizational factors influencing intraoral scanner selection among dentists, evaluate awareness of ISO specifications [...] Read more.
Background/ Objectives: Selecting an intraoral scanner (IOS) has evolved into a complex clinical decision that extends beyond technical performance alone. This study aimed to investigate the clinical, technical, educational, and organizational factors influencing intraoral scanner selection among dentists, evaluate awareness of ISO specifications and evidence-based criteria, and propose a conceptual framework to support evidence-based technology selection. Methods: A cross-sectional questionnaire-based study was conducted among dentists practicing in Greece. A total of 86 questionnaires were returned from 271 invited participants (response rate: 31.73%). Valid sample sizes varied across analyses according to item applicability and analyzable responses. The questionnaire assessed demographic and professional characteristics, intraoral scanner use, selection criteria, awareness of ISO specifications, educational background, and evidence-related factors. Composite indices were developed to evaluate the principal decision-making dimensions. Results: Technical and clinical performance-related criteria received the highest importance ratings in intraoral scanner selection. Although participants generally reported good perceived knowledge of intraoral scanners, only 35.7% reported awareness of ISO specifications. Dentists familiar with ISO standards assigned significantly greater importance to specification-related and safety-related criteria, suggesting an association between reported ISO awareness and greater emphasis on specification- and safety-related selection criteria. Continuing professional education represented the predominant source of knowledge acquisition. Based on the integration of these findings, the EBIOS Pilot Framework (Evidence-Based Intraoral Scanner Selection Framework) is proposed as a preliminary conceptual synthesis of factors potentially relevant to evidence-informed intraoral scanner selection. Conclusions: The findings suggest that intraoral scanner selection may be understood as a multidimensional decision-making process involving technical performance, scientific evidence, professional education, and standardized quality criteria. The proposed EBIOS Pilot Framework provides a preliminary conceptual basis for future validation and refinement in larger, independent populations. Full article
Show Figures

Figure 1

232 pages, 10451 KB  
Article
Learning Nonparametric Conditional Single-Index U-Processes for Missing Locally Stationary Functional Random Fields with Stochastic Spatial Design
by Salim Bouzebda
Symmetry 2026, 18(9), 1453; https://doi.org/10.3390/sym18091453 - 29 Aug 2026
Abstract
We develop a design-conditional limit theory for kernel estimators of conditional U-functionals based on locally stationary functional random fields observed at irregular random locations and under incomplete response observation. The covariates take values in a separable Hilbert space, the responses are allowed [...] Read more.
We develop a design-conditional limit theory for kernel estimators of conditional U-functionals based on locally stationary functional random fields observed at irregular random locations and under incomplete response observation. The covariates take values in a separable Hilbert space, the responses are allowed to take values in a general Polish space, and the target is indexed by a class of symmetric kernels of a fixed order. Functional localization is induced by single-index semi-metrics, while spatial localization is performed on the rescaled observation domain. Missing responses are incorporated through a complete-case construction under a Missing At Random condition and a uniform-positivity assumption. The resulting estimator is a ratio of spatially weighted U-statistics with random tuplewise observation indicators. The asymptotic analysis must account simultaneously for four sources of complexity: dependence within the spatial field, nonstationarity across an expanding domain, concentration in an infinite-dimensional covariate space, and the random thinning generated by missing responses. Conditioning on the sampling locations removes the randomness of the spatial design weights but does not eliminate dependence among the observations. We therefore derive a design-conditional projection decomposition adapted to the triangular-array structure of the model. The leading component is represented by a spatially dependent complete-case empirical process, whereas the higher-order canonical terms are controlled uniformly over the response kernels, functional-target points, single-index directions, and rescaled spatial locations. The proofs combine stationary tangent-field approximations for locally stationary random fields, large-block–small-block decompositions, coupling arguments under spatial absolute regularity, small-ball probability estimates, and entropy bounds for the joint indexing class. These arguments yield a uniform stochastic expansion in which the empirical fluctuation, the spatial–functional smoothing bias, and the local-stationarity approximation error appear as distinct contributions. In particular, the local-stationarity remainder has no counterpart in the strictly stationary theory and quantifies the cost of replacing the observed nonstationary field with its stationary tangent approximation. Under the MAR and positivity conditions, complete-case sampling reduces the effective local information and modifies the covariance structure, but it does not change the formal order of the uniform-convergence rate. Under strengthened moment, mixing, entropy, and negligibility conditions, we establish weak convergence of the normalized conditional U-process in the corresponding supremum-norm function space to a tight centered Gaussian process. The limiting covariance is determined by the complete-case first-order projection and consequently retains the effect of the observation propensity and the spatial dependence structure. We also introduce a complete-case leave-tuple-out spatial prediction criterion for bandwidth selection and prove oracle optimality over admissible bandwidth families. The general theory applies to conditional rank association, discrimination probabilities, set-indexed conditional distribution functionals, and related pairwise statistical-learning criteria. Simulation experiments and applications to spatial environmental and epidemiological data illustrate the finite-sample implications of the theory and the stabilizing role of single-index localization. Viewed through the lens of data-driven science, the framework addresses a fundamental asymmetry between the information carried by irregular, locally heterogeneous functional covariates and the selectively observed response tuples. By combining design conditioning, complete-case normalization, tangent-field localization, and single-index dimension reduction, the proposed approach resolves this inferential asymmetry at the level of the model by matching estimation and uncertainty quantification to the information actually available locally, without imposing artificial stationarity or complete-data symmetry. Full article
(This article belongs to the Special Issue Symmetry and Asymmetry in Data-Driven Science)
32 pages, 3480 KB  
Article
Explainable Domain-Adaptive CNN–Transformer for Bidirectional Cross-Domain Bearing Fault Diagnosis
by Muhammad Javed, Suhang Ding, Hongxia Yan, Lei Ma and Teerath Kumar
Electronics 2026, 15(17), 3896; https://doi.org/10.3390/electronics15173896 - 28 Aug 2026
Viewed by 90
Abstract
Industry 5.0 requires resilient, adaptive, and trustworthy manufacturing systems capable of maintaining reliable diagnostic performance across heterogeneous industrial environments. However, data-driven fault diagnosis models often experience substantial performance degradation when transferred across machines, operating conditions, and data acquisition platforms because of domain distribution [...] Read more.
Industry 5.0 requires resilient, adaptive, and trustworthy manufacturing systems capable of maintaining reliable diagnostic performance across heterogeneous industrial environments. However, data-driven fault diagnosis models often experience substantial performance degradation when transferred across machines, operating conditions, and data acquisition platforms because of domain distribution shifts. This study proposes an explainable domain-adaptive CNN–Transformer framework for unsupervised bidirectional cross-domain bearing fault diagnosis using the Case Western Reserve University (CWRU) and Paderborn University (PU) datasets. The framework integrates one-dimensional convolutional layers for extracting local high-frequency vibration patterns, Transformer encoders for modelling long-range temporal dependencies, and Maximum Mean Discrepancy (MMD)-based feature-distribution alignment for learning transferable domain-invariant representations. Under the Unsupervised Domain Adaptation (UDA) protocol, the source domain supplies labelled samples for classification learning, whereas the target domain contributes unlabelled features only for MMD-based alignment; target labels are withheld from training and model selection and are used only for final evaluation. Conventional 1D-CNN and bidirectional long short-term memory baselines achieve over 90% accuracy in-domain but fall to 82.14% and 79.88%, respectively, for CWRU→PU, corresponding to domain-drop magnitudes of 12.07 and 12.99 percentage points. The proposed framework achieves 98.63% and 96.82% in-domain accuracy on CWRU and PU, respectively, and 92.46% for CWRU→PU and 94.18% for PU→CWRU, with domain-drop magnitudes of 6.17 and 2.64 percentage points. Ablation results confirm the complementary contributions of convolutional feature extraction, Transformer-based temporal modelling, and domain alignment. Furthermore, attention, saliency, and feature-importance analyses show that the model focuses on fault-relevant vibration regions and informative diagnostic characteristics, including kurtosis, root-mean-square (RMS), and crest factor, improving prediction transparency. These findings support accurate, transferable, and interpretable vibration-based condition monitoring across heterogeneous bearing datasets. Full article
40 pages, 9689 KB  
Article
FTU-Seek: Foundation Model-Guided Hard-Negative Learning for Sparse Functional Tissue Unit Segmentation
by Zonghao Liu, Lei Su, Jiguang Yu, Xuqing Geng, Louis Shuo Wang, Jianmin Wang and Jingfeng Liu
Biomedicines 2026, 14(9), 1935; https://doi.org/10.3390/biomedicines14091935 - 28 Aug 2026
Viewed by 135
Abstract
Background/Objectives: Functional tissue units (FTUs), including tertiary lymphoid structures (TLSs), blood vessels, and glands, encode localized immune, vascular, and epithelial organization in histopathology. Accurate quantification of these structures is important for studying tissue architecture and disease-associated tissue organization. However, FTUs are frequently [...] Read more.
Background/Objectives: Functional tissue units (FTUs), including tertiary lymphoid structures (TLSs), blood vessels, and glands, encode localized immune, vascular, and epithelial organization in histopathology. Accurate quantification of these structures is important for studying tissue architecture and disease-associated tissue organization. However, FTUs are frequently sparse, heterogeneous, and surrounded by large amounts of morphologically similar background tissue, making automated segmentation in whole-slide images (WSIs) challenging. We therefore developed FTU-Seek, a pathology foundation model-guided framework that treats morphology-aware negative-patch selection as a key component of sparse FTU segmentation. Methods: FTU-Seek uses frozen multi-depth features from the UNI pathology foundation model to train a patch-level classifier that distinguishes FTU-containing from FTU-absent tissue. Target-absent patches are subsequently ranked according to their predicted target-containing probabilities, and the highest-scoring hard negatives are selected through a static TopK strategy to construct compact segmentation training sets. The framework was evaluated using five-fold cross-validation and internal test cohorts across TLS, blood-vessel, and gland segmentation tasks, with an additional independent 30-WSI held-out cohort for TLS. Positive-only, all-tissue, random-negative, and matched random TopK sampling strategies served as comparators. Segmentation-derived phenotypes were further explored in external TCGA cohorts. Results: The patch-level classifiers achieved mean validation AUCs of 95.92%, 90.11%, and 98.16% for TLS, blood vessel, and gland classification, respectively. For TLS segmentation, the pre-specified Top1000 configuration retained 27.6% of the all-tissue training workload and achieved a slide-level Dice of 76.69 ± 11.89% on the independent 30-WSI held-out cohort. Compared with matched random Top1000 sampling, it improved Dice by 3.72 percentage points (95% CI, 2.05–5.38). Blood vessel and gland segmentation achieved performance approaching all-tissue training while reducing the retained training workload by approximately one-half and one-third, respectively. Compared with matched random sampling, classifier-guided hard-negative selection produced the greatest improvements for sparse and morphologically ambiguous FTUs. Exploratory TCGA analyses further showed associations of TLS phenotypes with overall survival, vascular phenotypes with overall survival and microvascular invasion, and glandular phenotypes with clinicopathological characteristics. Conclusions: FTU-Seek demonstrates that pathology foundation models can support sparse FTU segmentation not only through feature representation but also through morphology-aware construction of segmentation training sets. By prioritizing informative hard negatives, the framework reduces redundant segmentation-training workload while maintaining competitive segmentation performance and supporting quantitative tissue phenotyping from routine histopathology. Full article
(This article belongs to the Special Issue Human Stem Cells in Disease Modelling and Treatment (2nd Edition))
Show Figures

Figure 1

36 pages, 7707 KB  
Article
Differential Privacy-Based Location and Trajectory Data Protection for Utility-Preserving Location-Based Services
by Qihao Yu, Fang Liu, Xianghui Meng and Junjun Ma
Sensors 2026, 26(17), 5456; https://doi.org/10.3390/s26175456 - 28 Aug 2026
Viewed by 117
Abstract
The widespread use of location-based services (LBSs) has led to the continuous collection of user location and trajectory data, increasing the risk of privacy leakage and creating a persistent tradeoff between privacy protection and data utility. To address this problem in discrete location [...] Read more.
The widespread use of location-based services (LBSs) has led to the continuous collection of user location and trajectory data, increasing the risk of privacy leakage and creating a persistent tradeoff between privacy protection and data utility. To address this problem in discrete location query scenarios, this paper proposes a single-point location privacy protection method based on Q-R tree retrieval and differential privacy, termed QRDPP. QRDPP combines the adaptive spatial partitioning capability of a Q-tree with the minimum bounding rectangle (MBR)-based indexing capability of an R-tree. It applies an improved geometric privacy budget allocation strategy to leaf nodes and an arithmetic allocation strategy to non-leaf nodes, followed by Laplace perturbation of the corresponding location data and node information. For continuous trajectory query scenarios, this paper proposes a spatiotemporal generalization and differential privacy method, termed STG-DPTP, to address inadequate temporal protection, inappropriate generalization, and trajectory distortion. STG-DPTP performs hierarchical spatiotemporal clustering, separately models temporal and spatial distributions using Gaussian kernel density estimation, dynamically optimizes bandwidth parameters through Bayesian optimization, selects representative candidate subsets using the exponential mechanism, and generates protected trajectories through constrained sampling. Experiments on the GeoLife dataset evaluate the proposed methods in terms of query accuracy, computational efficiency, spatial trajectory similarity, reconstruction error, adversarial uncertainty, and temporal preservation. The results show that QRDPP improves the utility and efficiency of privacy-preserving spatial queries, while STG-DPTP better preserves the spatial distribution, trajectory structure, and temporal characteristics of the original data under the adopted differential privacy framework. Full article
(This article belongs to the Section Sensor Networks)
Show Figures

Figure 1

38 pages, 2738 KB  
Article
Causal Machine Learning for Heterogeneous Cost Effects in Mutual Funds: A Double Machine Learning and Causal Forest Approach
by László Vancsura
AI 2026, 7(9), 333; https://doi.org/10.3390/ai7090333 - 28 Aug 2026
Viewed by 258
Abstract
The cost–performance relationship in mutual funds is a longstanding open question in financial economics, particularly when costs are assumed to exert a single, linear effect on returns. This study proposes an integrated causal machine learning framework to revisit this question using a panel [...] Read more.
The cost–performance relationship in mutual funds is a longstanding open question in financial economics, particularly when costs are assumed to exert a single, linear effect on returns. This study proposes an integrated causal machine learning framework to revisit this question using a panel of Hungarian open-ended public investment funds across all major asset classes—equity, bond, absolute yield, misc, money market, real estate, and commodity—covering 2017–2024. Six machine learning algorithms are benchmarked for return prediction, and Double Machine Learning, with fund-level cluster-robust inference and year fixed effects, is applied to estimate the effect of the Total Expense Ratio (TER) on next-year returns, under the identifying assumptions stated in the paper, while flexibly controlling for a set of observed fund-level confounders (size, NAV dynamics, volatility, past and cumulative performance, and fund age) without imposing a linear functional form. To move beyond average effects, a Causal Forest model—tuned using an out-of-fold, effect size-neutral selection criterion—estimates heterogeneous treatment effects across funds, and SHAP-based interpretation uncovers the mechanisms underlying this heterogeneity. The results show that, once the outcome is measured in the year following the one in which TER is observed and panel dependence is properly accounted for, the average TER effect is not robustly different from zero at the full-sample level; where a statistically robust effect emerges, it is negative rather than positive, concentrated in equity and absolute-yield funds, and largely confined to the period after 2022, which coincided with the war in Ukraine, rising interest rates, and heightened market volatility, although the research design does not identify which, if any, of these developments drove the change. Average-effect models are shown to conceal this heterogeneity, and the results are further shown to be sensitive to two methodological choices that might otherwise appear secondary—the timing convention linking cost and return, and the criterion used to select among competing heterogeneous-effects specifications—underscoring the importance of making such choices explicit. These findings demonstrate the added value of combining predictive and causal machine learning, together with identification-robust and panel-robust inference, for uncovering heterogeneity that conventional econometric approaches overlook and offer a transferable methodological template for causal machine learning applications in finance and other high-dimensional decision-making domains. Full article
(This article belongs to the Section AI Systems: Theory and Applications)
Show Figures

Figure 1

42 pages, 2590 KB  
Article
Fidelity-Constrained Rule-Based Knowledge Distillation of Random Forests for Compact Breast Cancer Grade Classification
by Mert Büyükdede and Esma Gülfem Aktaş
Appl. Sci. 2026, 16(17), 8574; https://doi.org/10.3390/app16178574 (registering DOI) - 28 Aug 2026
Viewed by 73
Abstract
Breast cancer is a highly heterogeneous disease at the histopathological and molecular levels, and histological grade is a key prognostic indicator reflecting tumor biology. Although Random Forest models effectively capture nonlinear patterns in high-dimensional gene expression data, the decision logic of large tree [...] Read more.
Breast cancer is a highly heterogeneous disease at the histopathological and molecular levels, and histological grade is a key prognostic indicator reflecting tumor biology. Although Random Forest models effectively capture nonlinear patterns in high-dimensional gene expression data, the decision logic of large tree ensembles remains difficult to inspect directly. This study proposes a knowledge-distillation framework that transfers the probabilistic decision behavior of a Random Forest teacher model to a compact, rule-based student model for distinguishing low- from high-histological-grade breast tumors using 20,385 gene-expression features from 1892 METABRIC samples, with clinical information used for sample matching and histological-grade label definition. Root-to-leaf decision paths extracted from the teacher were converted into a binary rule-activation matrix, followed by proximal group sparsification and fidelity-constrained Top-K rule selection. For the principal 400-tree teacher configuration, the student retained 65.0 ± 13.7 rules and achieved an ROC-AUC of 0.833 ± 0.023 and an average precision of 0.843 ± 0.019, while maintaining a raw teacher–student Pearson fidelity of 0.9607 ± 0.0112. As teacher capacity increased, the number of active decision paths increased substantially, whereas the selected rule budget remained within a comparatively narrow range and increased only modestly. The compact student also showed approximately 46% lower measured end-to-end inference latency and a 76% smaller serialized deployment size than the 400-tree teacher. Within the present METABRIC evaluation, these findings indicate that the structural complexity of a tree ensemble and the complexity required to approximate its probabilistic behavior need not scale proportionally, and that the proposed framework can provide a favorable trade-off among predictive performance, teacher fidelity, and deployment compactness. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

Back to TopTop