Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (2,517)

Search Parameters:
Keywords = pre-trained neural networks

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
27 pages, 26649 KB  
Article
Evaluating Deep Learning Local Features for RGB-Thermal Image Matching and 3D InfraRed Thermography
by Luca Morelli, Neil Sutherland, Francesco Ioli, Alfonso Vitti, Stuart Marsh, Jon Mills, Paul Bryan and Fabio Remondino
Geomatics 2026, 6(4), 83; https://doi.org/10.3390/geomatics6040083 - 29 Jul 2026
Abstract
InfraRed Thermography (IRT), a non-invasive, non-contact, and non-destructive testing (NDT) technique, has become an established tool in the assessment of a building’s behavior and energy performance. However, the inherent low spatial resolution of thermal infrared (TIR) cameras has led recent work to fuse [...] Read more.
InfraRed Thermography (IRT), a non-invasive, non-contact, and non-destructive testing (NDT) technique, has become an established tool in the assessment of a building’s behavior and energy performance. However, the inherent low spatial resolution of thermal infrared (TIR) cameras has led recent work to fuse thermographic and geometric data to generate accurate 3D representations of buildings encapsulating temperature information. Whilst existing data fusion methods have relied on sensors in fixed relative orientation (RO), the co-registration of independent TIR and RGB blocks using ground control points (GCPs), or the reprojection of TIR images onto additional geometric or parametric models, approaches that directly match multi-modal images remain limited. In principle, if multi-modal tie points were available, it would be possible to directly align the RGB block with the TIR block; however, such matching is extremely challenging due to the substantial differences in radiometric properties. The main contribution of this paper is to demonstrate the applicability of off-the-shelf deep learning-based image matching algorithms, originally trained on mono-modal datasets, to multi-modal matching tasks for InfraRed Thermography 3D-Data Fusion (IRT-3DDF). We conduct a comparative evaluation of the principal algorithms developed in recent years, with particular emphasis on 3D accuracy and computational efficiency, under the hypothesis that, owing to the inherently local nature of the problem they address, these algorithms can generalize from a mono-modal training domain to a multi-modal application domain. The results are benchmarked against existing hand-crafted open-source multi-modal reference methods. Importantly, the proposed method is fully-automatic, obviating the need for sensor pre-calibration, manual co-registration, or associated positioning information. Results demonstrate that DL-based image matching, using pre-trained neural networks outside of their expected training domain, provides a viable approach for IRT-3DDF capable of co-registering blocks of multi-modal images across varying scales, settings, sensors, and subjects. Our results indicate accuracy in 3D is up to seven times better than multi-modal hand-crafted algorithms, while hand-crafted mono-modal methods fail to co-register images in their entirety. Full article
Show Figures

Figure 1

38 pages, 10402 KB  
Article
Topological Data Analysis for Characterising Earthquake Damage Patterns in Urban Building Clusters: A Novel Computational Framework with Benchmark Validation
by Enio Deneko, Marjo Hysenlliu, Klodian Dhoska and Andres Annuk
Buildings 2026, 16(15), 2963; https://doi.org/10.3390/buildings16152963 - 25 Jul 2026
Viewed by 255
Abstract
The spatial pattern of building damage produced by an earthquake carries information that classical building-by-building vulnerability indices cannot capture. This study presents one of the first frameworks to use Topological Data Analysis (TDA), a set of methods that quantify the “shape” of data, [...] Read more.
The spatial pattern of building damage produced by an earthquake carries information that classical building-by-building vulnerability indices cannot capture. This study presents one of the first frameworks to use Topological Data Analysis (TDA), a set of methods that quantify the “shape” of data, to characterise the spatial topology of seismic damage across an urban building inventory. Using the geo-referenced centroids of buildings as a point cloud, a sequence of connectivity graphs (a Vietoris–Rips filtration) is built at increasing distance scales, and persistent homology is used to track which spatial features appear and disappear. From this we extract four interpretable descriptors: Betti numbers (the numbers of connected building clusters and of enclosed gaps), persistence entropy (a measure of how disordered the damage pattern is), total persistence (the combined lifespan of all topological features), and the Wasserstein-2 distance (how far the post-earthquake pattern has moved from the intact pre-earthquake pattern). These descriptors form a physics-informed feature vector that is used to predict the building-cluster damage state. The developed framework was trained, tested, and validated on 1490 buildings over seven post-earthquake scenarios. Lognormal fragility parameters were estimated with maximum likelihood estimation, and an Artificial Neural Network (ANN) and a Random Forest (RF) were retrained on the same 593-building training dataset for comparison. On the 847-building benchmark, the TDA framework reached 93.3% accuracy (95% CI: 91.4–94.9%), F1 = 0.921 (0.902–0.940), and AUC = 0.933, using a stratified 70/15/15 split (training = 593, validation = 127, test = 127). This is a 6.0-percentage-point gain over the retrained ANN and a 12.1-percentage-point gain over the HAZUS-MH index (McNemar p = 0.017). Damage was recorded on the six EMS-98 states DS0–DS5, with DS4 and DS5 merged into a single class to give a five-class taxonomy, and building-type-specific inter-storey drift ratio thresholds were validated against EN 1998-3 (Eurocode 8 Part 3). Exact Rips computation is practical only for clusters up to about 2000 buildings; for larger populations, a CGAL (Computational Geometry Algorithms Library)-based sparse approximation with O(N log N) cost is recommended. It seems that the topological descriptions of the damage field may provide predictive information above and beyond that given by density and ground motion intensity and offer a reproducible tool for post-earthquake screening. Full article
(This article belongs to the Section Building Structures)
Show Figures

Figure 1

23 pages, 8304 KB  
Article
Enhancing the Explainability of the MRI-Based Brain Tumour Detection with Image Preprocessing
by Aykut Ismailov, Ina Zheleva, Petia Georgieva and Vladimir Dimitrov Hristov
Appl. Sci. 2026, 16(14), 7307; https://doi.org/10.3390/app16147307 - 21 Jul 2026
Viewed by 174
Abstract
Accurate and interpretable brain tumour detection from magnetic resonance imaging (MRI) is important for the reliable use of computer-assisted diagnostic systems. This study examines whether image preprocessing can improve the localisation quality of explanations generated by convolutional neural network (CNN) classifiers while preserving [...] Read more.
Accurate and interpretable brain tumour detection from magnetic resonance imaging (MRI) is important for the reliable use of computer-assisted diagnostic systems. This study examines whether image preprocessing can improve the localisation quality of explanations generated by convolutional neural network (CNN) classifiers while preserving high classification performance. Two pre-trained CNN architectures, ResNet50 and DenseNet121, were fine-tuned using the BRISC 2025 dataset, which contains 6000 annotated contrast-enhanced T1-weighted MRI images: 5000 training images and 1000 test images. The dataset includes four classes: glioma, meningioma, pituitary tumour, and healthy brain images. The original classification layers were replaced with custom fully connected heads designed for four-class classification. Model explanations were generated using Grad-CAM, Integrated Gradients, and LIME. Their localisation quality was evaluated against the available tumour segmentation masks using Intersection over Union (IoU), the Dice coefficient, and the Pointing Game metric. Tests show that both models (99.1% for ResNet50 and 99.2% for DenseNet121) perform well in terms of validation accuracy, but the explanation maps often operate in regions outside the clinically relevant area. To combat this issue, an image preprocessing pipeline utilising Otsu threshold masking, hole filling, and brightness–contrast jittering was implemented to filter noise from the background and isolate the focus area—brain region. After preprocessing, the validation accuracy of the DenseNet121 model was 99.6%, and the average Grad-CAM explainability metrics improved from 21.49% to 23.08% IoU, from 30.21% to 32.49% Dice coefficient, and from 48.33% to 51.67% Pointing Game score. The results indicate that conventional image preprocessing can moderately improve the spatial agreement between explanation maps and annotated tumour regions without reducing classification accuracy. Full article
Show Figures

Figure 1

7 pages, 564 KB  
Proceeding Paper
Fine-Tuning a Denoiser for Low-Dose CT Image Reconstruction
by Tim Selig, Thomas März, Martin Storath and Andreas Weinmann
Eng. Proc. 2026, 150(1), 38; https://doi.org/10.3390/engproc2026150038 - 21 Jul 2026
Viewed by 121
Abstract
A commonly employed imaging modality that relies on ionizing radiation is Computed Tomography (CT). While lowering the radiation dose is beneficial for patient health, it can result in reduced image quality. Therefore, improving low-dose CT (LDCT) reconstruction is a significant area of research. [...] Read more.
A commonly employed imaging modality that relies on ionizing radiation is Computed Tomography (CT). While lowering the radiation dose is beneficial for patient health, it can result in reduced image quality. Therefore, improving low-dose CT (LDCT) reconstruction is a significant area of research. The LoDoPaB-CT benchmark evaluates LDCT reconstruction methods, where many top methods use UNet-type architectures. We explore a two-stage approach for LDCT reconstruction: the first stage employs traditional filtered backprojection (FBP), while the second stage performs CT image enhancement. Our training strategy involves pretraining a neural network to denoise natural grayscale images, which are corrupted by Gaussian noise, followed by fine-tuning the network for CT image enhancement using LDCT and normal-dose CT (NDCT) pairs. Experiments on various small subsets of the LoDoPaB-CT dataset demonstrate the effectiveness of our method, showing that less task-specific data are required for training. Full article
Show Figures

Figure 1

23 pages, 11758 KB  
Article
Revisiting Deep Learning-Based Semantic Segmentation on Large-Scale Hydraulic-Structure LiDAR Point Clouds: A Spatial Surrogate Modeling Perspective
by Tianyang Chen, Wenwu Tang, Shen-En Chen, Craig Allan and Navanit Sri Shanmugam
Remote Sens. 2026, 18(14), 2413; https://doi.org/10.3390/rs18142413 - 20 Jul 2026
Viewed by 184
Abstract
Geospatial Artificial Intelligence (GeoAI) and the rapid advancement of 3D data acquisition technologies (e.g., LiDAR) have enabled scalable semantic interpretation of georeferenced large-scale 3D point clouds across different applications. 3D deep learning-based semantic segmentation on large-scale 3D point clouds typically relies on data [...] Read more.
Geospatial Artificial Intelligence (GeoAI) and the rapid advancement of 3D data acquisition technologies (e.g., LiDAR) have enabled scalable semantic interpretation of georeferenced large-scale 3D point clouds across different applications. 3D deep learning-based semantic segmentation on large-scale 3D point clouds typically relies on data partitioning and sampling strategies to address computational constraints, resulting in a substantial proportion of points not being directly predicted by deep neural networks. These points not directly predicted by the model require a processing step for label propagation, which is commonly handled using simple spatial proximity-based rules. While this simplification may be acceptable for some applications, it becomes critical in tasks that require precise spatial measurements and accurate object delineation, where propagation errors can directly affect downstream analyses. To bridge this research gap, this study views this process as a spatial surrogate modeling problem, where predictions from deep learning models are used to infer labels for underrepresented points based on spatial relationships. We adopt inverse distance weighting (IDW) as a transparent, deterministic, and training-free spatial post-processing strategy to examine whether explicitly incorporating spatial relationships improves segmentation outcomes. We evaluate the proposed method within an existing 3D semantic segmentation workflow for bridge inspection, a practical application requiring accurate spatial measurement and reliable object delineation. In the experiments, we use self-collected terrestrial LiDAR point clouds of bridges and associated hydraulic structures and systematically assess how neighborhood size and distance-decay parameters affect model performance on the segmentation task. Results show that this spatial post-processing step provides a modest enhancement over the conventional nearest-neighbor propagation baseline, with notable gains especially on spatially sparse or geometrically complex classes. The results also reveal class-dependent spatial effects, suggesting that different semantic classes exhibit distinct spatial dependencies during label propagation. These findings highlight the practical importance of accounting for spatial context when propagating semantic information to those points not directly predicted by deep neural networks. This is particularly important for downstream applications that require highly accurate spatial measurement and object delineation. The practical value of this post-processing step becomes more apparent in data-scarce application domains such as hydraulic-structure inspection, where large labeled point-cloud datasets and public benchmarks remain limited. In such settings, users may rely on domain-specific pre-trained models, while the original training data may not be publicly available for retraining or extensive model modification. Using a practical case study in the hydraulic domain, this study shows that deterministic post-inference label propagation can improve complete point-wise prediction from an existing model, thereby supporting the reuse of available models for domain-specific applications. The resulting response surfaces support two practical uses: site-specific calibration when limited labeled target data are available, and empirically informed initial settings, a rule of thumb, for comparable bridge-LiDAR applications when target-site labels are unavailable. Full article
Show Figures

Figure 1

28 pages, 8539 KB  
Article
AUKAT: Conditional VAE-Driven Augmentation and Neural Modeling of Enzyme Turnover Numbers
by Mengmeng Liu, Xialong Ni and Michal Brylinski
Biomolecules 2026, 16(7), 1049; https://doi.org/10.3390/biom16071049 - 18 Jul 2026
Viewed by 326
Abstract
Accurate prediction of enzyme turnover numbers (kcat) is essential for applications in systems biology, metabolic engineering, and drug discovery, yet remains challenging due to the limited availability and uneven distribution of experimental data. Here, we present AUKAT, an [...] Read more.
Accurate prediction of enzyme turnover numbers (kcat) is essential for applications in systems biology, metabolic engineering, and drug discovery, yet remains challenging due to the limited availability and uneven distribution of experimental data. Here, we present AUKAT, an integrated framework that combines conditional generative modeling with deep neural prediction to improve kcat estimation. A conditional variational autoencoder generates synthetic training instances in embedding space, followed by a selection pipeline that retains samples with strong agreement across independent evaluators, thereby ensuring data reliability. A hybrid convolutional neural network and transformer-based architecture is then used to predict kcat from substrate, enzyme functional, and species embeddings. Incorporating synthetic data improved predictive performance for both random forest and neural network models in five-fold cross-validation, with larger gains observed for the neural network architecture. Benchmarking against DLKcat demonstrated comparable predictive accuracy on the standard test set, while evaluation on stricter unseen subsets indicated improved generalization for low-similarity substrates and enzymes. Feature importance analysis further showed that AUKAT leverages substrate, enzyme functional, and species information in a more balanced manner rather than relying predominantly on a single feature source. In addition, AUKAT-human, a specialized model trained using a pre-training and fine-tuning strategy, achieved improved prediction accuracy for human enzyme kinetics. Overall, AUKAT provides a scalable approach for enzyme kinetics prediction and offers a practical solution to data scarcity in biochemical modeling. Full article
Show Figures

Figure 1

14 pages, 1005 KB  
Article
Noninvasive Two-Phase Foaling Prediction in Thoroughbred Mares Using Thermal Imaging and AI-Based Behavioral Analysis
by Hisashi Nabenishi, Nagisa Taki, Shoji Nishibayashi and Tomoyuki Ishii
Animals 2026, 16(14), 2221; https://doi.org/10.3390/ani16142221 - 17 Jul 2026
Viewed by 174
Abstract
Early and accurate detection of foaling is important for reducing perinatal mortality and labor demands in commercial breeding farms. This study evaluated a fully noninvasive foaling prediction system integrating thermal imaging with artificial intelligence-based behavioral analysis in Thoroughbred mares. A total of 115 [...] Read more.
Early and accurate detection of foaling is important for reducing perinatal mortality and labor demands in commercial breeding farms. This study evaluated a fully noninvasive foaling prediction system integrating thermal imaging with artificial intelligence-based behavioral analysis in Thoroughbred mares. A total of 115 pregnant mares across 13 farms were monitored during the 2024 foaling season. Locomotor activity and body surface temperature relative to ambient temperature were extracted from thermal images, while posture changes and tail-raising behavior were identified from visible-light images using a trained prediction model. Prediction was performed at 5 min intervals using a pretrained neural network model implemented in Python, which generated logistic probabilities based on locomotor activity, body surface temperature, posture change frequency, and tail-raising behavior. Data obtained during the 5 h preceding foaling were analyzed. Locomotor activity and surface temperature significantly increased 70–90 min before foaling (p < 0.05), whereas posture changes and tail-raising behavior markedly increased 25–45 min before foaling. These temporal patterns suggest a biologically interpretable two-phase pre-foaling process derived from sequential changes in behavioral and thermal indicators. The model based on locomotor activity and surface temperature achieved an 80.0% detection rate with a mean lead time of 186 ± 20 min. Incorporating posture changes and tail-raising parameters improved detection to 94.8% and reduced the detection-to-foaling interval to 89 ± 14 min. The system achieved high predictive performance under commercial conditions without requiring invasive or wearable devices. Full article
Show Figures

Figure 1

26 pages, 18614 KB  
Article
Sensor-Modality-Aware Human Activity Recognition with the Convolutional Tsetlin Machine: Interpretable and Resource-Efficient Neuro-Symbolic Learning
by Olga Tarasyuk, Anatoliy Gorbenko, Oleksandr Gordieiev, Artem Akulynichev, Rishad Shafik and Alex Yakovlev
Sensors 2026, 26(14), 4482; https://doi.org/10.3390/s26144482 - 15 Jul 2026
Viewed by 343
Abstract
Human activity recognition (HAR) based on smartphone and wearable sensor data is commonly addressed using statistical learning methods and deep neural networks that often provide strong predictive performance, but at the expense of limited interpretability and substantial computational and energy requirements. Such limitations [...] Read more.
Human activity recognition (HAR) based on smartphone and wearable sensor data is commonly addressed using statistical learning methods and deep neural networks that often provide strong predictive performance, but at the expense of limited interpretability and substantial computational and energy requirements. Such limitations reduce their suitability for deployment in practical sensing environments where model decisions must be transparent, verifiable and executable on resource-constrained devices. In this work, we investigate the Convolutional Tsetlin Machine (CTM) for multimodal HAR using only the raw inertial signals (9 × 128) of the UCI-HAR dataset, rather than its pre-computed 561-feature representation. The Tsetlin Machine is a novel neuro-symbolic machine learning approach that offers two important advantages over many conventional machine learning methods: (i) it learns logic-based decision rules that support human inspection and provide a transparent basis for analyzing model decisions, and (ii) it operates with comparatively low computational complexity, making it well suited to efficient and low-power on-device learning. The proposed study systematically analyses the contribution of different feature modalities by decomposing the inertial signals space into semantically defined subsets according to: (i) sensor source: accelerometer and gyroscope; (ii) signal group: gyroscope angular velocity, body and total acceleration (including gravity); (iii) coordinate axis: x, y and z. A separate CTM classifier was trained for each modality and its combinations in order to determine the relative discriminative value of each modality group for activity classification. In addition to predictive performance, the study emphasizes the interpretability of the CTM model ensured by expressing each decision in the form of propositional clauses, thereby enabling visualization and direct inspection of the modality-specific patterns supporting each activity class. Owing to its symbolic structure and modest computational demands, the CTM provides a principled framework for the design of explainable, resource-efficient and deployable HAR systems. The proposed work therefore contributes toward trustworthy multimodal sensing by jointly addressing predictive performance, interpretability and suitability for embedded and mobile platforms. Full article
(This article belongs to the Special Issue Multimodal Ubiquitous Sensing for Human-Centered Healthcare)
Show Figures

Figure 1

24 pages, 11435 KB  
Article
A Deep Learning Framework for EEG-Based Decoding of Visually Imagined Arrows with Different Colors and Directions
by Rami Alazrai, Oula Hatahet, Sahar Qaadan, Youssef Alothman and Mohamed Bader-El-Den
Biosensors 2026, 16(7), 383; https://doi.org/10.3390/bios16070383 - 14 Jul 2026
Viewed by 475
Abstract
Brain–computer interface (BCI) systems have demonstrated significant potential across medical, educational, and entertainment domains. Recently, visual imagery (VI) has emerged as an alternative to traditional motor imagery (MI) paradigms, offering a broader spectrum of control signals for dexterous assistive devices. In this study, [...] Read more.
Brain–computer interface (BCI) systems have demonstrated significant potential across medical, educational, and entertainment domains. Recently, visual imagery (VI) has emerged as an alternative to traditional motor imagery (MI) paradigms, offering a broader spectrum of control signals for dexterous assistive devices. In this study, we propose a novel BCI framework for classifying visually imagined arrows defined by different colors and directions. The proposed framework employs the Choi–Williams time–frequency distribution (CW-TFD) to construct a joint time–frequency–spatial representation (TFSR) of EEG signals. The resulting TFSR is converted into grayscale images and provided as input to a newly designed convolutional neural network (CNN), which performs 16-class decoding of visually imagined arrows defined by combined color and direction attributes. A new EEG dataset was collected from 16 subjects who imagined 16 distinct arrows comprising four colors and four directions. The framework achieved an average classification accuracy of 95.05% and a Cohen’s kappa score of 0.947 across the 16 classes. To comprehensively evaluate the proposed approach, three comparative analyses were conducted. First, multiple time–frequency representations were assessed for VI-based EEG decoding. Second, the proposed CNN architecture was benchmarked against several state-of-the-art pre-trained deep learning models. Third, the framework was compared with conventional machine learning classifiers using handcrafted features. Results demonstrate that the constructed CWD-based TFSR combined with the proposed CNN consistently outperforms alternative representations and classification models. These findings demonstrate the feasibility of decoding an expanded set of visually imagined color–direction arrow commands in a subject-specific EEG-based BCI setting, supporting further development of calibrated VI-based BCI systems for assistive and interactive applications. Full article
Show Figures

Figure 1

22 pages, 29683 KB  
Article
Transfer Learning and Optimized Machine Learning Techniques for Multiclass Diabetic Retinopathy Classification Using Retinal Images
by Mohammad Reza Yousefi, Ali Bakrani, Elias Ebrahimzadeh and Amin Dehghani
Diagnostics 2026, 16(14), 2189; https://doi.org/10.3390/diagnostics16142189 - 14 Jul 2026
Viewed by 259
Abstract
Background/Objectives: Diabetic Retinopathy (DR) is a prevalent and severe complication of diabetes, caused by prolonged hyperglycemia that damages retinal microvasculature and may ultimately lead to vision loss or blindness. While convolutional neural networks (CNNs) have shown promise in automating DR detection via retinal [...] Read more.
Background/Objectives: Diabetic Retinopathy (DR) is a prevalent and severe complication of diabetes, caused by prolonged hyperglycemia that damages retinal microvasculature and may ultimately lead to vision loss or blindness. While convolutional neural networks (CNNs) have shown promise in automating DR detection via retinal imaging, traditional approaches often suffer from limited diagnostic accuracy, long training times, and reliance on small or imbalanced datasets. Objective: This study evaluates an integrated transfer-learning using adaptive training strategies for multiclass retinal image classification. Methods: The proposed framework integrates transfer learning, feature-space dimensionality reduction, and adaptive training strategies based on an ImageNet pretrained ResNet50 backbone to improve training stability, computational efficiency, and multiclass retinal image classification performance. Results: The proposed Transfer Learning (TL)-based model was trained and evaluated on a large, publicly available dataset of retinal images, achieving an overall accuracy of 84%, maximum class-specific accuracy of 89%, sensitivity of up to 97%, and an F1-score of 92%. These results demonstrate reasonable overall classification performance under constrained data conditions. Conclusions: The proposed framework demonstrates the feasibility of integrating transfer learning and adaptive training strategies for multiclass retinal image classification under constrained benchmark conditions. However, the study is limited by the use of heavily downsampled retinal images, and further validation on high-resolution clinical datasets is required before practical deployment. Future methodological refinement and validation on high-resolution clinical datasets may support development of computer-assisted retinal image analysis systems. Full article
(This article belongs to the Special Issue Innovative Advances in Diagnosis Through Artificial Intelligence)
Show Figures

Figure 1

25 pages, 3429 KB  
Article
TabPFN-Based Prediction of Concrete Compressive Strength
by Zhihao Zhao, Jinjin Wang, Guohui Ma and Mingjie Han
Buildings 2026, 16(14), 2781; https://doi.org/10.3390/buildings16142781 - 13 Jul 2026
Viewed by 324
Abstract
The use of supplementary cementitious materials such as fly ash can reduce environmental impacts and improve the sustainability of concrete construction. However, the nonlinear interactions among mixture design parameters make accurate prediction of concrete compressive strength challenging. In this study, TabPFN, a pre-trained [...] Read more.
The use of supplementary cementitious materials such as fly ash can reduce environmental impacts and improve the sustainability of concrete construction. However, the nonlinear interactions among mixture design parameters make accurate prediction of concrete compressive strength challenging. In this study, TabPFN, a pre-trained foundation model for tabular data, was applied to predict the compressive strength of fly ash concrete and compared with tuned Random Forest, support vector regression, an artificial neural network, LightGBM, CatBoost, Ridge regression, and Abrams empirical regression. A dataset containing 1062 samples and eight mixture-level variables was used for model development and evaluation. Predictive performance was assessed using the coefficient of determination, mean absolute error, and root mean square error over 100 repeated random splits. The results showed that TabPFN achieved the best overall performance, with an average coefficient of determination of 0.9329, a mean absolute error of 3.2758 MPa, and a root mean square error of 4.6678 MPa. Compared with the strongest tuned gradient-boosting baseline, CatBoost, TabPFN reduced the mean absolute error and root mean square error by 0.8768 MPa and 0.8560 MPa, respectively. Furthermore, repeated-split conformal prediction demonstrated reliable uncertainty quantification, with an average prediction interval coverage probability of 0.9615 and a mean prediction interval width of 23.4554 MPa. SHAP analysis identified the water-to-cement ratio, mortar strength, and water-to-binder ratio as important variables, while additional multicollinearity and feature ablation analyses indicated that correlated ratio variables should be interpreted cautiously. The results indicate that TabPFN provides an accurate, robust, and uncertainty-aware framework for preliminary prediction of 28-day fly ash concrete compressive strength. Full article
(This article belongs to the Section Building Materials, and Repair & Renovation)
Show Figures

Figure 1

26 pages, 6268 KB  
Article
LFODet: Lightweight Few-Shot Object Detection with Meta-Learning in Remote Sensing Images
by Haoran Wu, Xuan Fang, Haonan Xiong and Xiaomei Yang
Sensors 2026, 26(14), 4371; https://doi.org/10.3390/s26144371 - 9 Jul 2026
Viewed by 342
Abstract
Balancing detection accuracy with model lightweightness remains a key challenge in remote sensing object detection. Although convolutional neural networks have improved performance, they typically require large-scale datasets, making few-shot detection of novel classes difficult. To tackle this, we propose LFODet, a lightweight few-shot [...] Read more.
Balancing detection accuracy with model lightweightness remains a key challenge in remote sensing object detection. Although convolutional neural networks have improved performance, they typically require large-scale datasets, making few-shot detection of novel classes difficult. To tackle this, we propose LFODet, a lightweight few-shot object detection network based on meta-learning. It uses two parallel branches to rapidly adapt to novel classes with limited samples while maintaining performance on base classes. For efficient feature representation, we integrate Semantic Ghost Channel Attention (GCA) and Fine-Grained Ghost Spatial Attention (GSA) to enhance semantic discriminability and spatial detail preservation. Moreover, we leverage Ghost convolutions to reduce computational complexity. The model is trained in three stages: base-class pre-training, meta-learner optimization, and few-shot fine-tuning. Experiments on DIOR and NWPU VHR-10 demonstrate that LFODet achieves stable and balanced performance across various few-shot learning scenarios. As validated on these benchmark datasets, this work provides a practical solution for resource-constrained remote sensing applications requiring rapid adaptation to new targets. Full article
(This article belongs to the Section Remote Sensors)
Show Figures

Figure 1

20 pages, 9011 KB  
Article
Inverse Design of Ultra-Wideband Microstrip Filters Based on Conditional Diffusion Networks
by Rongzhen Xu, Zhongfang Ren, Haoshun Zhang and Haipeng Wang
Electronics 2026, 15(14), 3014; https://doi.org/10.3390/electronics15143014 - 9 Jul 2026
Viewed by 257
Abstract
Traditional design methods for ultra-wideband (UWB) filters rely on complex electromagnetic (EM) simulations and iterative parameter optimization, consuming significant computational resources and design time. Current inverse design approaches predominantly employ generative adversarial networks (GANs) and convolutional neural networks (CNNs). However, these models often [...] Read more.
Traditional design methods for ultra-wideband (UWB) filters rely on complex electromagnetic (EM) simulations and iterative parameter optimization, consuming significant computational resources and design time. Current inverse design approaches predominantly employ generative adversarial networks (GANs) and convolutional neural networks (CNNs). However, these models often produce blurry topological boundaries, rendering them inadequate for UWB filters that demand extremely precise modeling of intricate features, such as microscopic stubs and exceedingly narrow gaps. To overcome this bottleneck, this paper proposes an inverse design framework integrating a conditional diffusion model with an evolutionary algorithm. The conditional diffusion model is uniquely suited for UWB inverse design, directly incorporating target scattering parameters as explicit conditions during training strictly guides the generation trajectory. This mechanism enables the model to synthesize high-fidelity, fine-grained pixelated patterns that traditional networks fail to achieve. During the inverse design process, the conditional diffusion model generates candidate topology-guided target S-parameters. Subsequently, an evolutionary algorithm, coupled with a pre-trained ResNet fast evaluator, iteratively searches the latent space to pinpoint the optimal geometric structure matching the target EM response. Validation through multiple UWB filter examples demonstrates that simulated S-parameters exhibit excellent agreement with target curves, achieving an efficient and precise design from EM performance to geometric structure. The two final fabricated inverse-designed filters exhibit 3 dB passbands ranging from 5.67 to 12.40 GHz and 7.35 to 14.60 GHz, achieving high fractional bandwidths of 74.53% and 66.06%, respectively. Full article
Show Figures

Figure 1

31 pages, 53017 KB  
Article
Lightweight Raw Echo Image Preprocessing for Long-Range Airborne Streak Tube Imaging LiDAR Using Adaptive Frequency-Domain Noise Suppression
by Chaowei Dong, Rongwei Fan, Zhaodong Chen, Zhiwei Dong, Deying Chen, Pengfei Hao and Lansong Cao
Remote Sens. 2026, 18(14), 2281; https://doi.org/10.3390/rs18142281 - 8 Jul 2026
Viewed by 232
Abstract
Long-range airborne streak tube imaging lidar (ASTIL) raw echo images are degraded by atmospheric speckle, detector noise, and weak-return fluctuations, which can bias centroid localization before range calculation. This study presents a lightweight preprocessing method combining row–column geometry-aware echo region pre-classification with frequency-domain [...] Read more.
Long-range airborne streak tube imaging lidar (ASTIL) raw echo images are degraded by atmospheric speckle, detector noise, and weak-return fluctuations, which can bias centroid localization before range calculation. This study presents a lightweight preprocessing method combining row–column geometry-aware echo region pre-classification with frequency-domain histogram-based adaptive suppression. Candidate regions are extracted from normalized streak images, classified by row–column morphology, filtered using local magnitude-spectrum percentile thresholds, and fused with a background-constrained weighted strategy. Simulated echo images, simulated point clouds, and 6 km airborne data were used for validation. In selected building roof control regions, the mean elevation root mean square error (RMSE) decreased from 0.34 m to 0.29 m, the mean absolute error (MAE) from 0.30 m to 0.26 m, and the mean roof elevation standard deviation from 0.19 m to 0.15 m, corresponding to an approximately 21% reduction in roof-level point cloud thickness. The results show that preprocessing before centroid extraction can improve roof-level vertical consistency without neural-network training or complex point cloud post-processing. Full article
Show Figures

Figure 1

16 pages, 5773 KB  
Article
Deep Learning-Based Multi-Class Pediatric Wrist Fracture Subtype Classification: A Pilot Study Comparing Convolutional Neural Network Architectures
by Rohan A. Phadke, Samer G. Salman, Zane G. Salman, Sai M. Yedupati, Joshua Ong, Alireza Tavakkoli, Sainyam Galhotra, Ajay Tripuraneni and James Rizkalla
J. Imaging 2026, 12(7), 307; https://doi.org/10.3390/jimaging12070307 - 8 Jul 2026
Viewed by 377
Abstract
Pediatric wrist fractures are among the most prevalent musculoskeletal injuries in children. Fracture subtype, including buckle/torus, greenstick, and Salter–Harris physeal injuries, directly influences management and prognosis. Subspecialty radiographic expertise required for subtype classification is not universally available in emergency or resource-limited settings. Deep [...] Read more.
Pediatric wrist fractures are among the most prevalent musculoskeletal injuries in children. Fracture subtype, including buckle/torus, greenstick, and Salter–Harris physeal injuries, directly influences management and prognosis. Subspecialty radiographic expertise required for subtype classification is not universally available in emergency or resource-limited settings. Deep learning (DL) offers an automated approach to fracture subtype recognition from plain radiographs. This pilot study evaluated convolutional neural network (CNN)-based five-class pediatric wrist fracture classification using the GRAZPEDWRI-DX dataset.A total of 940 pediatric wrist radiographs from GRAZPEDWRI-DX (figshare ID 14825193) were labeled using Arbeitsgemeinschaft fur Osteosynthesefragen (AO) pediatric codes into five classes: no fracture, buckle/torus, greenstick, Salter–Harris physeal fracture, and other fracture. Contrast-limited adaptive histogram equalization (CLAHE) and letterbox resizing to 224 × 224 pixels were applied. Patient-level stratified splits (70/15/15%) prevented data leakage. Three ImageNet-pretrained architectures (DenseNet-169, ResNet-50, and EfficientNet-B4) underwent two-phase transfer learning. Performance was assessed by balanced accuracy, macro F1, macro area under the receiver operating characteristic curve (AUROC), and Cohen’s kappa.DenseNet-169 achieved the highest balanced accuracy (0.371; 95% confidence interval [CI]: 0.289–0.448), macro F1 (0.334; 95% CI: 0.251–0.416), and macro AUROC (0.669), with Cohen’s kappa of 0.269 on the held-out test set (n = 139) under initial five-epoch pilot training conditions. All three networks exceeded a majority-class (no-information) baseline (balanced accuracy 0.20). Extending training to 50 epochs (approximately 2100 mini-batch iterations) with GPU acceleration substantially improved DenseNet-169 to a balanced accuracy of 0.532 (95% CI: 0.451–0.614), macro F1 of 0.516, and macro AUROC of 0.815, with statistically significant pairwise architecture differences (McNemar p < 0.01); per-class sensitivity was highest for no-fracture detection (0.969) and lowest for buckle/torus fractures (0.393). Gradient-weighted class activation mapping (Grad-CAM) confirmed anatomically coherent model saliency at the distal radial metaphysis and physeal plate.DenseNet-169 achieved the best five-class classification performance among evaluated architectures under pilot training conditions, and extended training substantially improved accuracy, although classification accuracy remained below clinically usable thresholds. These results establish a reproducible, patient-stratified DL pipeline and a benchmark for full-dataset training and future methodological development, rather than a clinically deployable tool. Full article
(This article belongs to the Special Issue Medical Computer Vision: Innovations and Clinical Impact)
Show Figures

Graphical abstract

Back to TopTop