Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (265)

Search Parameters:
Keywords = Manual Curation

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
27 pages, 4479 KB  
Article
Clinician-Guided Deep Learning Segmentation of Skull Base Pneumatization on Computed Tomography Using 3D Slicer and MONAI Label
by Cristian-Norbert Ionescu, Gergő Ráduly, Marian Pop, Karin Ursula Horváth, Dan Iovănescu and Gheorghe Mühlfay
Life 2026, 16(8), 1284; https://doi.org/10.3390/life16081284 - 3 Aug 2026
Abstract
Skull base pneumatization is anatomically variable and clinically relevant to temporal bone and transsphenoidal surgical corridors, but manual volumetric segmentation is time-consuming. This retrospective pilot study evaluated a clinician-guided deep learning workflow for mastoid and sphenoid sinus compartment segmentation on bone computed tomography. [...] Read more.
Skull base pneumatization is anatomically variable and clinically relevant to temporal bone and transsphenoidal surgical corridors, but manual volumetric segmentation is time-consuming. This retrospective pilot study evaluated a clinician-guided deep learning workflow for mastoid and sphenoid sinus compartment segmentation on bone computed tomography. Images were curated and annotated in 3D Slicer using MONAI Label and separate three-dimensional SegResNet models. The mastoid development dataset comprised 122 side-cases, with 28 reserved side-cases; the sphenoid dataset comprised 117 development and 17 reserved examinations. The best internal-validation Dice scores were 0.8660 for the mastoid model and 0.8750 for the sphenoid model. In AI-assisted correction cohorts, mean Dice ranged from 0.9539 to 0.9691 for mastoid and from 0.9347 to 0.9426 for sphenoid. In independently annotated subsets, AI-to-expert Dice was 0.8083–0.8097 for mastoid and 0.8339–0.8473 for sphenoid, while interobserver Dice was 0.8003 and 0.9136, respectively. AI assistance reduced mean segmentation time by 86.5% for mastoid and 81.4% for sphenoid. These findings support clinician-supervised AI segmentation as an efficient starting point for volumetric assessment. Full article
(This article belongs to the Special Issue Cranial Base Tumors: Pathogenesis, Diagnosis, and Treatments)
Show Figures

Figure 1

20 pages, 1844 KB  
Article
Clinically Validated XAI for Calcified Plaque Segmentation in Coronary CT Angiography
by Julius Siaulys, Agne Paulauskaite-Taraseviciene, Antanas Jankauskas, Gintare Sakalyte and Dovydas Verikas
J. Clin. Med. 2026, 15(15), 6039; https://doi.org/10.3390/jcm15156039 - 3 Aug 2026
Abstract
Background: Accurate segmentation of calcified plaques in coronary computed tomography angiography (CCTA) images is critical for reliable assessment of coronary atherosclerotic burden and for supporting interpretation of luminal stenosis, yet it remains a challenging task due to annotation inconsistencies, blooming artifacts and [...] Read more.
Background: Accurate segmentation of calcified plaques in coronary computed tomography angiography (CCTA) images is critical for reliable assessment of coronary atherosclerotic burden and for supporting interpretation of luminal stenosis, yet it remains a challenging task due to annotation inconsistencies, blooming artifacts and low contrast at lesion boundaries. These limitations may affect both automated model performance and the clinical trustworthiness of AI systems. This study explores the impact of annotation refinement on segmentation performance and model explainability, as well as the influence of representation learning on explanation quality. Methods: We trained a deep convolutional neural network to segment calcified plaques in coronary arteries using a dataset of expert-labeled CT slices. Initial training on radiologist-provided annotations yielded suboptimal results. To address this, annotations were manually revised and validated by radiologists. In addition to a standard ImageNet-pretrained model, we evaluated a self-supervised representation learning approach using DINOv2. Grad-CAM was used to generate visual explanations for model predictions before and after annotation refinement. Results: Models trained with refined annotations achieved notably improved segmentation accuracy, with clearer delineation of calcified regions. Grad-CAM localization analysis demonstrated improved concentration of model attention within plaque and vessel regions. Furthermore, models incorporating DINOv2 representations produced more spatially coherent attention maps, with improved anatomical localization consistent with coronary vessel regions and reduced off-target activations, as qualitatively confirmed by expert radiologists. Conclusions: Our findings emphasize the importance of high-quality, validated annotations in developing accurate and interpretable AI models for medical imaging. In addition, the results suggest that representation learning influences the reliability and clinical relevance of explainability outputs. The combination of manual annotation refinement, expert validation, and improved feature representations provides a practical workflow for human-in-the-loop AI development in cardiovascular imaging. This study demonstrates that annotation quality is a critical and often underestimated determinant of XAI reliability, and suggests that explainability methods can serve as feedback tools for iterative, clinician-guided dataset curation in cardiovascular imaging. Full article
(This article belongs to the Special Issue Cardiac Imaging: Emerging Techniques and Clinical Applications)
Show Figures

Figure 1

27 pages, 2839 KB  
Article
Semantic Clustering for Automated Few-Shot Exemplar Selection in LLM-Based Formative Feedback for Middle-School Mathematics: A Feasibility Study
by Yuv Raj Pant, Haitham Y. Adarbah, Afzel Noore, Dunren Che, Aden Ahmed and Robert Ayala
AI 2026, 7(8), 280; https://doi.org/10.3390/ai7080280 - 24 Jul 2026
Viewed by 263
Abstract
Large Language Models (LLMs) show promise for supporting formative assessment by generating feedback on students’ written mathematical reasoning. However, practical use in educational settings remains constrained by the need to manually curate representative few-shot exemplars for prompt construction. This study examines whether unsupervised [...] Read more.
Large Language Models (LLMs) show promise for supporting formative assessment by generating feedback on students’ written mathematical reasoning. However, practical use in educational settings remains constrained by the need to manually curate representative few-shot exemplars for prompt construction. This study examines whether unsupervised semantic clustering can automate few-shot exemplar selection for LLM-generated formative feedback in middle-school mathematics. As a controlled methodological feasibility study, we generated and refined 100 exam-realistic constructed responses for a Grade 7 inequality task aligned with middle-school mathematics standards. Student responses were embedded using Sentence-BERT, projected into a lower-dimensional space using Uniform Manifold Approximation and Projection (UMAP), and clustered with Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) to identify dominant reasoning patterns and ambiguous responses. Representative centroid and boundary exemplars from the resulting clusters were then used to construct few-shot prompts for the Llama 3.3 70B model, which generated feedback for the remaining 93 responses. Six independent mathematics instructors evaluated the AI-generated feedback using a structured 0–5 usability rubric. Across all instructor evaluations, 538 of the 558 instructor ratings (96.42%) were 3–5, representing feedback ranging from fair, requiring moderate edits, to excellent, ready to send. More specifically, 482 of the 558 instructor ratings (86.38%) were scores of 4–5, indicating feedback requiring no edits or only minor revisions. The remaining 20 ratings (3.58%) were scores of 0–2, while 13 unique feedback messages received at least one low rating. Across all instructor evaluations, 20 of 558 ratings (3.58%) were assigned scores of 0–2, while 13 unique feedback messages received at least one low rating. Qualitative analysis of low-scoring cases revealed recurring failure modes, including hallucinated completeness in concise solutions, failed arithmetic verification, and false logic flagging for atypical reasoning patterns. These findings suggest that clustering-based exemplar selection may reduce manual prompt-engineering effort while supporting usable LLM-generated formative feedback in a controlled mathematics setting. However, the present study does not compare clustering against alternative exemplar-selection strategies, and therefore conclusions should be interpreted as evidence of feasibility rather than comparative superiority. Full article
(This article belongs to the Special Issue How Is AI Transforming Education?)
Show Figures

Figure 1

27 pages, 11969 KB  
Article
ULSTM: Multi-Scale and Full-Level Temporal Consistency for Traffic Anomaly Detection
by Borja Pérez, Mario Resino, Jaime Godoy, Abdulla Al-Kaff and Fernando García
Smart Cities 2026, 9(7), 120; https://doi.org/10.3390/smartcities9070120 - 22 Jul 2026
Viewed by 182
Abstract
Urban traffic anomaly detection is essential for intelligent transportation systems, particularly in smart city environments where fast identification of abnormal events can improve road safety and traffic management. This work proposes a novel ULSTM-driven architecture that explicitly models temporal dependencies across consecutive traffic [...] Read more.
Urban traffic anomaly detection is essential for intelligent transportation systems, particularly in smart city environments where fast identification of abnormal events can improve road safety and traffic management. This work proposes a novel ULSTM-driven architecture that explicitly models temporal dependencies across consecutive traffic frames to achieve more stable and temporally coherent reconstructions. The proposed framework leverages sequential spatio-temporal representations to improve the distinction between normal traffic patterns and anomalous events. To further enhance reliability, we introduce a Hybrid Weighted Fusion strategy that synergistically combines structural, perceptual and pixel-wise metrics. The framework’s parameters are optimized using a Discrete Dirichlet Sampling approach, achieving a peak F1 Score of 70.28%. Evaluations were conducted on a manually curated traffic anomaly dataset with frame-level annotations. Experimental results demonstrate that the ULSTM framework significantly outperforms frame-independent generative models by suppressing high-frequency reconstruction noise, providing a robust solution for real-world smart city deployments. While highly effective in complex scenarios, the proposed framework is strictly applicable to highly dynamic traffic environments with active motion, as static background ensembles can degrade performance. Full article
(This article belongs to the Section Smart Urban Mobility, Transport, and Logistics)
Show Figures

Figure 1

38 pages, 27713 KB  
Article
Visual-Type Detection of Yuxian Paper-Cut Opera Figures for Digital Preservation of Intangible Cultural Heritage: A YOLOv8-SCSA-Based Approach
by Zhengqi Huang, Nan Ji, Xunrong Ye, Yi Zhang and Qiang Wang
Electronics 2026, 15(14), 3230; https://doi.org/10.3390/electronics15143230 - 22 Jul 2026
Viewed by 182
Abstract
Yuxian paper cutting is a distinctive regional form within UNESCO-recognized Chinese paper cutting. Its opera figure images remain difficult to organize at scale because interpretation and cataloguing rely heavily on manual expertise, while automated tools for preliminary annotation and retrieval remain limited. This [...] Read more.
Yuxian paper cutting is a distinctive regional form within UNESCO-recognized Chinese paper cutting. Its opera figure images remain difficult to organize at scale because interpretation and cataloguing rely heavily on manual expertise, while automated tools for preliminary annotation and retrieval remain limited. This study formulates a domain-specific object detection task and constructs a curated dataset of 926 original images containing 976 annotated instances across three operational visual categories: armor-clad, robe-wearing, and female-role figures. A YOLOv8n-based detector is adapted by integrating the Spatial and Channel Synergistic Attention (SCSA) module and the Inner-IoU regression strategy. On the test set, the resulting configuration achieves 95.8% precision, 92.9% recall, 98.8% mAP@0.5, and 89.6% mAP@0.5:0.95, exceeding the YOLOv8n baseline by 3.0, 3.6, 2.9, and 2.9 percentage points, respectively. It also achieves a forward-pass throughput of 92.3 FPS under the reported GPU configuration. Among the evaluated configurations, the combined SCSA placement with Inner-IoU achieved the highest overall detection performance under the current dataset and experimental protocol. The results support its use for preliminary visual-category annotation, retrieval, and organization of Yuxian paper-cut opera figure resources. However, the findings are based on a small domain-specific dataset and require further validation across alternative data partitions and broader heritage collections. Full article
Show Figures

Figure 1

15 pages, 1022 KB  
Article
Can AI Reliably Identify Marine Microplastics in Wildlife? Assessing Multi-Modal Foundation Models for Polymer Classification with Minimal Training
by Gabriela Fernandez, Domenico Vito, Siddharth Suresh-Babu, Dipsy Booth, Sayali Sanjay Shelke, Kawther Kaziz, Robert Yabumoto and Mohamed Banni
Int. J. Environ. Res. Public Health 2026, 23(7), 929; https://doi.org/10.3390/ijerph23070929 - 20 Jul 2026
Viewed by 263
Abstract
While existing AI-based microplastic monitoring studies predominantly rely on task-specific fine-tuned models, this study evaluates whether general-purpose multimodal foundation models with no spectroscopic instrumentation, minimal fine-tuning, and minimal computational resources can serve as accessible, low-cost tools for polymer classification under ecologically realistic field [...] Read more.
While existing AI-based microplastic monitoring studies predominantly rely on task-specific fine-tuned models, this study evaluates whether general-purpose multimodal foundation models with no spectroscopic instrumentation, minimal fine-tuning, and minimal computational resources can serve as accessible, low-cost tools for polymer classification under ecologically realistic field conditions with samples of plastic debris collected from coastal Sousse, Tunisia, a Mediterranean region experiencing anthropogenic pollution pressures. A curated dataset of 1080 high-resolution images was developed, representing six polymer groups (HDPE, LDPE, PA, PET, PP, PS, and mixed plastics). Fragments were imaged under standardized lighting conditions against natural sand backgrounds to preserve environmental realism. Each image was manually annotated using polygon-based boundaries to generate pixel-level segmentation masks and associated class labels, providing expert-validated ground truth for quantitative evaluation. Multimodal LLMs were evaluated using a composite scoring framework. Spatial accuracy was assessed using mean Intersection-over-Union (mIoU) against expert annotations, while polymer classification performance was measured using macro-averaged F1 scores across all categories. Model reliability was further evaluated through prompt stability testing and robustness analyses under controlled environmental perturbations designed to assess consistency across varying coastal imaging conditions. Results indicate that multimodal foundation models can distinguish plastic fragments from sand backgrounds, although performance varied across polymer classes and environmental perturbations. The weighted composite framework provides a structured approach for comparing model performance according to ecological monitoring objectives rather than computational metrics. These findings contribute to understanding the potential utility and current limitations of AI-based approaches for marine microplastic analysis and provide insights into coastal pollution patterns and ecosystem health within a One Health monitoring context. Full article
Show Figures

Figure 1

27 pages, 3582 KB  
Article
FaSDiNet: An End-to-End Efficient Network for Oracle Bone Fragment Rejoining Driven by Full-Image Candidate Screening
by Yang Hong, Rui Tian, Yuqing Zhou, Yunlong Feng and Chuanming Song
Appl. Sci. 2026, 16(14), 7236; https://doi.org/10.3390/app16147236 - 20 Jul 2026
Viewed by 262
Abstract
Oracle bone rejoining is central to oracle bone studies and digital cultural heritage curation, but many computer-assisted methods still rely on manually selected contour curves, candidate fracture edges, or local preprocessing. We introduce the Fast Siamese Diffusion Network for Oracle Bone Fragment Rejoining [...] Read more.
Oracle bone rejoining is central to oracle bone studies and digital cultural heritage curation, but many computer-assisted methods still rely on manually selected contour curves, candidate fracture edges, or local preprocessing. We introduce the Fast Siamese Diffusion Network for Oracle Bone Fragment Rejoining (FaSDiNet), a FasterNet-based Siamese framework with a latent-diffusion-based data augmentation module for full-fragment-image candidate screening and assisted rejoining of oracle bone fragments. FaSDiNet combines latent diffusion augmentation, a FasterNet-based Siamese feature encoder, and joint optimization of contrastive loss and binary cross-entropy loss to improve pairwise matching under limited training data. On the self-built OBF dataset, FaSDiNet achieves Top-1, Top-5, Top-10, and Top-15 recall values of 23.68%, 60.53%, 76.32%, and 84.21%, respectively. Among the evaluated OBF methods, its strongest comparative performance appears at broader candidate ranges, where it obtains the best Top-5, Top-10, and Top-15 recall values. On the public COBD dataset, FaSDiNet achieves Top-10 and Top-15 recall values of 93.8% and 96.9%, respectively, showing competitive candidate-retention performance at larger candidate ranges. These quantitative results indicate that FaSDiNet is especially effective for preserving truly rejoinable fragments within manageable ranked candidate lists, supporting efficient expert-assisted screening and offering a practical visual-computing workflow for oracle bone collation. Full article
Show Figures

Figure 1

21 pages, 2804 KB  
Article
First Complete Mitochondrial Genomes of the Invasive Mussel Perna viridis from Brazil and the Southwestern Atlantic
by André Oliveira Souza Lima, Rafael Schroeder, Gabriela S. Delabary, Gilberto C. Manzoni and Mayara C. Beltrão
Biology 2026, 15(14), 1199; https://doi.org/10.3390/biology15141199 - 20 Jul 2026
Viewed by 340
Abstract
The Asian green mussel Perna viridis is expanding across coastal systems of the Americas, creating demand for curated molecular references to support biosecurity, aquaculture monitoring, and comparative invasion studies. Here, we report the first complete mitochondrial genomes of invasive P. viridis from Brazil [...] Read more.
The Asian green mussel Perna viridis is expanding across coastal systems of the Americas, creating demand for curated molecular references to support biosecurity, aquaculture monitoring, and comparative invasion studies. Here, we report the first complete mitochondrial genomes of invasive P. viridis from Brazil and the southwestern Atlantic, generated from two specimens collected in Santa Catarina, currently the southernmost documented sector of the Brazilian range. Long-range PCR, Oxford Nanopore sequencing, read-supported assembly, manual annotation, comparative mitogenomics, nucleotide-diversity analysis, protein-level divergence, and phylogenomic inference were used to generate and evaluate complete mitochondrial references. The two circular assemblies were 16,015 bp and 16,011 bp and contained 38 annotated features, including 13 protein-coding genes, 23 tRNAs, and two rRNAs, all on the H strand, with an identical A+T content of 67.5%. Pairwise comparisons among the four complete P. viridis mitogenomes analyzed showed high overall similarity but retained 20–125 SNPs/mismatches and one to 12 indel events. Gene order was conserved within P. viridis but differed from P. canaliculus and P. perna, and phylogenomics placed the Brazilian assemblies within the P. viridis clade. These curated mitogenomes provide initial complete mitochondrial references for future marker testing and comparative studies of P. viridis in the southwestern Atlantic, but the two-specimen dataset should not be used to infer population origin, invasion routes, or connectivity. Full article
(This article belongs to the Special Issue Advances in Aquatic Ecological Disasters and Toxicology)
Show Figures

Figure 1

32 pages, 3228 KB  
Article
A Data Curation Framework for Unstructured Real-World Turkish Breast Imaging Reports
by Seda Yıldırım, Erkan Ülker and Necdet Poyraz
Appl. Sci. 2026, 16(14), 7189; https://doi.org/10.3390/app16147189 - 17 Jul 2026
Viewed by 203
Abstract
Breast imaging reports are typically stored as unstructured free-text documents, which limits their use in clinical analytics, research, and downstream computational applications. These challenges are particularly pronounced in Turkish because of its agglutinative linguistic structure, orthographic variability, and the limited availability of standardized [...] Read more.
Breast imaging reports are typically stored as unstructured free-text documents, which limits their use in clinical analytics, research, and downstream computational applications. These challenges are particularly pronounced in Turkish because of its agglutinative linguistic structure, orthographic variability, and the limited availability of standardized clinical corpora. This study presents a data curation framework for transforming heterogeneous real-world Turkish breast imaging reports into structured and machine-readable datasets, with a specific focus on the extraction and standardization of BI-RADS labels already recorded in routine clinical reports. The proposed pipeline integrates domain-specific preprocessing, text normalization, hybrid conclusion-section segmentation, multi-phase BI-RADS label extraction, duplicate and near-duplicate report handling, and modality-based separation within a transparent rule-guided workflow. The framework was applied to two real-world breast imaging datasets comprising 35,104 reports in Dataset 1 and 27,215 reports in Dataset 2 after overlap and duplicate control around the May 2023 transition period. The datasets included ultrasonography, mammography, and magnetic resonance imaging records. The framework achieved conclusion-section segmentation coverage rates of 96.71% and 99.94%, respectively, and BI-RADS label extraction coverage rates of 95.04% and 100.00%, respectively. These coverage values indicate the proportion of reports for which the rule-guided pipeline produced extractable outputs and should not be interpreted as precision, recall, F1-score, or accuracy against an independent gold-standard corpus. The curation process reduced report-level textual redundancy, improved structural consistency, and enabled systematic separation of report narratives from diagnostic assessment labels. Manual review of selected subsets was used as an internal plausibility check rather than as a substitute for independent radiologist-annotated gold-standard validation. By addressing the challenges of structuring free-text breast imaging reports in a low-resource language setting, this study provides a transparent and adaptable methodological basis for future clinical NLP, machine learning, and real-world healthcare analytics. Full article
Show Figures

Figure 1

23 pages, 5011 KB  
Article
Field Application of Terrestrial and Vessel-Based LiDAR with AI-Assisted Point-Cloud Processing for Sustainable Coastal Feature Extraction and Shoreline Management
by Joonkyu Park, Keunwang Lee and Joonghyeok Heo
Sustainability 2026, 18(14), 7258; https://doi.org/10.3390/su18147258 - 16 Jul 2026
Viewed by 239
Abstract
Accurate characterization of coastal environments requires high-resolution spatial data and robust analytical workflows which are capable of capturing the complexity of intertidal surfaces and engineered shoreline structures. This study presents a field application of terrestrial and vessel-based LiDAR with AI-assisted point-cloud processing to [...] Read more.
Accurate characterization of coastal environments requires high-resolution spatial data and robust analytical workflows which are capable of capturing the complexity of intertidal surfaces and engineered shoreline structures. This study presents a field application of terrestrial and vessel-based LiDAR with AI-assisted point-cloud processing to support coastal mapping and shoreline analysis in a complex tidal flat environment. Field measurements were conducted in a geomorphologically dynamic estuarine system dominated by wide tidal flats and diverse natural and artificial coastal features. A comparative assessment of terrestrial laser scanning (TLS) and vessel-based mobile mapping system (MMS) data revealed distinct platform characteristics: TLS provided long-range, high-fidelity geometric measurements suitable for broad intertidal zones, whereas vessel-based MMS offered rapid and continuous coverage but exhibited range-related limitations in offshore and distal areas. To extract key coastal features, an AI-assisted workflow was implemented in Trimble Business Center (TBC), primarily using the TLS dataset as the input for feature extraction. The vessel-based MMS data were used to evaluate complementary acquisition characteristics, including coverage continuity, scanning range, accessibility, and visualization of coastal features, rather than as direct input for joint TLS–MMS AI classification. A manually curated training dataset consisting of 80 object-level point-clouds—cars, streetlights, powerlines, and fences—was used to train a custom extraction model, while TBC’s built-in AI tools were employed to automatically separate ground surfaces, vegetation, and built structures from the TLS point-cloud. The TLS-based AI-assisted workflow produced a multilayered representation of the TLS data, enabling detailed delineation of intertidal flats, engineered shoreline structures, and adjacent artificial objects. Quantitative evaluation based on comparison with manually annotated ground truth yielded an overall F1-score of approximately 0.87, indicating practical extraction performance for the evaluated object types. The results highlight the complementary strengths of TLS and vessel-based MMS data and the practical applicability of TLS-based AI-assisted point-cloud processing in complex coastal settings. The results indicate that TLS-based AI-assisted point-cloud processing, supported by vessel-based MMS comparison, can provide an operationally useful approach for coastal feature extraction, shoreline monitoring, and coastal infrastructure assessment in complex tidal flat environments. By improving the acquisition, interpretation, and management of high-resolution coastal spatial information, the proposed workflow can support sustainable shoreline monitoring, coastal infrastructure maintenance, and evidence-based coastal environmental management in vulnerable tidal flat environments. Full article
Show Figures

Figure 1

21 pages, 32395 KB  
Article
OSM-CLIP: Enhancing Remote Sensing Image–Text Representation Learning with OpenStreetMap Data
by Alessio Pierdominici, Riccardo Ricci, Mohammed Alruqimi and Farid Melgani
Appl. Sci. 2026, 16(14), 7002; https://doi.org/10.3390/app16147002 - 13 Jul 2026
Viewed by 366
Abstract
Remote sensing vision–language models, such as RemoteCLIP and GeoRSCLIP, have advanced image–text representation learning. However, they rely on manually curated caption datasets that are expensive to scale and provide only global image-level supervision. In this paper, we introduce OSM-CLIP, a framework that exploits [...] Read more.
Remote sensing vision–language models, such as RemoteCLIP and GeoRSCLIP, have advanced image–text representation learning. However, they rely on manually curated caption datasets that are expensive to scale and provide only global image-level supervision. In this paper, we introduce OSM-CLIP, a framework that exploits the freely available, continuously growing annotations of OpenStreetMap (OSM) to provide regionally scalable, patch-level supervision for remote sensing image-text learning. We construct a large-scale dataset of over 265,000 satellite images covering the contiguous United States, each automatically paired with fine-grained geographic annotations scraped from OSM and mapped to individual image patches. A contrastive loss operating at the patch level associates each image region with its corresponding OSM textual description, enabling the model to learn spatially grounded representations without any manual labeling effort. After fine-tuning on standard remote sensing captioning datasets, OSM-CLIP achieves an average improvement of 10.81% in zero-shot classification, 5.06% in text-to-image retrieval (R@1), and 3.87% in image-to-text retrieval (R@1) over existing methods across 13 classification and 4 retrieval benchmarks. Our results demonstrate that freely available geographic annotations can serve as a powerful source of supervision for remote sensing vision–language models in regions with high-quality OSM coverage. Full article
Show Figures

Figure 1

20 pages, 3624 KB  
Article
Real-Time Post-Earthquake Structural Crack Segmentation Using a Quadrupedal Robotic Inspection Platform
by Kemal Hacıefendioğlu, Volkan Kahya, Ali Motamedi and Ayşecan Bostan
Appl. Sci. 2026, 16(14), 6922; https://doi.org/10.3390/app16146922 - 10 Jul 2026
Viewed by 248
Abstract
Post-earthquake structural inspections are critical for public safety and recovery, yet traditional manual assessments are slow, hazardous, and resource-intensive. This paper proposes a novel system that integrates a quadrupedal robot with a deep learning (DL) vision model to rapidly detect structural cracks and [...] Read more.
Post-earthquake structural inspections are critical for public safety and recovery, yet traditional manual assessments are slow, hazardous, and resource-intensive. This paper proposes a novel system that integrates a quadrupedal robot with a deep learning (DL) vision model to rapidly detect structural cracks and damage in the aftermath of earthquakes. A Unitree Go2 quadruped robot, equipped with cameras and sensors, is paired with a YOLOv8 instance segmentation network for near-real-time crack detection and localization. The approach addresses key limitations of manual post-disaster inspections by enabling operator-supervised, near-real-time visual crack screening in hazardous or hard-to-reach areas. The YOLOv8 model is trained on a curated dataset of crack and damage images to support crack detection and segmentation performance, and its advanced segmentation capabilities allow precise delineation of damaged regions. The integrated system is validated on a laboratory-scale concrete–steel frame with simulated damage. Preliminary results demonstrate that the Unitree Go2 quadruped robot can navigate and inspect structural elements while the AI model identifies and segments cracks and surface damage under near-real-time laboratory operating conditions. This work highlights the potential of combining advanced legged robotics and state-of-the-art DL for structural health monitoring (SHM), offering a preliminary visual screening tool that can support operator awareness and help prioritize areas requiring expert structural inspection. Full article
(This article belongs to the Section Robotics and Automation)
Show Figures

Figure 1

25 pages, 15591 KB  
Article
A Comparative Benchmark of Real-Time Detectors for Canopy Image-Based Blueberry Detection Toward Precision Orchard Management
by Xinyang Mu, Yuzhen Lu and Boyang Deng
Sensors 2026, 26(14), 4373; https://doi.org/10.3390/s26144373 - 10 Jul 2026
Viewed by 276
Abstract
Computer vision with artificial intelligence (AI) offers a promising tool for blueberry growers to accomplish orchard tasks such as harvest maturity assessment and yield estimation, which otherwise would be labor-intensive and prone to error. However, blueberry detection in natural environments remains challenging due [...] Read more.
Computer vision with artificial intelligence (AI) offers a promising tool for blueberry growers to accomplish orchard tasks such as harvest maturity assessment and yield estimation, which otherwise would be labor-intensive and prone to error. However, blueberry detection in natural environments remains challenging due to variable natural lighting, frequent occlusions by leaves and branches, and motion blur due to environmental factors and imaging devices. AI models such as deep learning-based object detectors promise to address these challenges, but they are data-driven, demanding a large-scale, diverse dataset that captures the complexities of real-world orchard conditions. Deployment of these models in practical scenarios often faces limited computing resources, highlighting the importance of achieving the right accuracy/speed/memory trade-off in model selection. This study presents a novel comparative benchmark analysis of advanced real-time object detectors, including YOLO (You Only Look Once) (v8–v12) and RT-DETR (Real-Time Detection Transformers) (v1–v2) families, consisting of 36 model variants, evaluated on a newly curated large dataset for blueberry detection. This dataset contained 661 canopy images collected with smartphones during the 2022–2023 seasons, consisting of 85,879 manually annotated instances (including 36,256 ripe and 49,623 unripe blueberries) that represent a broad range of lighting conditions, occlusions, and fruit maturity stages. Among the YOLO models, YOLOv12m achieved the best accuracy with a mAP@50 of 93.3%, while RT-DETRv2-X obtained a mAP@50 of 93.6%, the highest among all RT-DETR variants. The inference time varied with the model scale and complexity, and the mid-sized models appeared to offer a good balance between accuracy and speed. To further improve fruit detection performance, all models were fine-tuned using Unbiased Mean Teacher-based semi-supervised learning (SSL) with 1644 cross-source unlabeled canopy images acquired from ground-based machine vision platforms. SSL resulted in accuracy improvements of up to 2.0%, with RT-DETR-v2-X achieving the highest mAP@50 of 95.5%. These findings highlight the efficacy of SSL for leveraging cross-domain unlabeled data, although further research is needed to fully exploit its benefits. The curated dataset and developed software programs are publicly available to facilitate further research and practical deployment. Full article
(This article belongs to the Special Issue Feature Papers in Smart Agriculture 2026)
Show Figures

Figure 1

29 pages, 5478 KB  
Article
An AI-Based Framework for Automated Radiographic Bone Loss Measurement Using Segmentation and Geometric Landmark Modeling
by Mohammad Abdel-Majeed, Iyad Jafar, Omar AL-Karadsheh, Shorouq Al-Awawdeh, Siraj Zabadi and Mahdi Flefl
Algorithms 2026, 19(7), 562; https://doi.org/10.3390/a19070562 - 8 Jul 2026
Viewed by 351
Abstract
Accurate assessment of radiographic bone loss (RBL) is essential for periodontal diagnosis and staging; however, manual measurement from dental radiographs is labor-intensive, time-consuming and subject to inter- and intra-examiner variability. Existing AI-based methods primarily formulate bone loss assessment as classification, landmark prediction, or [...] Read more.
Accurate assessment of radiographic bone loss (RBL) is essential for periodontal diagnosis and staging; however, manual measurement from dental radiographs is labor-intensive, time-consuming and subject to inter- and intra-examiner variability. Existing AI-based methods primarily formulate bone loss assessment as classification, landmark prediction, or direct segmentation of thin anatomical structures, limiting measurement interpretability and robustness. This study proposes clinically interpretable two-phase framework for automated and clinically interpretable RBL estimation from periapical radiographs. The framework explicitly separates anatomical structure recognition from geometric measurement, improving transparency and reducing error propagation. In the first phase, deep learning models segment key anatomical structures, including the crown, root, third root and alveolar bone. In the second phase, a deterministic geometric algorithm extracts clinically relevant landmarks, including the cemento–enamel junction (CEJ), bone crest, and root apex, and computes root length, CEJ–bone crest distance, and radiographic bone loss following established periodontal measurement principles. The framework was evaluated on a curated dataset of annotated radiographs. DS-TransUNet achieved the best segmentation performance. Quantitative evaluation yielded mean absolute errors of 0.81 mm for CEJ–bone crest distance, 0.71 mm for root length, and 5.89% for RBL estimation. Bland–Altman analysis demonstrated minimal systematic bias (−1.03%) and good agreement with expert measurements across different disease severities, supporting the framework’s potential as an objective and clinically applicable tool for periodontal bone loss assessment. Full article
Show Figures

Figure 1

16 pages, 2835 KB  
Article
Automated Peak Annotation in Time-of-Flight Secondary Ion Mass Spectrometry via a Physics-Informed Probabilistic Framework
by Jiahua Chen, Yujie Cao, Xingyu Jiang, Chunpeng Wu, Qing Hao, Yun Hu and Jiahui Liu
Molecules 2026, 31(13), 2388; https://doi.org/10.3390/molecules31132388 - 7 Jul 2026
Viewed by 311
Abstract
Peak annotation in Time-of-Flight Secondary Ion Mass Spectrometry (ToF-SIMS) is a persistent bottleneck that typically requires the manual assignment of chemical formulas to hundreds of fragment ion peaks per spectrum. This work describes a physics-informed probabilistic framework that automates this task by combining [...] Read more.
Peak annotation in Time-of-Flight Secondary Ion Mass Spectrometry (ToF-SIMS) is a persistent bottleneck that typically requires the manual assignment of chemical formulas to hundreds of fragment ion peaks per spectrum. This work describes a physics-informed probabilistic framework that automates this task by combining five chemically motivated constraints—Gaussian mass accuracy, element composition priors, isotope pattern matching, nitrogen rule parity, and graded valence bounds—into a multiplicative belief score. We evaluate the framework on 643 ground-truth peaks from 151 compounds spanning both positive and negative ion modes, and we explicitly distinguish two regimes. As a scoring task—when the correct formula is present in the candidate list—the framework attains 52.3% Top-1 and 76.4% Top-3 accuracy, a 4.9-fold improvement over mass-only scoring. In fully automated end-to-end deployment, where candidates are generated de novo, Top-1 accuracy is 26.3%; the limiting factor is candidate generation rather than scoring, as only 46.5% of ground-truth formulas are currently produced by the database and combinatorial generator. Leave-One-Compound-Out Cross-Validation (59 compounds, 525 peaks) yields 51.8% Top-1 accuracy with fixed domain-knowledge weights, confirming generalization stability. Ablation analysis identifies element composition priors as the dominant non-mass constraint (−27.7 percentage points when removed), followed by isotope matching (−10.3 pp) and the nitrogen rule (−5.3 pp). The framework requires no labeled training spectra—relying instead on physically motivated priors and curated fragment databases—provides interpretable per-constraint scores (which represent relative rankings rather than calibrated probabilities), and supports polarity-specific configurations, offering a practical computational foundation for automated ToF-SIMS spectrum interpretation. Full article
(This article belongs to the Special Issue Application of Mass Spectrometry Techniques in Analytical Chemistry)
Show Figures

Figure 1

Back to TopTop