Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (607)

Search Parameters:
Keywords = mechanic visual perception

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
23 pages, 4379 KB  
Article
A Geometry-Conditioned Symmetry-Aware Domain-Robust Observation Correction Front-End for Anti-UAV Visual Perception
by Yinlong Yuan, Liang Hua and Yun Cheng
Symmetry 2026, 18(8), 1378; https://doi.org/10.3390/sym18081378 - 16 Aug 2026
Abstract
Although the bounding boxes produced by an object detector provide real-time target localization cues for anti-UAV visual perception, they remain susceptible to geometric deviations under long-range small-target conditions, complex backgrounds, motion blur, and cross-domain environmental variations. Consequently, these detector outputs cannot always serve [...] Read more.
Although the bounding boxes produced by an object detector provide real-time target localization cues for anti-UAV visual perception, they remain susceptible to geometric deviations under long-range small-target conditions, complex backgrounds, motion blur, and cross-domain environmental variations. Consequently, these detector outputs cannot always serve directly as stable observations for state estimation, trajectory prediction, and interception control. To address this issue, this paper proposes CDBR-Net, a causally conditioned domain-robust observation correction network for post-detection UAV bounding-box refinement. CDBR-Net employs a shared encoder, disentangled multi-branch representations, and a quality-aware gating mechanism to jointly produce a corrected observation box, a robust representation, and an observation uncertainty estimate. To preserve this conditional environment-transformation symmetry without suppressing geometry-induced symmetry breaking, CDBR-Net constructs geometry-conditioned cross-domain sample pairs and imposes cross-domain consistency and geometry-sensitivity preservation constraints. After training on 8749 post-detection observations, CDBR-Net is evaluated on 997 aligned validation observations from three simulated scene domains. It reduces the YOLO bounding-box mean absolute error (MAE) by 6.8%, from 0.002924 to 0.002725, and increases the intersection over union (IoU) by 0.007979, from 0.852167 to 0.860146. It further reduces MAE by 1.7% and increases IoU by 0.002031 relative to the YOLO + MLP Residual baseline. Ablation studies demonstrate the complementary roles of geometry-conditioned cross-domain consistency and geometry-sensitivity preservation. These results indicate that CDBR-Net provides a more stable and geometrically consistent post-detection observation interface for anti-UAV visual perception. Full article
(This article belongs to the Section A: Computer Science)
Show Figures

Figure 1

20 pages, 2196 KB  
Article
A Novel Fast Detection and Localization Method for the ‘Sucui No.1 Pear’ Based on YOLOv11-Pear
by Denghui Li, Jun Li, Fahui Wang, Li Wang, Yafei Yang, Guoqiang Wang and Xujun Zhai
Agriculture 2026, 16(16), 1728; https://doi.org/10.3390/agriculture16161728 - 12 Aug 2026
Viewed by 209
Abstract
To address the visual perception challenges in automated harvesting of the ‘Sucui No.1 Pear’, this study proposes a fast detection and 3D localization method based on an improved YOLOv11 architecture and an RGB-D camera. First, a multi-scene ‘Sucui No.1 Pear’ dataset containing 5842 [...] Read more.
To address the visual perception challenges in automated harvesting of the ‘Sucui No.1 Pear’, this study proposes a fast detection and 3D localization method based on an improved YOLOv11 architecture and an RGB-D camera. First, a multi-scene ‘Sucui No.1 Pear’ dataset containing 5842 RGB-D image pairs and 31,559 labeled instances was constructed. Second, a lightweight YOLOv11-pear detection model optimized for pear fruit was developed; by reconstructing the feature pyramid network and introducing a global attention mechanism, using K-means++ clustering to optimize prior anchor boxes, and designing a composite loss function (Varifocal Loss + CIoU Loss + DFL), the detection accuracy was improved while maintaining lightweight design. The model adopts a “detect first, then fuse” strategy, achieving robust 3D coordinate calculation based on the median depth of the bottom region of the detection box. Experimental results show that YOLOv11-pear achieves 93.8% mAP@0.5 and 65.5% mAP@0.5:0.95 on the independent test set, with a precision of 94.2% and a recall of 91.5%. The model has only 5.8 M parameters and achieves real-time inference at 38.7 FPS on the Jetson Orin NX edge platform, with a mean absolute error of 9.9 mm for 3D localization. In severely occluded and complex lighting scenarios, the mAP@0.5 reaches 80.9% and 87.5%, respectively. After integrating the vision system into the harvesting robot platform, end-to-end closed-loop testing in a real orchard achieved an 88% harvesting success rate and 5% fruit damage rate. The average time for the visual perception stage was 1.3 s, accounting for 9.4% of the harvesting cycle. This research provides a high-precision, lightweight, and deployable vision solution for automated harvesting of the ‘Sucui No.1 Pear’, and has important reference value for promoting the development of intelligent fruit harvesting technology. Full article
(This article belongs to the Section Agricultural Product Quality and Safety)
Show Figures

Figure 1

24 pages, 13353 KB  
Article
Morphological Buffering in Neuro-Architecture: Interactive Effects of Color, Shape, and Area Proportion on Affective Processing
by Xiaoxiao Dou, Yan Zhang, Yannan Zhang, Mengyao Kang, Qiangqiang Fan, Xiaona Xie, Yufan Sun and Mengyao Li
Behav. Sci. 2026, 16(8), 1364; https://doi.org/10.3390/bs16081364 - 9 Aug 2026
Viewed by 185
Abstract
While architectural visual features significantly influence human emotion, the spatiotemporal neuro-mechanisms underlying their interactive effects remain poorly under-researched. This study systematically decoupled the dynamics of indoor spatial perception by investigating the interactive impacts of color combination, color shape, and area proportion to find [...] Read more.
While architectural visual features significantly influence human emotion, the spatiotemporal neuro-mechanisms underlying their interactive effects remain poorly under-researched. This study systematically decoupled the dynamics of indoor spatial perception by investigating the interactive impacts of color combination, color shape, and area proportion to find out the effects on participants’ emotion. By integrating high-resolution electroencephalography (EEG), specifically analyzing the N200, late positive potential (LPP), Frontal Alpha Asymmetry (FAA), and Beta bands with Self-Assessment Manikin (SAM) ratings and Liking scores in immersive virtual indoor environments, data from 61 valid participants were analyzed. The results revealed a hierarchical neuro-affective processing structure: morphology dominated early cognitive processing, where curved geometries acted as a pre-attentive cognitive buffer, acting as a cognitive buffer by significantly attenuating the early-stage visual stress induced by angular, high-contrast configurations. Conversely, color combination primarily drove late-stage sustained environmental arousal, tracked by the LPP. Notably, area proportion demonstrated a significant amplification effect, significantly exacerbating pre-existing neuro-affective biases. Crucially, cross-modal triangulation revealed a selective inter-modal coupling: while sustained cortical activation (ΔLPP) significantly predicted subjective emotional arousal, early physiological buffering (ΔN200/ΔFAA) was functionally dissociated from subjective aesthetic preference (ΔLiking). This decoupling provides preliminary evidence consistent with a dual-system processing framework, suggesting that while curves may facilitate a preconscious biological safe baseline, ultimate aesthetic appreciation is likely subject to top-down cognitive over-riding. Ultimately, these findings transition indoor spatial color design from intuitive practice to evidence-based neuro-aesthetics, highlighting the critical role of morphological buffering and EEG metrics in optimizing human emotional well-being. Full article
(This article belongs to the Section Cognition)
Show Figures

Figure 1

23 pages, 6923 KB  
Article
Fast-YOLO11n: A Lightweight and Efficient Apple Detection Model for Complex Orchard Environments
by Jinan Gu, Zhongkai Shen, Juan Liu and Xinyu Jiang
Agriculture 2026, 16(16), 1697; https://doi.org/10.3390/agriculture16161697 - 7 Aug 2026
Viewed by 290
Abstract
Accurate and real-time apple detection in complex orchard environments is essential for robotic harvesting but remains challenging because of illumination variation, foliage occlusion, and limited computational resources. This study proposes Fast-YOLO11n, a lightweight detector derived from the nano variant of You Only Look [...] Read more.
Accurate and real-time apple detection in complex orchard environments is essential for robotic harvesting but remains challenging because of illumination variation, foliage occlusion, and limited computational resources. This study proposes Fast-YOLO11n, a lightweight detector derived from the nano variant of You Only Look Once 11 (YOLO11n) and integrating three complementary components. A Fast-C3k2 module based on partial convolution (PConv) reduces redundant computation while preserving cross-layer feature transmission. A focal modulation (FM) mechanism enhances target-related responses and suppresses background interference under occlusion and uneven illumination. In addition, a parallel downsampling module, termed ADown, retains local geometric details and multi-scale semantic information during downsampling. Experiments were conducted on a field-collected orchard dataset comprising 2240 images and 22,673 annotated apple instances under diverse lighting, scale, and occlusion conditions. Fast-YOLO11n achieved mean average precision values of 75.76% across intersection-over-union (IoU) thresholds of 0.50–0.95 (mAP@50–95) and 91.29% at an IoU threshold of 0.50 (mAP@50), while operating at 366.19 frames per second (FPS) with 2.51 million parameters and 6.00 billion floating-point operations (FLOPs). Compared with the YOLO11n baseline, it improved mAP@50–95 and mAP@50 by 2.39 and 1.39 percentage points, respectively, while reducing the parameter count and FLOPs by 2.71% and 5.36%. Ablation experiments demonstrated the individual and combined effects of the three modules on detection performance and computational efficiency. The proposed model provides a favorable balance between detection accuracy and computational efficiency, indicating its potential for real-time orchard perception on resource-constrained platforms. Full article
Show Figures

Figure 1

28 pages, 11281 KB  
Article
A Biologically Inspired Unsupervised Artificial Visual System for Motion-Direction Detection via Hierarchical Agglomerative Clustering
by Yingjie Zhang, Zhiyu Qiu, Tianqi Chen, Yuki Todo and Zheng Tang
Electronics 2026, 15(16), 3516; https://doi.org/10.3390/electronics15163516 - 7 Aug 2026
Viewed by 203
Abstract
Motion perception in biological vision is influenced by experience-dependent processes, during which direction-selective responses emerge and are further refined over time, although the underlying mechanisms remain partially understood. This study proposes a biologically inspired Hierarchical Agglomerative Clustering (HAC)-based unsupervised artificial visual system (AVS) [...] Read more.
Motion perception in biological vision is influenced by experience-dependent processes, during which direction-selective responses emerge and are further refined over time, although the underlying mechanisms remain partially understood. This study proposes a biologically inspired Hierarchical Agglomerative Clustering (HAC)-based unsupervised artificial visual system (AVS) for motion-direction detection. The model consists of a local motion-processing layer and a global motion inference layer. At the local level, motion-direction responses are extracted using retina-inspired local motion detection neurons (LMDN) to capture pixel-level spatiotemporal variation. These local direction responses are clustered using prototype-aware HAC to infer the global motion direction at the global level. Additionally, an unsupervised learning mechanism inspired by the concepts of neural plasticity and critical periods has been established. Extensive computer simulations are conducted under varying noise conditions and object scales. The results indicate that the proposed AVS maintained robust global motion-direction detection accuracy across diverse test conditions and exhibited biologically interpretable features resembling selected features of biological visual systems. In conclusion, this HAC-based unsupervised AVS provides a computational model for studying the formation of motion representations in visual systems and supports the development of artificial vision systems. Full article
Show Figures

Figure 1

34 pages, 10258 KB  
Article
FruitDet: A Multi-Module Lightweight Detector for Young Apple Fruits Under Day–Night Orchard Conditions
by Jipeng Chen, Jinzheng Yu, Langyu Tang, Rong Zhang, Jinyan Li, Hongda Chen, Zhiyuan Zhang, Yang Liu and Hongfei Yang
Agriculture 2026, 16(15), 1684; https://doi.org/10.3390/agriculture16151684 - 5 Aug 2026
Viewed by 257
Abstract
Reliable perception of young apple fruits in natural orchards is a prerequisite for automated thinning and intelligent orchard management, yet remains difficult in real field conditions due to small fruit size, dense distribution, branch–leaf occlusion, background similarity, and severe illumination degradation at night. [...] Read more.
Reliable perception of young apple fruits in natural orchards is a prerequisite for automated thinning and intelligent orchard management, yet remains difficult in real field conditions due to small fruit size, dense distribution, branch–leaf occlusion, background similarity, and severe illumination degradation at night. This study presents FruitDet, a lightweight multi-module detector designed for robust day–night young apple fruit detection in complex orchard environments. A field dataset was established in a high-density apple orchard in Aksu, Xinjiang, covering daylight and low-light night-time scenes with diverse occlusion, scale, and illumination variations. To improve detection robustness without sacrificing computational efficiency, FruitDet combines three complementary mechanisms: an inverted-bottleneck-based multi-scale feature enhancement module for preserving small-fruit details, a channel–spatial attention module for suppressing foliage and illumination interference, and a lightweight Transformer-based context module for modeling long-range dependencies between fruits and surrounding orchard structures. In daytime scenes, FruitDet achieved 91.904% precision, 77.557% recall, 83.254% mAP50, and 66.427% mAP50–95; in night-time scenes, it maintained 90.107% precision, 75.135% recall, 80.544% mAP50, and 64.719% mAP50–95. Compared with mainstream detectors including YOLOv5n, YOLOv8n, YOLO11n, YOLO26n, Faster R-CNN, RT-DETR, and RT-DETRv2, FruitDet consistently delivered higher accuracy across lighting conditions. Ablation, visualization, public-dataset testing, and edge-deployment experiments verified that the proposed modules jointly improve small-object representation, background discrimination, low-light robustness, and real-time applicability. With 2.960 M parameters, 3.726 G FLOPs, and approximately 180 FPS, FruitDet offers a practical and efficient visual perception approach for Young fruit monitoring was conducted under both daytime and night-time orchard conditions covered in this study. All-weather orchard monitoring and robotic young-fruit thinning. The shareable data are available Full article
(This article belongs to the Special Issue Advances in Precision Agriculture in Orchard)
Show Figures

Figure 1

25 pages, 859 KB  
Article
Object Detection and Scene Perception for Connected and Autonomous Vehicles Using LM-JEPA
by Abhishek Gupta and Ajmery Sultana
Sensors 2026, 26(15), 4894; https://doi.org/10.3390/s26154894 - 3 Aug 2026
Viewed by 212
Abstract
This paper presents the latent model-joint embedding predictive architecture (LM-JEPA), a resource-efficient collaborative perception framework for connected and autonomous vehicles that integrates latent predictive representation learning with lightweight multi-modal reasoning. Autonomous driving in urban and highway environments requires accurate scene understanding under strict [...] Read more.
This paper presents the latent model-joint embedding predictive architecture (LM-JEPA), a resource-efficient collaborative perception framework for connected and autonomous vehicles that integrates latent predictive representation learning with lightweight multi-modal reasoning. Autonomous driving in urban and highway environments requires accurate scene understanding under strict latency, energy, and communication constraints, limiting the practicality of large language model (LLM) and vision–language model (VLM)-based approaches in edge deployments. To address this, LM-JEPA encodes heterogeneous inputs including camera, LiDAR, radar, and map data into a unified latent space using a joint embedding predictive architecture, enabling efficient perception and reasoning without token-level inference. Unlike existing latent-space learning approaches that primarily learn predictive visual embeddings for single-modal perception, the proposed framework integrates multi-modal latent reasoning and adaptive sensor fusion to support collaborative perception under resource-constrained vehicular edge environments. The collaborative perception framework introduces a context-adaptive multi-modal fusion mechanism that dynamically weights sensor and model contributions, along with selective latent transmission and adaptive decoding for resource-aware operation. A lightweight VLM is integrated with an edge-assisted vehicular pipeline to support real-time on-vehicle inference with adaptive offloading based on latency and energy constraints, while a latent-space reasoning module enables cooperative decision-making. Experiments on BDD100K and nuScenes-QA show that LM-JEPA improves perception accuracy by 5% and reduces latency by approximately 7% over LLM and VLM baselines, while achieving up to 25% improvement in scene understanding, 20% higher intersection success rates, improved highway merging, and approximately 15% reduction in the transmitted model parameters. Full article
(This article belongs to the Special Issue Vehicular Sensing for Improved Urban Mobility: 2nd Edition)
Show Figures

Figure 1

30 pages, 1132 KB  
Article
An Artificial Intelligence-Driven UAV and Ground Sensor Fusion Framework for Crop Growth Assessment in Smart Agriculture
by Puxing Gao, Keyue Wang, Yunuo Li, Jiayue Zhang, Qingyu Li, Wenjie Lu and Yihong Song
Agriculture 2026, 16(15), 1650; https://doi.org/10.3390/agriculture16151650 - 31 Jul 2026
Viewed by 343
Abstract
With the rapid development of artificial intelligence, UAV remote sensing, and agricultural Internet of Things technologies, crop growth monitoring is evolving from manual inspection and single-source analysis toward intelligent decision-making based on multisource perception. However, existing methods still suffer from limited robustness under [...] Read more.
With the rapid development of artificial intelligence, UAV remote sensing, and agricultural Internet of Things technologies, crop growth monitoring is evolving from manual inspection and single-source analysis toward intelligent decision-making based on multisource perception. However, existing methods still suffer from limited robustness under environmental variations, insufficient integration between UAV imagery and sparse ground sensor observations, and weak capability for transforming predictions into practical agricultural management recommendations. This study proposes a UAV–ground sensor collaborative lightweight framework for crop growth assessment and agricultural decision support. The proposed framework integrates UAV RGB and multispectral imagery with ground sensor observations through a region-level aerial–ground alignment mechanism and a sensor-guided attention fusion module, enabling environmental conditions to enhance visual feature interpretation. Furthermore, a fact-constrained decision module is developed to generate management recommendations based on crop status, environmental risks, and field information. Experimental results demonstrate that the proposed method achieves superior performance in crop growth classification and yield-trend prediction, reaching Accuracy, Precision, Recall, and F1-score values of 92.47%, 91.86%, 91.39%, and 91.62%, respectively, with an RMSE of 0.381 and an R2 of 0.902. The lightweight framework requires only 6.18M parameters and 0.91G FLOPs, achieving 39.56 ms inference latency and 25.28 FPS on edge devices. The proposed framework also improves decision reliability, achieving an expert agreement rate of 89.34% and a risk identification accuracy of 90.18%. Economic analysis indicates that the proposed framework reduces labor cost, water consumption, and fertilizer input by 49.7%, 26.7%, and 23.0%, respectively, while increasing net benefit by 46.1% compared with conventional field management practices. These results demonstrate that the proposed method provides an accurate, interpretable, and deployable AI-driven solution for intelligent crop management in smallholder and medium-sized farming systems. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Figure 1

59 pages, 1990 KB  
Article
A Modular Reference Architecture and Co-Simulation Platform for Software-Defined Vehicles in a Software-Defined Internet of Vehicles Framework
by Zhenqian Li, Valentin Ivanov and Jochen Seitz
Appl. Sci. 2026, 16(15), 7518; https://doi.org/10.3390/app16157518 - 28 Jul 2026
Viewed by 510
Abstract
The automotive industry is evolving toward Software-Defined Vehicles (SDVs) enabled by centralized computing, cloud integration, and Over-the-Air (OTA) updates. Yet, prevailing SDV and Internet of Vehicles (IoV) simulators often treat each vehicle as a single monolithic node, obscuring the interplay between internal vehicle [...] Read more.
The automotive industry is evolving toward Software-Defined Vehicles (SDVs) enabled by centralized computing, cloud integration, and Over-the-Air (OTA) updates. Yet, prevailing SDV and Internet of Vehicles (IoV) simulators often treat each vehicle as a single monolithic node, obscuring the interplay between internal vehicle modules and the surrounding infrastructure in dense urban scenarios. This work proposes a modular SDV reference architecture embedded in a Software-Defined Internet of Vehicles (SD-IoV) framework together with a Software-in-the-Loop (SiL) co-simulation testbed built on Objective Modular Network Testbed in C++ (OMNeT++), Simulation of Urban MObility (SUMO), and Vehicles in Network Simulation (Veins). The architecture decouples perception, communication, decision, and actuation into typed replaceable modules and instantiates them across six co-existing agent types: an SDV; two human-driver vehicle classes with cognition modelled as a multi-stage Eye–Ear–Brain–Hand–Foot pipeline with reaction-delay sampling; a public transport bus; a Roadside Unit (RSU); and a Traffic Light (TL). Three platform-level mechanisms connect the agents to the infrastructure: a single shared world model with a three-layer line-of-sight funnel that serves visual-sensor queries and reuses the building polygons of the wireless shadowing model; a dual-CPU mobile-fog node implementing a cycles-per-frequency workload model with explicit end-to-end latency decomposition; and a three-plane intersection coordination fabric that combines 802.11p wireless with a wired RSU-to-TL star and a wired peer mesh between adjacent TLs. The initial results confirm that the implemented message paths and module interactions behave as specified, including directional Signal Phase and Timing (SPaT) reception, cross-junction handover, bus-side fog-offload latency accounting, and passive identification of Vehicle-to-Everything (V2X)-silent vehicles. Several architecture elements are specified but deliberately not exercised in the present evaluation and remain design targets for future work: the Roadside Unit (RSU) route planning and fog computing companion (and any multi-tier offloading comparison), non-line-of-sight SPaT reception, and a safety violation detection layer. Within the above scope, the testbed is positioned as a reusable foundation for module-level SDV research and as a basis for future extensions such as Joint Communication and Sensing (JCAS), energy-aware driving, and Hardware-in-the-Loop (HiL) integration. Full article
(This article belongs to the Special Issue Intelligent Autonomous Vehicles: Development and Challenges)
Show Figures

Figure 1

17 pages, 10115 KB  
Article
RG-RGD: Task-Triggered Small-Target RGB-D Depth Refinement for Robotic Laser Ablation
by Bowen Si, Dayong Ning, Jiaoyi Hou, Yongjun Gong, Ming Yi, Fengrui Zhang and Zhilei Liu
Machines 2026, 14(8), 841; https://doi.org/10.3390/machines14080841 - 25 Jul 2026
Viewed by 216
Abstract
Robotic 3D manipulation of small targets, such as laser ablation of urban-tree fruit balls, requires locally reliable depth. Mainstream RGB-D depth-refinement methods usually optimize image-wide metrics, which can leave task-relevant regions under-resolved and limit their direct use in precision robotic operation. This paper [...] Read more.
Robotic 3D manipulation of small targets, such as laser ablation of urban-tree fruit balls, requires locally reliable depth. Mainstream RGB-D depth-refinement methods usually optimize image-wide metrics, which can leave task-relevant regions under-resolved and limit their direct use in precision robotic operation. This paper presents residual-gated RGB-D depth refinement (RG-RGD), a task-triggered local depth-refinement framework that reallocates computation toward task-relevant 3D geometry after a candidate target has been selected. The method contains three coupled designs: a self-play benefit-driven foveation mechanism that focuses network capacity on regions where refinement reduces geometric residuals; residual prediction with Bayesian measurement fusion that anchors predictions to available raw observations; and inertial measurement unit (IMU)-conditioned self-supervised training that improves inter-frame view consistency. Evaluated on the Visual Odometry with Inertial and Depth (VOID) benchmark, RG-RGD obtains competitive standard depth-completion metrics, including a mean absolute error (MAE) of 24.95 mm and an inverse mean absolute error (iMAE) of 10.85. On self-collected London plane fruit-ball sequences, the method reduces region-of-interest geometric error more strongly than full-image error. A qualitative demonstrative laser-ablation use case illustrates how refined local geometry can drive physical branch filtering, cutting-point selection, and gimbal-based execution. The results support task-triggered local depth refinement as a practical perception component for robotic manipulation. Full article
(This article belongs to the Section Robotics, Mechatronics and Intelligent Machines)
Show Figures

Figure 1

42 pages, 25950 KB  
Review
A Review of Research Status of Advanced Technologies and Equipment for Underground Crop Harvesting Based on Soil Stratification
by Jun Zhang, Jiahao Shen, Chirui Zhang, Gan Liu, Tiantian Jing and Zhong Tang
Appl. Sci. 2026, 16(15), 7436; https://doi.org/10.3390/app16157436 - 24 Jul 2026
Viewed by 381
Abstract
Mechanized harvesting of subsurface crops has long been confronted with the critical engineering dilemmas of high damage rates and high impurity rates. Traditional taxonomic classification methods based on botanical families and genera fail to provide effective guidance for the engineering research and development [...] Read more.
Mechanized harvesting of subsurface crops has long been confronted with the critical engineering dilemmas of high damage rates and high impurity rates. Traditional taxonomic classification methods based on botanical families and genera fail to provide effective guidance for the engineering research and development of harvesting machinery. From an engineering perspective, this paper proposes a novel classification logic that categorizes subsurface crops into three major types based on their soil burial depth and physical distribution characteristics: shallow-soil clustered growth type (0–20 cm), mid-soil scattered growth type (20–40 cm), and deep-soil vertically rooted type (>40 cm). The harvesting bottlenecks of representative crops within these strata, including potato, onion, peanut, sweet potato, cassava, and yam, are systematically elucidated. Furthermore, this review provides an in-depth analysis of the current state of frontier core technologies, such as bionic drag reduction excavation, flexible multi-stage separation, microscopic discrete element method (DEM) simulation, kinematic optimization, and AI-based visual perception. This paper aims to reveal the common bottlenecks in subsurface crop harvesting and prospect future developmental trends centered on the deep integration of machinery and agronomy as well as intelligent perception and adaptation, thereby providing a solid theoretical foundation and engineering reference for the innovation of global agricultural machinery. Full article
(This article belongs to the Section Agricultural Science and Technology)
Show Figures

Figure 1

31 pages, 6721 KB  
Article
PPO-GAT-Follow: Graph-Attention Reinforcement Learning for Robust Robot Person Following in Dense Crowds
by Xinyu Zhou, Yongliang Shi, Songhao Piao and Chao Gao
Sensors 2026, 26(15), 4711; https://doi.org/10.3390/s26154711 - 24 Jul 2026
Viewed by 252
Abstract
Robot person following (RPF) in dense crowds requires a mobile robot to maintain an appropriate relative position with respect to a moving target while avoiding surrounding pedestrians and satisfying rear-following and social constraints. This paper proposes PPO-GAT-Follow, an interaction-aware reinforcement learning framework for [...] Read more.
Robot person following (RPF) in dense crowds requires a mobile robot to maintain an appropriate relative position with respect to a moving target while avoiding surrounding pedestrians and satisfying rear-following and social constraints. This paper proposes PPO-GAT-Follow, an interaction-aware reinforcement learning framework for dense-crowd RPF under geometric visibility loss with available target-relative pose estimates. The follower, target pedestrian, and surrounding pedestrians are represented as graph nodes, and a graph attention encoder models their local interactions. A task-oriented reward mechanism jointly accounts for target maintenance, visibility preservation, collision avoidance, proximity-aware social compliance, rear position maintenance, post-arrival stabilization, and action stability. Experiments are conducted in IR-SIM under fixed-route and random-route settings, with comparisons against MPC, DWA, SFM, and an adapted SARL baseline. In the fixed-route setting with 12 background pedestrians, PPO-GAT-Follow achieves a task success rate of 98.8% and a collision rate of 1.1%, improving task success by 10.9 percentage points over MPC. In the random-route setting at the training density, it achieves 83.1% task success and an SPL of 0.815, outperforming MPC by 18.3 percentage points in task success; at this density, it also surpasses SARL in the main task-level metrics. Zero-shot evaluations across crowd densities, together with structural and reward ablations, reward weight sensitivity analysis, tolerance shift tests, multi-seed training, and stress testing under target pose noise and heterogeneous pedestrian dynamics, further demonstrate the effectiveness and reliability of the proposed framework. Gazebo-based validation also demonstrates system integration feasibility with localization, point cloud-based surrounding pedestrian perception, tracking, and UWB-like target-relative pose input. Nevertheless, visual target identification, re-identification, and perception-level occlusion recovery remain outside the scope of the present validation. Full article
(This article belongs to the Section Sensors and Robotics)
Show Figures

Figure 1

19 pages, 2926 KB  
Article
Physics-Guided Learning for Monocular Visual Object Localization in Indoor Environments
by Haorui Ge, Luzheng Bi, Weijie Fei and Jiarong Wang
Sensors 2026, 26(15), 4674; https://doi.org/10.3390/s26154674 - 23 Jul 2026
Viewed by 172
Abstract
Accurate object localization is essential for enabling autonomous operation of indoor robotic systems. As a low-cost, compact, and flexibly deployable solution, monocular visual object localization (MVOL) is highly applicable to lightweight embedded robotic platforms. However, conventional data-driven MVOL methods suffer from inherent limitations [...] Read more.
Accurate object localization is essential for enabling autonomous operation of indoor robotic systems. As a low-cost, compact, and flexibly deployable solution, monocular visual object localization (MVOL) is highly applicable to lightweight embedded robotic platforms. However, conventional data-driven MVOL methods suffer from inherent limitations of 2D-to-3D ill-posed mapping, which is caused by the incapability of constraining the spatial physical logic of real scenes, resulting in severe depth ambiguity, inaccurate scale estimation, and physically unreasonable predictions. To address these issues, this paper proposes a novel physics-guided monocular visual localization framework termed PC-IMVL for indoor scenarios. The PC-IMVL integrates deep visual perception with embedded physical modeling, which explicitly introduces spatial physical constraints into the network optimization process and builds a physical consistency-aware loss function to regularize 3D position and pose estimation. Combined with a lightweight tailored architecture, the framework enables efficient and reliable embedded deployment. Offline experiments and real-world online tests validate the effectiveness of the proposed method. PC-IMVL yields average absolute errors (AE) of 0.095–0.333 m, reducing the localization error of early fusion methods by more than 50%. Within a working distance of 3–4 m, it achieves a relative error (RE) of 2.4% and a horizontal viewing angle error (VAE) below 2°, outperforming existing state-of-the-art MVOL methods. The effectiveness of the physical guidance mechanism is verified. This work provides a practical high-precision localization solution for embedded indoor robotic systems. Full article
(This article belongs to the Section Environmental Sensing)
Show Figures

Figure 1

38 pages, 1156 KB  
Systematic Review
From Black Box to Clarity: A Systematic Review of Explainability Methods in Deep Convolutional Neural Networks
by Zina Tayari and Mourad Zaied
Mach. Learn. Knowl. Extr. 2026, 8(8), 220; https://doi.org/10.3390/make8080220 - 23 Jul 2026
Viewed by 523
Abstract
Deep neural networks (DNNs) have significantly advanced machine perception and reasoning; however, their lack of transparency in decision-making continues to pose a major challenge, particularly in high-stakes domains such as healthcare, finance, and law. This is especially concerning with the black-box nature of [...] Read more.
Deep neural networks (DNNs) have significantly advanced machine perception and reasoning; however, their lack of transparency in decision-making continues to pose a major challenge, particularly in high-stakes domains such as healthcare, finance, and law. This is especially concerning with the black-box nature of convolutional neural networks (CNNs), where the rationale for making a decision can be as important as the decision itself. This paper is driven by a question that is easier to ask than to answer: how can CNNs be made to explain themselves? To answer the question, we wrote a PRISMA-compliant systematic review of 154 studies published between 2017 and 2025. These studies were selected from 4421 studies retrieved through Web of Science, Scopus, IEEE Xplore, and ACM Digital Library. CNN-specific taxonomy was developed. This taxonomy organizes explainable artificial intelligence (XAI) methods on four axes: explanation timing, model dependency, output type, and target component. We found that there is a huge bias in the field regarding post hoc visual methods. Grad-CAM is the most widely cited visual explanation methodology, and within the model-agnostic framework, LIME and SHAP prevail. This research was also the first to analyze standard assessment methods. It was found that out of the 154 studies in the review, 98 used objective methods to evaluate fidelity, stability, or sensitivity. Conversely, fewer than ten of them used human-centered methods to evaluate how tasks were performed, how the users trusted the method, or how the users were prepared to interact with the system. We argue for a dual-reporting convention under which metrics should be reported together at least once, as per the family of metrics. The third contribution is an evidence-based challenge map, where we outline four issues: absence of standardized benchmarks, post hoc mechanism scalability limitations, vulnerability to adversarial perturbations, and the persistent gap between the technical descriptions and human understanding. For each challenge, we propose concrete directions: integrating causal reasoning, adopting participatory evaluation design, and building hybrid transparent architectures. We offer this review as a practical roadmap for researchers and practitioners working toward more explainable deep neural networks. Full article
(This article belongs to the Section Learning)
Show Figures

Figure 1

41 pages, 21931 KB  
Article
Decoding Tourists’ Landscape Perception Preferences in Historical and Cultural Heritage Parks Through Social Media Images: A Dual-Task Deep Learning Framework
by Changzhi Zhang, Yibei Wang, Liyuan Li, Junfeng Zhao and Shitong Peng
Buildings 2026, 16(14), 2918; https://doi.org/10.3390/buildings16142918 - 22 Jul 2026
Viewed by 571
Abstract
Historical and cultural heritage parks are important spaces for heritage conservation, cultural transmission, and public recreation. However, conventional landscape perception research mainly relies on questionnaires and interviews, making it difficult to capture tourists’ visual preferences at scale. This study proposes a dual-task attention-enhanced [...] Read more.
Historical and cultural heritage parks are important spaces for heritage conservation, cultural transmission, and public recreation. However, conventional landscape perception research mainly relies on questionnaires and interviews, making it difficult to capture tourists’ visual preferences at scale. This study proposes a dual-task attention-enhanced ResNet framework based on social media user-generated content (UGC) images to investigate tourists’ landscape perception preferences in historical and cultural heritage parks. Using Yellow Crane Tower Park, Guqintai, and Guishan Scenic Area in Wuhan, China, as case studies, 6221 images were collected from Ctrip, Xiaohongshu, Weibo, and field surveys. The framework jointly performs landscape element detection and aesthetic attribute classification through shared feature representation and attention mechanisms. The proposed model achieved a composite Macro-F1 score of 0.7641, demonstrating robust classification performance. The results show that Buildings and Structures exhibited the highest average prediction probability (0.5824), while Spatial Legibility was the dominant aesthetic attribute (0.5322), indicating a perception pattern characterized by cultural-symbol prominence and enhanced spatial cognition. Vegetation and road networks were positively associated with spatial mystery, whereas excessive visual complexity reduced spatial legibility. These findings demonstrate the value of combining deep learning with social media image analytics for cultural landscape perception research and provide practical insights for landscape planning, heritage conservation, and tourism management. Full article
(This article belongs to the Section Architectural Design, Urban Science, and Real Estate)
Show Figures

Figure 1

Back to TopTop