Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (143)

Search Parameters:
Keywords = neighbor query

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
28 pages, 13063 KB  
Article
DualGLEAN: Dual Allocation for VLM-Guided Generalized Category Discovery in Remote Sensing Images
by Hongfu Li, Yuxiang Xie, Jing Zhang, Yanming Guo and Xin Zhang
Remote Sens. 2026, 18(17), 3054; https://doi.org/10.3390/rs18173054 - 7 Sep 2026
Viewed by 181
Abstract
Generalized category discovery (GCD) aims to classify known categories while discovering novel ones in unlabeled data, yet existing methods lack mechanisms to correct boundary-ambiguous samples that receive noisy pseudo-labels, as they primarily rely on visual feature learning without external semantic guidance. Vision-language models [...] Read more.
Generalized category discovery (GCD) aims to classify known categories while discovering novel ones in unlabeled data, yet existing methods lack mechanisms to correct boundary-ambiguous samples that receive noisy pseudo-labels, as they primarily rely on visual feature learning without external semantic guidance. Vision-language models (VLMs) offer a natural source of cross-modal semantic correction. However, applying VLM-guided contrastive signals directly within the GCD training loop proves counterproductive because the locally-oriented InfoNCE loss conflicts geometrically with the globally oriented K-means objective in the shared backbone space. We identify the root cause as a dual resource allocation problem: the VLM-derived signal must be allocated to the correct feature subspace to avoid geometric conflict with K-means clustering (space allocation), and the limited VLM inference budget must be allocated to the correct samples to maximize discriminative return (budget allocation). These two decisions are coupled; failure on either renders the other ineffective. To resolve this, we propose DualGLEAN, a framework that addresses the dual allocation challenge through two coupled mechanisms: decoupled contrastive alignment (DCA), which routes the VLM-guided neighbor contrastive loss to a dedicated projector space while preserving the backbone space for global clustering, and compound uncertainty querying (CUQ), a three-stage filtering metric that jointly evaluates predictive entropy, boundary proximity, and local label inconsistency to direct VLM queries exclusively to truly boundary-critical samples. Extensive experiments on the AID and RSSDIVCS datasets demonstrate that DualGLEAN achieves strong performance, improves four diverse GCD baselines as a plug-in module, generalizes across seven VLM backbones, introduces zero additional trainable parameters to the base GCD network, and incurs a total VLM API cost of only CNY 2.45 per full training run on the AID dataset under the default search-scope configuration, with the cost scaling linearly with the query budget. Full article
Show Figures

Figure 1

17 pages, 7140 KB  
Article
Spatial-Aware Modulation for Implicit Neural Representations
by Chen Qing, Wenxin Zhang, Haoyu Wang, Dongshen Han, Mingming Zhang and Caiyan Qin
Appl. Sci. 2026, 16(17), 8870; https://doi.org/10.3390/app16178870 - 7 Sep 2026
Viewed by 139
Abstract
Implicit Neural Representations (INRs) provide a flexible and resolution-independent formulation for continuous signal representation. Despite their strong representation ability, standard INRs usually predict each queried coordinate independently, making it difficult to explicitly exploit the local coherence widely observed in natural signals. For images [...] Read more.
Implicit Neural Representations (INRs) provide a flexible and resolution-independent formulation for continuous signal representation. Despite their strong representation ability, standard INRs usually predict each queried coordinate independently, making it difficult to explicitly exploit the local coherence widely observed in natural signals. For images and volumetric data, neighboring locations often share correlated responses in smooth regions, while sharp variations mainly appear around spatial transitions. Ignoring such local dependency may reduce learning efficiency and weaken the reconstruction of spatially consistent details. To address this limitation, we propose Spatial-Aware Implicit Neural Representation (SA-INR), which enhances INRs by introducing local feature aggregation into the hidden representation space. Motivated by local feature coherence, SA-INR aggregates neighboring coordinate features through a learnable spatial-aware local operator. The aggregation weights are initialized as a uniform mean filter, providing a smooth local bias during early optimization. As training proceeds, the aggregation weights are updated by reconstruction supervision and become adaptive to spatial content. To preserve coordinate-specific information and avoid over-smoothing, the aggregated feature is further integrated with the original feature through a residual connection. Extensive experiments on image representation, CT reconstruction, and image denoising demonstrate that SA-INR consistently improves reconstruction fidelity across different INR backbones and reconstruction tasks. These results suggest that explicitly modeling local feature interaction is an effective way to enhance continuous signal representation. Full article
Show Figures

Figure 1

20 pages, 12123 KB  
Article
GFE-Net: Geometry-Enhanced Feature Extraction Network for Semantic Segmentation of Large-Scale LiDAR Point Clouds
by Hui Liu, Guangming Zhang, Chuang Chen and Zhihan Shi
Remote Sens. 2026, 18(17), 2990; https://doi.org/10.3390/rs18172990 - 3 Sep 2026
Viewed by 234
Abstract
Accurate semantic segmentation of large-scale outdoor LiDAR point clouds remains a challenging endeavor, primarily due to ambiguous class transitions at object interfaces, non-uniform sampling density across the surveyed area, and shared geometric signatures among distinct object categories. This paper proposes GFE-Net (Geometry-Enhanced Feature [...] Read more.
Accurate semantic segmentation of large-scale outdoor LiDAR point clouds remains a challenging endeavor, primarily due to ambiguous class transitions at object interfaces, non-uniform sampling density across the surveyed area, and shared geometric signatures among distinct object categories. This paper proposes GFE-Net (Geometry-Enhanced Feature Extraction Network), a hierarchical encoder–decoder architecture that systematically improves per-point feature characterization through three complementary design contributions: First, to mitigate the shortcomings of conventional fixed-neighborhood queries in regions of variable point density, a Structure-Guided Neighborhood Adaptation (SGNA) module is devised. At its core lies a morphology-driven contextual gating (MCG) unit that synthesizes neighbor-wise calibration weights from hierarchical shape descriptors fused with elevation difference statistics, allowing the network to preferentially amplify morphologically congruent neighbors while dampening spurious or cross-boundary contributions. Second, to strengthen semantic discrimination beyond what spatial locality alone affords, a Local–Global Interactive Enhancement (LGIE) module is presented. The LGIE module simultaneously distills precise local structure through Euclidean-space neighborhood graphs and captures scene-wide co-activation patterns via compact bilinear factorization of the latent feature space, merging both streams through a residual refinement mechanism that markedly improves inter-class separability. Third, to enforce label consistency at object interfaces without relying on post-processing heuristics, a Neighborhood Prediction Consistency (NPC) loss is introduced. Built upon a Gaussian distance-decay weighting kernel, the NPC loss assigns progressively stronger penalties to label mismatches between a query point and its geometrically proximate neighbors, thereby promoting spatially coherent predictions and attenuating boundary noise. GFE-Net is rigorously benchmarked on two widely adopted large-scale datasets—S3DIS and SensatUrban—yielding OA/mIoU of 89.6%/73.1% and 93.3%/61.1%, respectively. These results demonstrate competitive performance under the reported protocols. Detailed ablation studies and computational profiling further substantiate the efficacy of each individual component. Full article
(This article belongs to the Section Remote Sensing Image Processing)
Show Figures

Figure 1

30 pages, 4824 KB  
Article
A Distributed Storage and Indexing Framework Based on Hierarchical Space-Time Grid Encoding for Digital Cities
by Huangchuang Zhang, Weiming Xing and Kai Zhang
ISPRS Int. J. Geo-Inf. 2026, 15(9), 398; https://doi.org/10.3390/ijgi15090398 - 1 Sep 2026
Viewed by 258
Abstract
The rapid prolife ration of digital twin cities has resulted in an unprecedented increase in the volume, heterogeneity, and dynamics of space-time data, posing significant challenges to scalable storage, efficient indexing, and high-performance query processing in distributed environments. Existing distributed frameworks generally treat [...] Read more.
The rapid prolife ration of digital twin cities has resulted in an unprecedented increase in the volume, heterogeneity, and dynamics of space-time data, posing significant challenges to scalable storage, efficient indexing, and high-performance query processing in distributed environments. Existing distributed frameworks generally treat data partitioning and indexing as independent processes, making it difficult to simultaneously preserve space-time locality, achieve balanced data distribution, and support efficient multidimensional queries. To address these challenges, this paper proposes a distributed space-time storage and indexing framework based on hierarchical space-time grid encoding. The proposed framework integrates Hadoop distributed file system (HDFS) and HBase to establish a unified architecture for data partitioning, storage organization, indexing, and query processing. Specifically, a space-time grid–based partitioning strategy is designed to preserve space-time locality while maintaining load balance across distributed storage nodes. Furthermore, a hybrid interleaved space-time encoding scheme together with an optimized HBase RowKey is developed to construct a unified space-time index, enabling efficient space-time range queries and space-time k-nearest neighbor (KNN) queries in large-scale distributed environments. Extensive experiments conducted on the T-Drive trajectory dataset demonstrate that the proposed framework consistently outperforms several representative distributed space-time indexing approaches in terms of query efficiency, scalability, and overall system performance while maintaining stable indexing performance under increasing data volumes. The proposed framework provides an efficient and scalable solution for the management and retrieval of massive space-time data and offers a practical infrastructure for digital twin city applications. Full article
(This article belongs to the Special Issue Urban Digital Twins Empowered by AI and Dataspaces)
Show Figures

Figure 1

24 pages, 19703 KB  
Article
STG: Structured Topology of Gridpoints for Occluded Pedestrian Detection
by Tian Qiu, Jifeng Shen and Xin Zuo
Sensors 2026, 26(17), 5496; https://doi.org/10.3390/s26175496 - 30 Aug 2026
Viewed by 226
Abstract
Pedestrian detection in crowds is a challenging problem in computer vision. Existing occlusion-handling methods heavily rely on expensive visible-box annotations to locate visible body parts, posing severe limitations in label acquisition cost and open-world generalization. To break through this limitation, we propose a [...] Read more.
Pedestrian detection in crowds is a challenging problem in computer vision. Existing occlusion-handling methods heavily rely on expensive visible-box annotations to locate visible body parts, posing severe limitations in label acquisition cost and open-world generalization. To break through this limitation, we propose a novel Structured Topology of Gridpoints (STG) framework. Operating strictly under standard full-box annotations without any extra visibility supervision, STG aims to achieve implicit, fine-grained local semantic compensation. Specifically, we formulate a coarse-to-fine reasoning paradigm consisting of three interactive stages. To mitigate the high spatial complexity and eliminate background redundancy, we first introduce a Saliency-Aware Feature Filtering (SAFF) mechanism, which leverages gridpoint heatmaps to filter out low-confidence pedestrian candidates. Second, a query-guided Across-Instance Feature Interaction (AIFI) model is designed to utilize inter-instance spatial relationships to propagate missing context from highly visible individuals to their occluded neighbors. Finally, we devise a prior-guided Inner-Instance Gridpoints Interaction (I2GI) model to achieve fine-grained structured part-level feature completion, which dynamically aggregates vital localized cues from diverse human parts to reconstruct holistic pedestrian representations. Extensive experiments on the CityPersons, CrowdHuman, and WiderPerson datasets demonstrate the effectiveness and efficiency of our proposed method. Specifically, STG achieves a log-average miss rate of 7.41% on Reasonable and 32.05% on Heavy Occlusion subsets of CityPersons, while running at up to 16 FPS, outperforming existing part-based methods under full-box supervision. Full article
(This article belongs to the Special Issue Image Processing and Analysis for Object Detection: 3rd Edition)
Show Figures

Figure 1

35 pages, 2448 KB  
Article
Neural Signed Distance Surrogates for Cut-Cell Finite-Volume Solvers
by Ammar Qarariyah and Suhail Odeh
Computation 2026, 14(9), 198; https://doi.org/10.3390/computation14090198 - 25 Aug 2026
Viewed by 246
Abstract
Cut-cell and ghost-cell finite-volume methods require an accurate boundary distance field, available in closed form only for simple shapes and otherwise obtained through costly nearest-neighbor search. Prior work embeds learned distance fields as an auxiliary component inside a larger neural architecture; this study [...] Read more.
Cut-cell and ghost-cell finite-volume methods require an accurate boundary distance field, available in closed form only for simple shapes and otherwise obtained through costly nearest-neighbor search. Prior work embeds learned distance fields as an auxiliary component inside a larger neural architecture; this study instead tests the numerical consequences of substituting the distance function alone within an otherwise unmodified classical discretization. We propose a compact trained neural network as a drop-in surrogate for this field, evaluated against exact and classical alternatives, including a KD-tree baseline, on four problems of increasing difficulty ending with a three-dimensional mechanical flange. A truncation analysis bounds the surrogate’s intercept error and identifies the training accuracy needed to preserve the scheme’s formal order. The surrogate matches second-order accuracy in every example, with fitted orders of 2.0 to 2.1 in two dimensions and, in three dimensions, a wider but still order-consistent 1.9 to 2.5, the extra width traced to identified, shared mesh- and timestep-resolution effects rather than to the surrogate itself. Query cost overtakes the KD-tree beyond roughly one hundred thousand points, with speedups up to 26.1 times, and gradient fields are substantially smoother than nearest-point construction throughout, tracking geometric severity in two dimensions and feature scale in three. Full article
(This article belongs to the Section Computational Intelligence)
Show Figures

Figure 1

21 pages, 710 KB  
Article
Hard-Negative Prototype Rectification for Low-Support Cervical Cytology Classification
by Mehret Ephrem Abraha and Juntae Kim
Electronics 2026, 15(15), 3416; https://doi.org/10.3390/electronics15153416 - 2 Aug 2026
Viewed by 232
Abstract
Reliable cervical cytology classification remains difficult when rare diagnostic categories are represented by only a few labeled examples and exhibit substantial morphological overlap with neighboring classes. This study introduces HardNegRect, a lightweight inductive prototype-rectification module designed for low-support classification among fixed cervical cytology [...] Read more.
Reliable cervical cytology classification remains difficult when rare diagnostic categories are represented by only a few labeled examples and exhibit substantial morphological overlap with neighboring classes. This study introduces HardNegRect, a lightweight inductive prototype-rectification module designed for low-support classification among fixed cervical cytology categories. The method constructs a class-specific hard-negative reference from the most similar competing prototypes and uses this inter-class context to predict a bounded, gated residual correction to each support-derived prototype. Because rectification depends exclusively on support information, the approach preserves independent query processing and avoids transductive access to the test distribution. HardNegRect was evaluated on two public cervical cytology benchmarks using common fold assignments, support sizes, held-out query sets, and draw-level metric aggregation for frozen-feature, metric-based, and optimization-based comparators. The study also includes a controlled component ablation study, a neighborhood-sensitivity analysis, and an additional-seed stability analysis. On Mendeley LBC, the clearest benefit occurred in the one-shot setting, where HardNegRect achieved a Macro-F1 of 0.9862±0.0062 and an SCC F1 of 0.9655±0.0216. On SIPaKMeD, the default Khn=2 configuration achieved Macro-F1 values of 0.9596±0.0052, 0.9626±0.0048, and 0.9616±0.0048 for K=1,3,10, respectively, numerically exceeding the strongest comparator mean at each support size. The controlled component ablation study associates the additional one-shot gain on Mendeley LBC with inter-class prototype correction rather than with embedding transformation alone. Overall, HardNegRect provides a lightweight, parameter-efficient, and geometry-aware extension to prototype-based low-support cytology classification, while patient-grouped, source-grouped, repeated-seed, and multi-center validation remain necessary before clinical generalization can be established. Full article
(This article belongs to the Special Issue Feature Papers in Bioelectronics: 2025–2026 Edition)
Show Figures

Figure 1

33 pages, 10661 KB  
Article
Memory Pollution in Multi-Product Visual Anomaly Detection: Diagnosis and Mitigation
by Sergio Villanueva López, Emilio Soria-Olivas and Manuel Sánchez-Montañés
Mach. Learn. Knowl. Extr. 2026, 8(7), 219; https://doi.org/10.3390/make8070219 - 22 Jul 2026
Viewed by 1273
Abstract
Memory-bank methods such as PatchCore are widely used in industrial quality control for visual anomaly detection since they require no training, are fast to deploy and achieve strong accuracy. However, they are memory-intensive. Furthermore, a single production line typically involves different products or [...] Read more.
Memory-bank methods such as PatchCore are widely used in industrial quality control for visual anomaly detection since they require no training, are fast to deploy and achieve strong accuracy. However, they are memory-intensive. Furthermore, a single production line typically involves different products or cameras, so using a single anomaly detection method with a shared nearest-neighbor memory bank is attractive since it simplifies deployment and makes new products easy to add. Nevertheless, embeddings from different products/cameras can interfere during retrieval, causing what we call “memory pollution”. In this work, we study this effect through a new diagnostic framework, which involves: (1) a new metric, the wrong-neighbor rate (WNR), which measures how often a query’s nearest neighbor belongs to a different product; (2) an empirically validated phenomenon, “oracle inversion”, where querying only the product’s own data can underperform the shared bank under a fixed memory budget; (3) a first-order analytical model of the WNR, which predicts how pollution grows with product count and memory budget; and (4) a minimal training-free router that removes the effect of memory pollution. Our results show that our system performs robustly across different datasets and backbones, with up to 25× memory reduction, which makes our framework attractive for industrial applications. Full article
(This article belongs to the Section Learning)
Show Figures

Figure 1

19 pages, 5204 KB  
Article
Entropy-Driven Action Randomization: A Deployment-Time Defense for Environment Privacy in Deep Reinforcement Learning
by Xin Cheng, Jinchuan Tang and Shuping Dang
Electronics 2026, 15(14), 2998; https://doi.org/10.3390/electronics15142998 - 8 Jul 2026
Viewed by 339
Abstract
Deep reinforcement learning (DRL) agents implicitly memorize the structure of their training environment, allowing an attacker to reconstruct it by querying the action outputs of a deployed model and causing environment-privacy leakage. To address this, this study proposes Entropy-Driven Action Randomization (EDAR), a [...] Read more.
Deep reinforcement learning (DRL) agents implicitly memorize the structure of their training environment, allowing an attacker to reconstruct it by querying the action outputs of a deployed model and causing environment-privacy leakage. To address this, this study proposes Entropy-Driven Action Randomization (EDAR), a deployment-time action randomization defense. Inspired by the exponential mechanism from differential privacy, EDAR replaces deterministic greedy action selection with probabilistic action sampling based on a normalized Q-value utility function, without altering the trained policy. A state-adaptive dynamic privacy budget is further designed, guided by the Shannon entropy of the action distribution. A simulated-annealing-based policy inversion attack is used to quantify privacy leakage, measured by the normalized recovery rate (NRR). Experiments on GridWorld environments ranging from 7 × 7 to 13 × 13 show that, at a privacy level comparable to a strong fixed budget, the proposed dynamic budget raises the reward retention from 6.2% to 78.9% while keeping a comparable NRR reduction. These results indicate that adapting the privacy budget to per-state policy determinism yields a markedly better privacy–utility trade-off than existing training-perturbation-based methods. We clarify that the differential privacy property invoked here is a per-query, mechanism-level guarantee of indistinguishability between neighboring observation states under a fixed trained model. The protection of the environment structure itself is established empirically through the policy inversion attack. Experiments are conducted on discrete GridWorld navigation tasks; the conclusions are scoped to small value-based navigation settings. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Figure 1

36 pages, 1711 KB  
Article
GeoIR-Compiler: A Geospatial Intermediate Representation and Compilation Framework for Chinese Urban Spatial Question Answering
by Chaolin Zhang, Jiqiu Deng, Hui Zhang, Longbo Li, Liji Sun and Xiao Ma
ISPRS Int. J. Geo-Inf. 2026, 15(7), 310; https://doi.org/10.3390/ijgi15070310 - 8 Jul 2026
Viewed by 706
Abstract
Natural-language access to spatial databases requires relation interpretation, entity grounding, metric normalization, and database-specific execution semantics. Direct generation of Structured Query Language (SQL) by large language models (LLMs) can therefore return executable but spatially wrong SQL, especially for Chinese urban questions with aliases, [...] Read more.
Natural-language access to spatial databases requires relation interpretation, entity grounding, metric normalization, and database-specific execution semantics. Direct generation of Structured Query Language (SQL) by large language models (LLMs) can therefore return executable but spatially wrong SQL, especially for Chinese urban questions with aliases, abbreviated place names, and geometry-dependent predicates. This paper presents GeoIR-Compiler, a spatially specialized framework that maps a Chinese question to a typed geospatial intermediate representation (GeoIR), grounds mentions and attributes to database objects, and deterministically compiles the grounded representation into SQL for PostGIS, a spatial database extension for PostgreSQL. The contribution is the specialization of intermediate representations for Chinese urban spatial question answering through explicit spatial relations, metric constraints, grounding records, and PostGIS execution templates. We construct two controlled executable benchmarks, NJ-GeoIR-700 and WH-GeoIR-700, covering retrieval, topology, distance, nearest-neighbor, aggregation, compositional, and alias/noisy-mention queries. Across seven locally served backbones, GeoIR-Full reaches mean execution accuracies of 0.7271 on Nanjing and 0.7363 on Wuhan, outperforming Direct-SQL, Data-Augmented In-Context Learning (DAIL)-SQL-style, and Linking-SQL under the fixed evaluation protocol. Ablations are consistent with grounding contributing strongly to the observed gains, while verification mainly trades coverage for answer reliability. Full article
Show Figures

Figure 1

27 pages, 602 KB  
Article
NNFDA: A Digest-Based Integrity Verification Scheme for Enhancing Secure Queries in Loss-Tolerant TMWSNs
by Peng Li, Weipeng Wang, Wenxin Yang and Yang Pei
Electronics 2026, 15(13), 2950; https://doi.org/10.3390/electronics15132950 - 6 Jul 2026
Viewed by 321
Abstract
Tiered Mobile Wireless Sensor Networks (TMWSNs), consisting of mobile sensor nodes and storage nodes, are widely used in various fields due to their scalability, energy efficiency, and flexibility. Most existing secure query algorithms assume that data packets generated by sensor nodes can always [...] Read more.
Tiered Mobile Wireless Sensor Networks (TMWSNs), consisting of mobile sensor nodes and storage nodes, are widely used in various fields due to their scalability, energy efficiency, and flexibility. Most existing secure query algorithms assume that data packets generated by sensor nodes can always be delivered to storage nodes. This assumption does not hold in practice, where packets may be lost due to attacks or adverse communication conditions. This paper proposes a loss-tolerant wireless network model for TMWSNs and a novel threat model tailored to this scenario, in which packet-dropping attacks compromise the integrity of query results. To counter these attacks, we present a baseline integrity verification algorithm, the Neighbor Node-Forwarding Digest Algorithm (NNFDA). Each sensor generates a digest of its data and forwards it to neighboring nodes. These digests are then transmitted to storage nodes together with the neighbors’ data, thereby establishing a chained relationship among sensor data. The base station verifies query results using this relationship. The baseline algorithm, however, causes high communication overhead. To reduce this cost, we propose an improved version, NNFDA-BM (NNFDA with Bitmap), which optimizes digest generation and transmission. Experimental results show that NNFDA-BM verifies query result integrity effectively while achieving a significant reduction in communication overhead compared with the baseline algorithm. Full article
(This article belongs to the Special Issue Novel Methods Applied to Security and Privacy Problems, Volume II)
Show Figures

Figure 1

22 pages, 1564 KB  
Article
Multi-Hop Trajectory Prediction of Aircraft Taxiing Using Spatio-Temporal Knowledge Graph with Vector-Index Support
by Jing Shan, Jianan Yin, Beijing Zhou and Minghua Hu
Electronics 2026, 15(12), 2613; https://doi.org/10.3390/electronics15122613 - 12 Jun 2026
Viewed by 352
Abstract
Efficient multi-hop prediction over large-scale spatio-temporal knowledge graphs of aircraft taxiing trajectories remains challenging, as existing methods focus either on static multi-hop relations or on accuracy improvement for spatio-temporal single-hop predictions, leading to computational inefficiency. This paper proposes a vector-index-supported multi-hop prediction method. [...] Read more.
Efficient multi-hop prediction over large-scale spatio-temporal knowledge graphs of aircraft taxiing trajectories remains challenging, as existing methods focus either on static multi-hop relations or on accuracy improvement for spatio-temporal single-hop predictions, leading to computational inefficiency. This paper proposes a vector-index-supported multi-hop prediction method. First, a knowledge graph embedding technique that integrates spatio-temporal features maps the trajectory graph into a low-dimensional complex vector space. Then, a hierarchical query acceleration structure based on IndexIVFFlat is constructed. A clustering strategy guided by the distribution of trajectory data partitions the vector space into subspaces, and approximate nearest neighbor search within those subspaces rapidly prunes the candidate set to accelerate multi-hop retrieval. Experiments on real aircraft taxiing trajectory datasets and general benchmarks show that the proposed method substantially improves prediction efficiency while maintaining competitive accuracy. The results demonstrate that the vector index mechanism effectively balances accuracy and efficiency, and the efficiency has been improved by at least 56.65%. This work provides a key technical foundation for real-time analysis and intelligent prediction of large-scale aircraft taxiing trajectories. Full article
Show Figures

Figure 1

28 pages, 2738 KB  
Article
BCAR-Net: A Bidirectional Cross-Attention Network with Auxiliary Reconstruction for Tree Counting in Complex Forest Scenes Using Airborne RGB and LiDAR Data
by Xiaoyu Wu, Xijian Fan, Mengjiao Tang and Size Dai
Plants 2026, 15(12), 1762; https://doi.org/10.3390/plants15121762 - 6 Jun 2026
Viewed by 1029
Abstract
Accurate tree counting from remote sensing data is essential for forest inventory, biomass estimation, carbon accounting, and ecological monitoring. However, existing approaches predominantly rely on airborne RGB imagery and often struggle in complex forest scenes where neighboring crowns exhibit highly similar textures and [...] Read more.
Accurate tree counting from remote sensing data is essential for forest inventory, biomass estimation, carbon accounting, and ecological monitoring. However, existing approaches predominantly rely on airborne RGB imagery and often struggle in complex forest scenes where neighboring crowns exhibit highly similar textures and colors and where overlapping crown boundaries become ambiguous. To address this limitation, the LiDAR-derived Canopy Height Model (CHM) is introduced as a complementary modality that provides explicit cues on canopy height variation and vertical structure to support RGB-based analysis. Building on this, we propose BCAR-Net, a broker-guided RGB and depth (RGB-D) multimodal framework that couples bidirectional cross-modal interaction, adaptive tri-branch fusion, and auxiliary reconstruction within a two-stage optimization scheme. Specifically, a bidirectional cross-attention U-Net generates an intermediate broker RGB-D representation from paired RGB images and depth maps through symmetric bidirectional cross-attention between the two modalities and direction-aware gating. The original RGB image, depth map, and broker representation are then jointly encoded by three weight-sharing branches and adaptively aggregated by a spatial fusion gate for density-map regression. To regularize the fused latent feature, a multi-scale cross-attention reconstruction decoder provides auxiliary RGB and depth reconstruction supervision by querying multi-scale BCA-UNet encoder features through 2D cross-attention, and a reconstruction-oriented first stage replaces externally generated fused-image supervision, yielding a task-consistent optimization scheme. Experiments on the NEONTreeEvaluation benchmark show that BCAR-Net consistently outperforms single-modality settings and direct RGB-D concatenation multimodal baseline. Additional experiments on a public UAV RGB–LiDAR dataset provide a small-scale supplementary evaluation under a different acquisition setting, where BCAR-Net achieves modest but consistent improvements over RGB-only and depth-only baselines. These results demonstrate that the proposed framework offers an effective but computationally cautious solution for tree counting in complex forest environments. Full article
(This article belongs to the Special Issue Computer Vision Techniques for Plant Phenomics Applications)
Show Figures

Figure 1

23 pages, 4180 KB  
Article
PGformer: Fusing Kernelized Transformers and GCNs for Automated Proximity Graph Parameter Configuration
by Fangyi Shen, Zhentao Zhan, Junjie Wu and Xiaoliang Xu
Information 2026, 17(6), 561; https://doi.org/10.3390/info17060561 - 5 Jun 2026
Viewed by 425
Abstract
Approximate nearest neighbor search (ANNS) serves as the fundamental querying method in large-scale and high-dimensional vector datasets, for which proximity graph (PG)-based algorithms are the preferred solution, offering the best balance between query efficiency and accuracy. However, PG-based algorithms’ optimal performance across multiple [...] Read more.
Approximate nearest neighbor search (ANNS) serves as the fundamental querying method in large-scale and high-dimensional vector datasets, for which proximity graph (PG)-based algorithms are the preferred solution, offering the best balance between query efficiency and accuracy. However, PG-based algorithms’ optimal performance across multiple indicators depends on extensive manual parameter configuration. To this end, we propose PGformer, an end-to-end framework that predicts performance to identify optimal graph configurations. PGformer integrates attention-driven global representation learning and neighbor-aware embedding extraction to capture comprehensive structural features. For large-scale scenarios, we implement a linear-complexity attention mechanism that maintains global accuracy while reducing computational overhead. These learned representations are then used to jointly estimate multiple performance metrics, facilitating precise candidate selection. The experimental results show that PGformer achieves competitive and often improved recommendation quality compared with traditional baselines, particularly in large-scale settings. In the evaluated scenarios, PGformer reduces recommendation time by about 40% and selects configurations that are close to the exhaustive-search reference optimum under recall-constrained objectives. Nevertheless, PGformer still relies on offline profiling data and a discretized candidate configuration space, which motivates further investigation under broader deployment environments and stronger distribution shifts. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Figure 1

20 pages, 13595 KB  
Article
POI-Guided Heuristic Mapping for UAV Motion Planning with Bounded Distance Updates
by Yong Li, Lihui Wang, Xueyong Xu, Renzhi Huang and Yuhang Xu
Drones 2026, 10(5), 332; https://doi.org/10.3390/drones10050332 - 29 Apr 2026
Viewed by 501
Abstract
Safety-oriented UAV motion planning relies on distance-to-obstacle fields and their gradients, yet onboard mapping is typically limited to bounded local distance updates. Consequently, optimization may stall outside the updated band due to missing gradients, while enlarging the update range substantially increases computational cost. [...] Read more.
Safety-oriented UAV motion planning relies on distance-to-obstacle fields and their gradients, yet onboard mapping is typically limited to bounded local distance updates. Consequently, optimization may stall outside the updated band due to missing gradients, while enlarging the update range substantially increases computational cost. Our key insight is that motion-planning locality implies only a small subset of obstacles governs local trajectory refinement. We term this subset points of interest (POIs). Motivated by this observation, we develop a locality-aware sequential motion planning framework with a POI-driven feedback mechanism that continuously identifies and augments these trajectory-relevant obstacles during search and optimization. The mechanism tightly couples mapping, search, and optimization and enables safe trajectory refinement without requiring global distance updates. The framework adopts a heuristic mapping strategy that combines a long-term occupancy grid with bounded incremental distance updates and a POI-based short-term k-d tree, enabling efficient nearest-neighbor queries and gradient proxies beyond the update band. The search process generates a dynamically feasible initial trajectory in the long-term map while collecting POIs, which are then used to construct the short-term component. The trajectory is subsequently refined through iterative optimization loops, where newly exposed closest obstacles are incorporated into the POI set and the short-term map is updated until convergence. Safety is enforced through conservative collision checking against the inflated long-term occupancy map. Simulations in building and forest environments show that 99.7% of trials converge within two refinements in sparse scenes and none exceed four overall. Compared with FastPlanner and EgoPlanner, the proposed method achieves consistently larger obstacle clearances. Onboard experiments further validate its practicality under real sensing and computational constraints. Full article
Show Figures

Figure 1

Back to TopTop