Highlights
What are the main findings?
- RoadMark-AWAConv improves fine-grained segmentation of complex road marking structures.
- Geometry-constrained weight anchors improve local geometric feature extraction under non-uniform point distributions.
What are the implications of the main findings?
- The proposed method provides a geometry-adaptive solution for fine-grained road marking point cloud segmentation.
- It can support LiDAR-based road perception and high-definition map updating.
Abstract
Road marking point cloud segmentation is essential for autonomous driving perception and high-definition map updating. However, road markings are typically represented by elongated, sparse, and irregular point structures, while non-uniform point density, occlusions, and pavement noise further increase the difficulty of fine-grained segmentation. To address these problems, we propose RoadMark-AWAConv, an adaptive weight-anchor convolution method for fine-grained road marking segmentation. The method introduces a geometry-constrained annular-domain anchor initialization strategy, hierarchical radius-based neighborhood aggregation, and normal vector direction calibration to better capture local geometric features and improve robustness to non-uniform density, occlusions, and pavement noise. Experimental results on a road marking point cloud dataset containing 13 semantic classes show that RoadMark-AWAConv achieves 71.94% mIoU, outperforming PointNet++, RandLA-Net, Point Transformer V3, and DeLA.
1. Introduction
Road markings are fundamental road-surface elements that encode traffic rules and provide essential guidance for road users [1]. Accurate recognition of road markings is critical for autonomous vehicles and advanced driver assistance systems, as it directly supports lane-boundary perception, path planning, driving-rule understanding, and traffic risk mitigation. Compared with image-based perception, LiDAR point clouds provide high-precision and explicit three-dimensional geometric information while being less affected by illumination variations and adverse weather conditions [2]. Accurate and spatially explicit representations are also important for applications such as impervious-surface mapping [3], where reliable spatial information provides a basis for subsequent analysis. Within LiDAR acquisition specifically, recent work on bundle adjustment has improved point cloud quality and mapping robustness in complex environments [4], providing a more reliable geometric foundation for downstream tasks such as semantic segmentation. Therefore, road marking segmentation from LiDAR point clouds has become an important approach for acquiring fine-grained semantics of road markings. The resulting road marking point clouds not only provide reliable lane-level guidance for autonomous driving systems but also support the dynamic updating of high-definition maps [5], thereby improving the safety and reliability of vehicle perception and decision-making.
Despite recent progress, road marking point cloud segmentation remains challenging because of the distinctive geometric properties of road markings [6]. Road markings are typically represented by elongated, sparse, discrete, and locally incomplete point structures, while LiDAR point clouds often exhibit non-uniform density distributions. Vehicle occlusion, pavement wear, surface variation, and scanning noise may further weaken the continuity of local geometric features, although recent control-theoretic advances in LiDAR odometry accuracy offer a partial remedy at the acquisition stage [7]. In addition, different marking categories may have similar colors or local shapes, making fine-grained semantic discrimination difficult [8].
Existing point cloud semantic segmentation methods can be broadly divided into point-based, projection-based, and voxel-based methods. Point-based methods directly learn local features from raw point clouds and preserve the original three-dimensional geometry, but their computational cost can be high in large-scale outdoor scenes [9]. Projection-based methods enable efficient inference by transforming point clouds into two-dimensional representations; however, the projection process inevitably loses part of the original geometric information and may cause feature fragmentation for thin and elongated markings. Voxel-based methods provide a balance between efficiency and accuracy, yet voxel quantization may remove fine boundary details, and fixed grid structures are not flexible enough to represent irregular road marking patterns.
Conventional point convolution usually relies on fixed or uniformly initialized kernels. Such kernels cannot readily adjust their spatial sampling positions to different marking shapes, although useful geometric information is often concentrated around narrow boundaries, turning regions, and directional structures. Neighborhood aggregation based on a fixed number of nearest points is also sensitive to point-density variation. A neighborhood may cover an excessively large spatial region in sparse areas but only a small region in dense areas, leading to inconsistent local feature extraction. In addition, pavement undulation, noise, and occlusion may affect local normal estimation, which can reduce the stability of convolution when modeling elongated and planar road markings.
To address the above problems, we propose RoadMark-AWAConv, an adaptive weight-anchor convolution method designed for the geometric properties and density characteristics of road marking point clouds. The geometry-constrained annular-domain initialization places more anchors near the boundary of the local region to improve the representation of elongated contours and irregular structures. Hierarchical radius-based neighborhood aggregation controls the spatial extent of local feature extraction under non-uniform point density, while normal vector calibration provides a more stable orientation for anchor convolution in noisy or incomplete regions.
The main contributions of this study are summarized as follows:
- (1)
- An adaptive weight-anchor convolution method, named RoadMark-AWAConv, is proposed for fine-grained road marking point cloud segmentation. Learnable spatial anchors are introduced into local feature extraction to better represent elongated, sparse, and irregular road marking structures.
- (2)
- A geometry-constrained annular-domain anchor initialization strategy and hierarchical radius-based neighborhood aggregation are introduced. The boundary-dense and interior-sparse anchor distribution strengthens the representation of marking contours and directional structures, while hierarchical radius-based grouping reduces the influence of non-uniform point density on local feature extraction.
- (3)
- A normal vector calibration strategy is designed to provide more stable directional information for weight-anchor convolution under pavement variation, noise, and incomplete observations. Comparative experiments and ablation studies demonstrate the effectiveness of the proposed components for fine-grained road marking segmentation.
2. Related Work
2.1. Point Cloud Datasets
Existing point cloud datasets cover indoor scenes, object-level shape analysis, and large-scale outdoor environments. S3DIS, ScanNet, ModelNet, and ShapeNet [10,11,12,13] have been widely used for general three-dimensional understanding, while SemanticKITTI, Semantic3D, Paris-Lille-3D, and Toronto-3D [14,15,16,17] provide benchmarks for outdoor road-scene segmentation. However, these datasets are mainly designed for general-purpose scene understanding, and road markings are usually represented using coarse semantic labels.
Road-marking-oriented datasets provide more specific annotations. DeepLane, RoadNet, K-Lane, and other lane mapping benchmarks [18,19,20] focus mainly on lane boundaries and lane-line structures, while RdmkNet and Toronto-RDMK [21] further support road marking classification and segmentation. Jiang et al. introduced RailPC, a large-scale railway point cloud semantic segmentation dataset that further enriches outdoor LiDAR benchmarks for transportation-related scene understanding [22]. These datasets provide useful benchmarks for LiDAR-based road marking perception, but fine-grained road marking classes still contain elongated, sparse, and geometrically similar structures that place high demands on local geometric feature extraction.
2.2. Road Marking Extraction and Segmentation from LiDAR Point Clouds
Early LiDAR-based lane and road marking extraction methods mainly relied on the high reflectance of painted road surfaces. Candidate lane or marking points were commonly identified using manually selected reflectance thresholds and subsequently grouped according to their spatial distribution. Hernández et al. [23], for example, used DBSCAN to cluster high-reflectance points and generate continuous lane markings. Although these approaches are simple and computationally efficient, their performance depends strongly on threshold selection. Changes in pavement material, marking wear, sensor configuration, and acquisition distance can result in unstable reflectance values and reduce their adaptability to complex road scenes.
With the development of deep learning, data-driven approaches have gradually become an important direction for LiDAR-based lane and road marking perception. Bai et al. [18] proposed a multi-sensor lane detection network that combines bird’s-eye-view representations generated from LiDAR point clouds with front-view camera images. The complementary information from the two sensors improves lane detection under different traffic conditions. Martinek et al. [19] designed a convolutional neural network based on BEV representations generated from LiDAR data and evaluated reference-lane generation in highway scenes.
Subsequent studies extended LiDAR-based road marking perception from lane detection to the extraction and classification of multiple road marking types. Mi et al. [24] proposed a two-stage top-down method for extracting and modeling 12 types of road markings from mobile laser scanning point clouds. Du et al. [21] introduced a multi-level feature optimization network for road marking classification and segmentation in complex urban environments. These studies demonstrate the potential of deep learning for learning semantic and geometric information from road marking point clouds.
Nevertheless, road markings are often represented by thin, sparse, and locally discontinuous point structures. Their local features can also be affected by non-uniform point density, pavement noise, marking wear, and vehicle occlusion. Existing extraction and segmentation methods may therefore produce fragmented boundaries or confuse classes with similar local shapes. Reliable fine-grained segmentation requires a local feature operator that can adapt to the geometry and spatial distribution of road marking points.
2.3. Semantic Segmentation
LiDAR point cloud semantic segmentation methods can be broadly categorized according to their input representation into point-based methods, projection-based methods, and voxel-based methods.
Point-based methods, such as PointNet [25], KPConv [26], and Point Transformer [27], directly process raw three-dimensional points as network inputs. PointNet uses sequential multilayer perceptrons to aggregate local point features into a global feature vector. KPConv defines three-dimensional kernel points and performs convolution directly on input points. Based on this idea, KPConvX [28] introduces an attention mechanism into kernel points, thereby improving model capacity without substantially increasing the size of the three-dimensional convolutional network. PointMixer [29] adapts the MLP-Mixer [30] architecture to point cloud processing, while Point Transformer uses the Transformer architecture to compute features for query points within local neighborhoods obtained by k-nearest neighbors. Fu et al. [31] proposed RailDLA-Net, an intensity-aware deep local aggregation framework for railway point cloud semantic segmentation, demonstrating the effectiveness of local feature aggregation in complex and imbalanced transportation scenes. These methods can effectively capture local geometric structures from raw point clouds. However, their high computational cost limits their applicability to large-scale outdoor LiDAR segmentation. For road marking segmentation, this limitation becomes more evident because large road scenes contain massive point clouds, while road markings occupy only a small proportion of the overall scene. Boundary-aware refinement has also been explored outside the point cloud domain: Huang et al. [32] proposed CBRNet, a corner-guided boundary refinement network that improves the extraction accuracy of fine-structured objects in remote sensing imagery, and a recent survey [33] on deep learning-based image segmentation similarly underscores the value of multi-scale feature representation for thin, elongated structures such as road markings.
RoadMark-AWAConv differs from these local feature operators in the way geometric support is constructed and used. KPConv defines local convolution using spatial kernel points, whereas RoadMark-AWAConv introduces a boundary-dense annular initialization tailored to road marking geometry and combines this initialization with learnable spatial weighting. Deformable convolution [34] adjusts sampling locations through learned offsets from an initial kernel, while RoadMark-AWAConv directly measures the spatial correlation between neighboring points and weight anchors. Point Transformer aggregates local information through attention between points, whereas RoadMark-AWAConv retains an explicit anchor-based convolution structure.
Projection-based methods, such as RangeViT [35], RangeFormer [36], and RangeNet++ [37], project LiDAR points onto two-dimensional planes, enabling semantic segmentation to benefit from network architectures originally developed for two-dimensional image tasks. RangeNet++ is one of the first methods to address point cloud semantic segmentation using range images generated by spherical projection, for which a two-dimensional convolutional neural network is adapted. RangeViT applies pre-trained vision Transformer models to two-dimensional range images, demonstrating the effectiveness of image pre-training for range-view point cloud segmentation. RangeFormer introduces RangeAug, which enhances two-dimensional projected range images by generating multiple augmented inputs to improve model performance. Projection-based methods offer fast inference and efficient computation, but they inevitably suffer from information loss during projection. This is particularly problematic for thin and elongated road markings, where geometric discontinuities and local feature fragmentation may occur after projection.
Voxel-based methods, such as MinkowskiNet [38], SphereFormer [39], and SPVCNN [40], divide three-dimensional space into voxel grids for efficient computation. MinkowskiNet voxelizes LiDAR data into sparse cubic grids and applies sparse convolution for feature extraction. SphereFormer introduces radial windows during voxelization and uses a Transformer structure to aggregate long-range information. SPVCNN builds upon MinkowskiNet and incorporates point-wise multilayer perceptrons to enhance point-level feature representation. Voxel-based methods provide a balance between inference efficiency and segmentation accuracy. Nevertheless, voxelization may still cause the loss of fine geometric details, and fixed grid structures are not sufficiently flexible for modeling irregular, sparse, and elongated road marking patterns.
Overall, existing semantic segmentation methods have achieved strong performance in general point cloud understanding. Ruan et al. [41] proposed SKPNet, which combines dynamic snake convolution and channel-wise self-attention to improve the continuity and detail representation of thin structures, providing useful insights for fine-grained road marking segmentation. However, road marking point clouds present specific challenges, including elongated structures, sparse and discrete distributions, non-uniform point density, pavement texture noise, and vehicle occlusion. These characteristics require a segmentation method that can better adapt to local road marking geometry and point density variation.
3. Materials and Methods
RoadMark-AWAConv is proposed for point-wise semantic segmentation of road marking point clouds. The method follows an encoder–decoder architecture and incorporates adaptive weight-anchor convolution into the local feature extraction process. Three components are introduced to address the geometric characteristics of road markings: geometry-constrained annular-domain anchor initialization, hierarchical radius-based neighborhood aggregation, and adaptive normal vector calibration. The overall framework and the individual components are described in the following sections.
3.1. Problem Formulation
Given a road marking point cloud , each point contains three-dimensional coordinates and additional attributes such as RGB values and reflectance intensity. The objective of fine-grained road marking semantic segmentation is to learn a point-wise mapping that assigns one of C semantic labels to each input point. In this study, C = 13.
3.2. Overview of RoadMark-AWAConv
In road marking point cloud segmentation, conventional point convolution methods often suffer from insufficient robustness when confronted with vehicle occlusions, pavement texture noise, elongated structures, and large geometric variations among different road marking categories. Road markings are usually represented by thin, sparse, and locally discontinuous point structures. Their discriminative features are mainly distributed along narrow boundaries and irregular local regions, making them difficult to capture using fixed convolution kernels or density-insensitive neighborhood aggregation.
To address these challenges, this study proposes RoadMark-AWAConv, namely Adaptive Weight-Anchor Convolution, for fine-grained semantic segmentation of road marking point clouds. The proposed method learns geometry-adaptive anchor distributions to better fit the geometric structures of road markings and improve local feature aggregation in complex road scenes. A weight anchor is defined as a discrete node in Euclidean space that carries convolution weights and anchors a local feature extraction region. Unlike fixed convolution kernels, the spatial positions of the weight anchors are optimized during training, allowing the convolution kernel to better adapt to irregular road marking geometries. This design enables the network to dynamically adapt to irregular road marking geometries, thereby improving the representation of elongated and sparse marking structures.
The overall architecture of the proposed method is a deep learning network designed for point-wise semantic segmentation of road marking point clouds. As shown in Figure 1, RoadMark-AWAConv follows an encoder–decoder framework and integrates the proposed Adaptive Weight-Anchor Convolution module into the local feature extraction process.
Figure 1.
Overall architecture of the proposed RoadMark-AWAConv network.
Given an input point cloud, the encoder first performs hierarchical sampling and feature extraction. At each encoding stage, local neighborhoods are constructed around sampled points, and the proposed AWAConv module is used to aggregate local features. Instead of relying on conventional fixed convolution kernels, AWAConv employs learnable weight anchors and a spatial correlation function to dynamically assign feature weights within local neighborhoods. In this way, the encoder can capture the geometric characteristics of elongated road markings more effectively while suppressing interference from pavement noise and occluded regions.
After hierarchical feature extraction, the decoder gradually restores the point cloud resolution through interpolation-based upsampling. Skip connections are introduced between the encoder and decoder to fuse high-resolution spatial details from shallow layers with semantic features from deeper layers. This design helps preserve fine-grained geometric information, which is particularly important for thin and discontinuous road markings. In addition, the weight-attention calculation module shown in Figure 1 generates adaptive feature weights from learnable anchor distributions and spatial correlations, further enhancing the model’s ability to distinguish foreground road markings from complex backgrounds. Finally, a multilayer perceptron is used to predict point-wise semantic labels, completing the end-to-end mapping from raw point clouds to fine-grained road marking segmentation results.
3.3. Geometry-Constrained Weight-Anchor Feature Extraction
Road markings are commonly composed of linear or quasi-linear geometric structures, such as elongated rectangles, dashed lines, arrows, and pavement boundary markings. These structures make it difficult for conventional point-based methods to efficiently capture boundary details and maintain semantic continuity. To improve the representation of elongated road markings, RoadMark-AWAConv introduces a geometry-constrained weight-anchor initialization strategy that is specifically adapted to the geometric characteristics of road marking point clouds.
The initial weight anchors are organized into a boundary-dense and interior-sparse annular-domain structure. This structure is designed to emphasize boundary regions, where the most discriminative geometric features of road markings are usually located. Specifically, each local kernel contains 15 anchors arranged in three groups. One anchor is located at the center, four anchors are evenly distributed on an inner ring with a radius of 0.5 r, and ten anchors are evenly distributed on an outer ring with a radius of 0.9 r, where r denotes the neighborhood radius at the current encoder stage. All anchors are initialized on the local xy plane with their z coordinates set to zero. This configuration places more anchors near the outer part of the neighborhood while retaining spatial support in the interior region. Unlike conventional fixed kernel points whose spatial positions remain unchanged during feature learning, the annular layout in AWAConv serves as an initial geometric configuration. The anchor positions are subsequently optimized during training, allowing the convolution kernel to adjust its spatial response to different road marking structures. The feature extraction process consists of four main steps. First, the raw input point cloud is discretized through spatial sampling. Second, the 15 spatial anchors are distributed according to the above annular topology and are associated with the input points through a local spatial correspondence mechanism. Third, a differentiable distance metric is used to calculate the weight value of each anchor, and the resulting weight distribution can be visualized using a color gradient from low to high response values. Finally, weighted aggregation is performed to generate weighted output representations, thereby completing the feature transformation process. The set of weight anchors is defined as
where denotes the anchor set, represents the i-th anchor, is the spatial dimension, and is the number of anchors, which is set to 15 in this study. The anchors are distributed in an annular topology within the local spatial domain, providing directional coverage for local point structures and reducing orientation bias during feature extraction.
For each input point , the weight assigned to the j-th anchor is computed by a normalized distance-based function:
where denotes the correlation weight between input point and anchor , and is a learnable or predefined influence coefficient that controls the sensitivity of the weight distribution to spatial distance. A larger assigns higher responses to closer anchors, whereas a smaller produces smoother weight distributions.
The output weighted feature for point is then calculated as
where denotes the output weighted representation of , and (⋅) is a feature transformation function applied to the anchor representation. Through this weighted aggregation process, the output feature preserves the spatial distribution of the input point while incorporating adaptive feature enhancement from the weight anchors. As a result, the network can better capture boundary-sensitive and shape-sensitive features of road markings. The adaptive weight-anchor feature extraction process is illustrated in Figure 2.
Figure 2.
Illustration of adaptive weight-anchor feature extraction. The gray points denote initial anchors with constant scalar features, while the colored points represent adaptive weight anchors with learned weights.
3.4. Adaptive Weight-Anchor Convolution
The core of AWAConv is to replace conventional fixed convolution kernels with deformable spatial anchors, thereby enabling local geometry-adaptive feature aggregation for point clouds. Each anchor is represented as a learnable three-dimensional spatial coordinate and is associated with a weight matrix. These anchors are independent of the input point cloud and can be dynamically optimized during training to fit the irregular topology of road markings.
Given the non-uniform density distribution of road marking point clouds, AWAConv adopts a radius-based neighborhood rather than a fixed k-nearest-neighbor structure. This design alleviates feature distortion caused by local density variations and enables the network to aggregate features from geometrically meaningful local regions. Compared with fixed voxel-based convolution, the adaptive anchor design is more flexible in adapting to irregular local structures.
Given an input point cloud P and its point-wise features, the convolution operation at position x is formulated as
where denotes the input feature function, (⋅) is the convolution kernel function, is a neighboring point of , denotes the feature vector of point , and is the local neighborhood centered at .
The local neighborhood is defined by a radius constraint:
where is the radius threshold at encoder stage j. The radius is determined hierarchically according to the grid subsampling size at each encoder stage. Let denote the grid size at stage j. The initial grid size is set to 0.05 m and is doubled after each pooling stage according to . The corresponding neighborhood radius is defined as where α is fixed at 3.0. Therefore, the initial neighborhood radius is 0.15 m and increases with the encoder depth. The radius at each stage is preset before training and does not vary according to the density of an individual local region. Compared with fixed k-nearest-neighbor grouping, the radius-based neighborhood can better preserve local geometric consistency under non-uniform point density.
The convolution kernel is constructed from learnable anchors as follows:
where denotes the relative coordinate from the center point to its neighboring point , represents the n-th learnable or calibrated anchor, is the weight matrix associated with the n-th anchor, and is the number of anchors used in the convolution kernel. The function (⋅) measures the spatial correlation between the relative position and the anchor .
Each neighboring point interacts with all anchors rather than being assigned to a single anchor. For each neighboring point, the spatial correlations with all anchors are calculated, and these correlations determine the contributions of the corresponding anchor weights to the convolution response.
The correlation function is defined as
where β is a learnable scale parameter that controls the spatial influence range of each anchor. Its initial spatial influence range is set to approximately at each encoder stage. This continuous correlation function ensures stable gradient propagation and allows the convolution kernel to adapt smoothly to local geometric variations. Since anchor positions and scale parameters are learnable, AWAConv can dynamically adjust its spatial weighting pattern according to the local structure of road markings.
3.5. Adaptive Normal Vector Calibration
In road marking point cloud scenes, pavement undulations, non-uniform point density, and vehicle occlusions may cause deviations in local geometric directions. These deviations can weaken the ability of convolution operations to model elongated and planar road marking structures. To address this issue, this study introduces an Adaptive Normal Vector Calibration strategy, which dynamically calibrates local normal vectors and provides direction-aware constraints for anchor-based convolution.
For each point , principal component analysis is first applied to its local neighborhood to compute the covariance matrix. Eigenvalue decomposition is then performed, and the eigenvector corresponding to the smallest eigenvalue is selected as the initial normal vector . Based on the three eigenvalues of the local covariance matrix, the local planarity score is defined as
where , and denote the smallest, intermediate, and largest eigenvalues of the covariance matrix for point , respectively. A larger indicates stronger local planarity, while a smaller suggests that the local neighborhood is more irregular or more strongly affected by noise.
The final calibrated normal vector is obtained by adaptively fusing the initial normal vector with the average normal vector of neighboring points:
where denotes the calibrated normal vector of point , is the initial normal vector estimated by PCA, and is the average normal vector of neighboring points. When the local region exhibits strong planarity, the calibrated normal vector relies more on the point’s own normal direction; when the local structure is noisy or unstable, the neighborhood average normal vector contributes more to smoothing abnormal directions.
The calibrated normal vector dynamically guides the orientation of anchor convolution, thereby improving the continuity and robustness of road marking segmentation. By incorporating direction-aware geometric constraints, the proposed calibration strategy enhances the model’s ability to distinguish elongated markings from pavement noise and occluded background points.
4. Results
4.1. Experimental Dataset and Settings
The experiments were conducted on a fine-grained road marking point cloud dataset, acquired using a vehicle-mounted mobile laser scanning (MLS) system equipped with a Hesai XT32 LiDAR sensor (Hesai Technology Co., Ltd., Shanghai, China). Similar demands for adaptive, environment-aware scanning control have also been studied for other LiDAR platforms, such as UAV-based systems [42]. The road-marking points were divided into 13 semantic classes: Straight Arrow, Left Turn Arrow, Right Turn Arrow, Compound Arrow, Crosswalk, Yield Marking, Text, Guidance Line, Chevron Marking, Dashed Line, Pavement Edge Line, Solid Line, and Others.
Each point contains three-dimensional coordinates, RGB values, and reflectance intensity. The 13 classes include elongated, discontinuous, intersecting, and irregular road marking structures, which are used to evaluate fine-grained road marking segmentation. The mean Intersection over Union (mIoU) and per-class IoU were adopted as the main evaluation metrics. These metrics are widely used in point cloud semantic segmentation and can comprehensively evaluate the overall segmentation performance and category-level accuracy, especially for multi-class tasks with severe class imbalance and many small objects.
To ensure a comprehensive comparison, representative methods from different point cloud segmentation paradigms were selected as baselines, including PointNet++ [43], RandLA-Net [44], Point Transformer V3 [45], and DeLA [46]. These methods cover point-based, efficient large-scale point cloud, Transformer-based, and road-marking-oriented segmentation frameworks.
All experiments were implemented using PyTorch 2.1.0 and conducted on an NVIDIA GeForce RTX 3090 GPU (NVIDIA Corporation, Santa Clara, CA, USA). The AdamW optimizer was used for parameter optimization, with an initial learning rate of 0.001 and a step decay strategy. During training, data augmentation strategies, including random rotation, translation, and Gaussian noise perturbation, were adopted to improve model generalization. The total number of training epochs was set to 100. Local feature regions were constructed using radius-based neighborhoods, and the proposed geometry-constrained weight anchors were used for dynamic feature aggregation and convolution. The initial grid subsampling size dl0 was set to 0.05 m according to the overall point density of the dataset and doubled after each encoder downsampling stage. The neighborhood radius at stage j was determined by , with α empirically set to 3.0 and kept fixed in all experiments. Accordingly, the initial neighborhood radius was 0.15 m. These radius settings were fixed before training and remained unchanged during training and inference.
4.2. Semantic Segmentation Results
Table 1 reports the quantitative segmentation results of different methods on the road marking point cloud dataset. The mIoU and per-class IoU are used as evaluation metrics, and the best result in each column is highlighted in bold.
Table 1.
Quantitative segmentation results of different methods on the road marking point cloud dataset. The best result in each column is highlighted in bold.
Overall, the proposed RoadMark-AWAConv achieves the best mIoU of 71.94%, outperforming PointNet++ by 18.46 percentage points, RandLA-Net by 15.98 percentage points, Point Transformer V3 by 10.63 percentage points, and DeLA by 2.56 percentage points. These results demonstrate the effectiveness of the proposed method in fine-grained multi-class road marking segmentation.
From the perspective of category-level performance, the proposed method achieves the best results on several important road marking categories. For Right Turn Arrow, RoadMark-AWAConv obtains an IoU of 89.37%, which is substantially higher than that of DeLA. This indicates that the proposed method can better capture the asymmetric geometric structure of turning-arrow markings. For Text, the proposed method achieves an IoU of 88.33%, showing stronger modeling capability for discrete and structurally complex markings. For Crosswalk, RoadMark-AWAConv obtains an IoU of 96.62%, showing stable recognition performance for large-area and high-reflectance road markings, although RandLA-Net achieves a slightly higher IoU of 97.63%. In addition, the proposed method also achieves competitive or superior performance on Dashed Line and Others, with IoU values of 86.66% and 51.91%, respectively, demonstrating its adaptability to discontinuous markings and complex boundary-like categories.
Compared with the proposed method, PointNet++ and RandLA-Net show relatively limited performance on most fine-grained categories, especially on complex or elongated markings such as Left Turn Arrow and Chevron Marking. This suggests that conventional point-based networks have difficulty capturing fine geometric details of road markings. Point Transformer V3 achieves better overall performance than PointNet++ and RandLA-Net, but it still performs poorly on several key categories, such as Right Turn Arrow and Text, indicating that pure Transformer-based architectures may still be limited in local geometric detail modeling. DeLA, as a road-marking-oriented segmentation network, performs well on several categories such as Straight Arrow and Yield Marking, but it still shows limited performance on Right Turn Arrow. In contrast, RoadMark-AWAConv achieves more balanced performance across different road marking shapes and scales through adaptive weight-anchor convolution.
The qualitative results are shown in Figure 3. PointNet++ and RandLA-Net produce fragmented predictions in several regions containing thin or discontinuous road markings. Missing points and class confusion are especially visible around dashed lines, guidance markings, and regions where several marking classes are located close to one another.
Figure 3.
Overall and local qualitative comparison of semantic segmentation results on the road marking point cloud dataset. The red boxes indicate the selected local region, and the corresponding enlarged views are shown on the right. From top to bottom: RGB-rendered point cloud, ground truth, Point Transformer V3, RandLA-Net, PointNet++, DeLA, and Ours.
Point Transformer V3 preserves the main structures more completely than PointNet++ and RandLA-Net, but errors remain around complex arrows and locally irregular markings. DeLA produces relatively complete results for several large or regular structures, although unclear boundaries and class confusion can still be observed in some complex regions.
Compared with these methods, RoadMark-AWAConv generally produces more complete contours and fewer discontinuities in the displayed scene. The shapes of right-turn and compound arrows are better preserved, and elongated or dashed markings remain more continuous. In regions containing several adjacent road marking classes, the predictions are also closer to the ground truth.
The qualitative results are consistent with the quantitative comparison in Table 1. RoadMark-AWAConv does not improve every category, but it provides more stable results for markings with different shapes, orientations, and local point distributions.
The enlarged views on the right provide a closer comparison of the selected local region. Compared with the baseline methods, RoadMark-AWAConv better preserves the continuity of thin boundaries and elongated markings, with fewer local discontinuities and misclassified points.
4.3. Ablation Study
To evaluate the effectiveness of each core component in RoadMark-AWAConv, ablation experiments were conducted on the road marking point cloud dataset. As shown in Table 2, the baseline model removes all proposed improvement modules, and the three key components are gradually introduced: geometry-constrained annular-domain anchor initialization, hierarchical radius-based neighborhood aggregation, and Adaptive Normal Vector Calibration.
Table 2.
Ablation study of the proposed modules.
After introducing only the geometry-constrained annular-domain anchor initialization, the mIoU increases from 58.72% to 65.17%. This improvement indicates that the boundary-dense and interior-sparse annular anchor distribution can enhance the model’s ability to capture shape-sensitive features of elongated markings, such as arrows and chevron markings. It also helps overcome the limitation of conventional uniformly sampled anchors, which may fail to sufficiently cover boundary details of road markings.
When hierarchical radius-based neighborhood aggregation is further added, the mIoU increases by 3.26 percentage points to 68.43%. This demonstrates that the radius-based neighborhood construction is more suitable for non-uniform point density than fixed k-nearest-neighbor grouping. It improves the segmentation completeness of sparse and discrete structures, such as dashed lines and text markings, and reduces feature loss caused by locally sparse point distributions.
Finally, after incorporating Adaptive Normal Vector Calibration, the full model achieves an mIoU of 71.94%. This module dynamically adjusts the orientation of anchor convolution and improves the robustness of the model against vehicle occlusions, pavement noise, and other complex interferences. The improvement is particularly beneficial for categories that are easily affected by background points, such as Left Turn Arrow and Solid Line. The ablation results confirm that each proposed component contributes to the final segmentation performance, and their combination enables RoadMark-AWAConv to better model fine-grained road marking structures.
5. Discussion
The experimental results indicate that fine-grained road marking point cloud segmentation differs substantially from general road-scene point cloud segmentation. Road markings usually occupy only a small proportion of the entire road scene and often appear as elongated, sparse, discrete, and locally discontinuous structures. For such objects, a model is required not only to predict point-wise semantic labels but also to preserve boundary details, directional patterns, and topological continuity. Therefore, directly applying general-purpose point cloud segmentation networks may be insufficient for fine-grained road marking perception.
The performance differences among comparison methods further support this observation. PointNet++ and RandLA-Net can process raw point clouds or large-scale point clouds, but their ability to model small, discontinuous, and boundary-sensitive road markings is still limited, which may lead to fragmented predictions and category confusion. Point Transformer V3 has stronger global feature modeling capability, but road markings rely heavily on local geometric boundaries and subtle shape differences; therefore, a general attention structure may still be insufficient for stable local structure representation. DeLA, as a road-marking-oriented method, performs well on some regular categories, but it still shows limited performance on some turning-arrow categories, particularly Right Turn Arrow. These results suggest that road marking segmentation requires not only semantic discrimination but also more explicit local geometric constraints.
The advantage of RoadMark-AWAConv mainly comes from its adaptation to the geometric and density characteristics of road marking point clouds. The annular-domain anchor initialization encourages the model to focus more on boundary regions, which is beneficial for capturing the contours of elongated markings and arrow-like structures. The hierarchical radius-based neighborhood aggregation maintains a predefined spatial scale at each encoder stage while allowing the number of neighboring points to vary with the local point distribution, reducing feature loss caused by non-uniform density. In addition, normal vector direction calibration enhances the sensitivity of anchor convolution to local orientation changes, helping the model maintain better segmentation continuity under pavement noise, occlusion, and incomplete local observations. Compared with fixed kernels, fixed neighborhoods, or general attention mechanisms, the proposed method is more directly tailored to the geometric representation of road markings.
The current method still has several limitations. First, Left Turn Arrow remains a difficult category across all evaluated methods, and RoadMark-AWAConv obtains an IoU of 7.62% for this class. This result may be related to the limited number of samples, incomplete point distributions, or geometric similarity to other arrow classes. Second, classes with similar shapes or functions, such as different arrow types, Pavement Edge Line and Solid Line, may still be confused. Class imbalance can further affect categories with fewer points. Third, the current experiments are conducted using LiDAR point clouds from a single road marking dataset. Further evaluation is still needed to examine the generalization of the method across different acquisition systems, road environments, and point-density patterns. Future work will consider class-balanced learning, evaluation on additional road scenes, multi-scale neighborhood modeling, and the fusion of point clouds with image and road-topology information. Further evaluation of computational cost is also needed for real-time vehicle deployment.
6. Conclusions
This paper presents RoadMark-AWAConv, an adaptive weight-anchor convolution method for fine-grained semantic segmentation of road marking point clouds. The method introduces a geometry-constrained annular-domain anchor initialization strategy to improve the representation of elongated boundaries and irregular local structures. Hierarchical radius-based neighborhood aggregation is used to reduce the influence of non-uniform point density, while Adaptive Normal Vector Calibration provides more stable orientation information under pavement variation, noise, and incomplete observations.
Experiments on a road marking point cloud dataset containing 13 semantic classes show that RoadMark-AWAConv achieves mIoU of 71.94%, outperforming PointNet++, RandLA-Net, Point Transformer V3, and DeLA. The quantitative comparison, qualitative results, and ablation study demonstrate the effectiveness of the proposed components for fine-grained road marking segmentation. Future work will focus on class-balanced learning, multi-scale neighborhood modeling, computational efficiency, and the fusion of LiDAR point clouds with image and road-topology information.
Author Contributions
Conceptualization, X.H., T.J., S.L. and Y.W.; methodology, X.H., Y.G. and Z.H.; software, X.H., Y.G. and Z.X.; validation, Y.G. and Z.X.; formal analysis, X.H., Y.G., M.H., T.J. and S.L.; investigation, Y.G., M.H., T.J. and S.L.; resources, T.J., S.L. and Y.W.; data curation, Z.H.; writing—original draft preparation, X.H., T.J. and S.L.; writing—review and editing, X.H., Y.G., M.H., T.J., S.L. and Y.W.; visualization, Z.X. and Z.H.; supervision, T.J., S.L. and Y.W.; project administration, M.H., T.J., S.L. and Y.W.; funding acquisition, M.H., T.J., S.L. and Y.W. All authors have read and agreed to the published version of the manuscript.
Funding
This research was supported by the Natural Science Foundation of Jiangsu Province, China (grant number BK20240598), the National Natural Science Foundation of China (grant numbers 42401552 and 42671586), the Natural Science Foundation of the Higher Education Institutions of Jiangsu Province, China (grant number 24KJB420005), the Postgraduate Research and Practice Innovation Program of Jiangsu Province (grant number SJCX25_0704), and the grant from State Key Laboratory of Resources and Environmental Information System.
Data Availability Statement
Data underlying the results presented in this paper are not publicly available at this time but may be obtained from the authors upon reasonable request.
Acknowledgments
The authors acknowledge all the reviewers for their valuable comments.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Zhang, Y.; Lu, Z.; Zhang, X.; Xue, J.-H.; Liao, Q. Deep Learning in Lane Marking Detection: A Survey. IEEE Trans. Intell. Transp. Syst. 2022, 23, 5976–5992. [Google Scholar] [CrossRef] [Scilit]
- Bi, J.; Song, Y.; Jiang, Y.; Sun, L.; Wang, X.; Liu, Z.; Xu, J.; Quan, S.; Dai, Z.; Yan, W. Lane Detection for Autonomous Driving: Comprehensive Reviews, Current Challenges, and Future Predictions. IEEE Trans. Intell. Transp. Syst. 2025, 26, 5710–5746. [Google Scholar] [CrossRef] [Scilit]
- Huang, M.; Li, H.; Chen, N.; Lin, H.; Zhu, D.; Gong, D.; Chen, Y.; Altan, O.; Gong, J. Overcoming Optical Observation Limitations: Automatic Dense Time-Series Mapping of Impervious Surfaces in Cloudy and Snow-Covered Regions. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2026, 19, 21312–21333. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Nguyen, T.-M.; Cao, M.; Yuan, S.; Hung, T.-Y.; Xie, L. Graph Optimality-Aware Stochastic LiDAR Bundle Adjustment with Progressive Spatial Smoothing. IEEE Trans. Intell. Transp. Syst. 2025, 26, 19076–19091. [Google Scholar] [CrossRef] [Scilit]
- Luo, Z.; Gao, L.; Xiang, H.; Li, J. Road Object Detection for HD Map: Full-Element Survey, Analysis and Perspectives. ISPRS J. Photogramm. Remote Sens. 2023, 197, 122–144. [Google Scholar] [CrossRef] [Scilit]
- Mi, X.; Dong, Z.; Cao, Z.; Yang, B.; Cao, Z.; Zheng, C.; Stoter, J.; Nan, L. A Benchmark Approach and Dataset for Large-Scale Lane Mapping from MLS Point Clouds. Int. J. Appl. Earth Obs. Geoinf. 2024, 133, 104139. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Xu, X.; Liu, J.; Cao, K.; Yuan, S.; Xie, L. UA-MPC: Uncertainty-Aware Model Predictive Control for Motorized LiDAR Odometry. IEEE Robot. Autom. Lett. 2025, 10, 3652–3659. [Google Scholar] [CrossRef] [Scilit]
- Yang, B.; Liu, Y.; Dong, Z.; Liang, F.; Li, B.; Peng, X. 3D Local Feature BKD to Extract Road Information from Mobile Laser Scanning Point Clouds. ISPRS J. Photogramm. Remote Sens. 2017, 130, 329–343. [Google Scholar] [CrossRef] [Scilit]
- Chen, D.; Wang, Y.; Zhang, L.; Kang, Z. Enhanced Local Feature Learning with Simple Offset Attention for Semantic Segmentation of Large-Scale Point Clouds. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5705916. [Google Scholar] [CrossRef] [Scilit]
- Armeni, I.; Sener, O.; Zamir, A.R.; Jiang, H.; Brilakis, I.; Fischer, M.; Savarese, S. 3D Semantic Parsing of Large-Scale Indoor Spaces. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; IEEE: Piscataway, NJ, USA, 2016. [Google Scholar]
- Dai, A.; Chang, A.X.; Savva, M.; Halber, M.; Funkhouser, T.; Nießner, M. ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; IEEE: Piscataway, NJ, USA, 2017. [Google Scholar]
- Wu, Z.; Song, S.; Khosla, A.; Yu, F.; Zhang, L.; Tang, X.; Xiao, J. 3D ShapeNets: A Deep Representation for Volumetric Shapes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; IEEE: Piscataway, NJ, USA, 2015. [Google Scholar]
- Chang, A.X.; Funkhouser, T.; Guibas, L.; Hanrahan, P.; Huang, Q.; Li, Z.; Savarese, S.; Savva, M.; Song, S.; Su, H.; et al. ShapeNet: An Information-Rich 3D Model Repository. arXiv 2015, arXiv:1512.03012. [Google Scholar]
- Behley, J.; Garbade, M.; Milioto, A.; Quenzel, J.; Behnke, S.; Stachniss, C.; Gall, J. SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; IEEE: Piscataway, NJ, USA, 2019. [Google Scholar]
- Hackel, T.; Savinov, N.; Ladický, L.; Wegner, J.D.; Schindler, K.; Pollefeys, M. Semantic3D.net: A New Large-Scale Point Cloud Classification Benchmark. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2017, IV-1-W1, 91–98. [Google Scholar] [CrossRef] [Scilit]
- Roynard, X.; Deschaud, J.-E.; Goulette, F. Paris-Lille-3D: A Point Cloud Dataset for Urban Scene Segmentation and Classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA, 18–22 June 2018; IEEE: Piscataway, NJ, USA, 2018. [Google Scholar]
- Tan, W.; Qin, N.; Ma, L.; Li, Y.; Du, J.; Cai, G.; Yang, K.; Li, J. Toronto-3D: A Large-Scale Mobile LiDAR Dataset for Semantic Segmentation of Urban Roadways. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA, 14–19 June 2020; IEEE: Piscataway, NJ, USA, 2020. [Google Scholar]
- Bai, M.; Mattyus, G.; Homayounfar, N.; Wang, S.; Lakshmikanth, S.K.; Urtasun, R. Deep Multi-Sensor Lane Detection. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 1–5 October 2018; IEEE: Piscataway, NJ, USA, 2018. [Google Scholar]
- Martinek, P.; Pucea, G.; Rao, Q.; Sivalingam, U. Lidar-Based Deep Neural Network for Reference Lane Generation. In Proceedings of the IEEE Intelligent Vehicles Symposium (IV), Las Vegas, NV, USA, 19 October–13 November 2020; IEEE: Piscataway, NJ, USA, 2020. [Google Scholar]
- Paek, D.-H.; Kong, S.-H.; Wijaya, K.T. K-Lane: Lidar Lane Dataset and Benchmark for Urban Roads and Highways. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), New Orleans, LA, USA, 19–20 June 2022; IEEE: Piscataway, NJ, USA, 2022. [Google Scholar]
- Du, J.; Ma, L.; Li, J.; Qin, N.; Zelek, J.; Guan, H.; Li, J. RdmkNet & Toronto-RDMK: Large-Scale Datasets for Road Marking Classification and Segmentation. IEEE Trans. Intell. Transp. Syst. 2024, 25, 13467–13482. [Google Scholar] [CrossRef] [Scilit]
- Jiang, T.; Li, S.; Zhang, Q.; Wang, G.; Zhang, Z.; Zeng, F.; An, P.; Jin, X.; Liu, S.; Wang, Y. RailPC: A Large-scale Railway Point Cloud Semantic Segmentation Dataset. CAAI Trans. Intell. Technol. 2024, 9, 1548–1560. [Google Scholar] [CrossRef] [Scilit]
- Cáceres Hernández, D.; Hoang, V.-D.; Jo, K.-H. Lane Surface Identification Based on Reflectance Using Laser Range Finder. In Proceedings of the IEEE/SICE International Symposium on System Integration (SII), Tokyo, Japan, 13–15 December 2014; IEEE: Piscataway, NJ, USA, 2014. [Google Scholar]
- Mi, X.; Yang, B.; Dong, Z.; Liu, C.; Zong, Z.; Yuan, Z. A Two-Stage Approach for Road Marking Extraction and Modeling Using MLS Point Clouds. ISPRS J. Photogramm. Remote Sens. 2021, 180, 255–268. [Google Scholar] [CrossRef] [Scilit]
- Charles, R.Q.; Su, H.; Kaichun, M.; Guibas, L.J. PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; IEEE: Piscataway, NJ, USA, 2017. [Google Scholar]
- Thomas, H.; Qi, C.R.; Deschaud, J.-E.; Marcotegui, B.; Goulette, F.; Guibas, L. KPConv: Flexible and Deformable Convolution for Point Clouds. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; IEEE: Piscataway, NJ, USA, 2019. [Google Scholar]
- Zhao, H.; Jiang, L.; Jia, J.; Torr, P.; Koltun, V. Point Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; IEEE: Piscataway, NJ, USA, 2021. [Google Scholar]
- Thomas, H.; Tsai, Y.-H.H.; Barfoot, T.D.; Zhang, J. KPConvX: Modernizing Kernel Point Convolution with Kernel Attention. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; IEEE: Piscataway, NJ, USA, 2024. [Google Scholar]
- Choe, J.; Park, C.; Rameau, F.; Park, J.; Kweon, I.S. PointMixer: MLP-Mixer for Point Cloud Understanding. In Proceedings of the 17th European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022. [Google Scholar]
- Tolstikhin, I.O.; Houlsby, N.; Kolesnikov, A.; Beyer, L.; Zhai, X.; Unterthiner, T.; Yung, J.; Steiner, A.; Keysers, D.; Uszkoreit, J.; et al. MLP-Mixer: An All-MLP Architecture for Vision. In Proceedings of the 35th Conference on Neural Information Processing Systems (NeurIPS), Virtual Conference, 6–14 December 2021. [Google Scholar]
- Fu, J.; Liu, S.; Pang, J.; Jing, X.; Zhang, Y.; Wang, G.; Dai, W.; Jiang, T. RailDLA-Net: An Intensity-Aware Deep Local Aggregation Framework for Railway Point Cloud Semantic Segmentation. Remote Sens. 2026, 18, 2379. [Google Scholar] [CrossRef] [Scilit]
- Huang, M.; Cao, S.; Zhu, D.; Liu, X.; Luo, J.; Chen, Y.; Niu, J.; Zhang, L.; Huang, X.; Lin, H. CBRNet: Corner-Guided Boundary Refinement Network for High-Precision Building Extraction From Remote Sensing Imagery. IEEE Trans. Geosci. Remote Sens. 2026, 64, 4406216. [Google Scholar] [CrossRef] [Scilit]
- Wang, G.; Li, Z.; Weng, G.; Chen, Y. An Overview of Industrial Image Segmentation Using Deep Learning Models. Intell. Robot. 2025, 5, 143–180. [Google Scholar] [CrossRef] [Scilit]
- Dai, J.; Qi, H.; Xiong, Y.; Li, Y.; Zhang, G.; Hu, H.; Wei, Y. Deformable Convolutional Networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; IEEE: Piscataway, NJ, USA, 2017. [Google Scholar]
- Ando, A.; Gidaris, S.; Bursuc, A.; Puy, G.; Boulch, A.; Marlet, R. RangeViT: Towards Vision Transformers for 3D Semantic Segmentation in Autonomous Driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; IEEE: Piscataway, NJ, USA, 2023. [Google Scholar]
- Kong, L.; Liu, Y.; Chen, R.; Ma, Y.; Zhu, X.; Li, Y.; Hou, Y.; Qiao, Y.; Liu, Z. Rethinking Range View Representation for LiDAR Segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; IEEE: Piscataway, NJ, USA, 2023. [Google Scholar]
- Milioto, A.; Vizzo, I.; Behley, J.; Stachniss, C. RangeNet ++: Fast and Accurate LiDAR Semantic Segmentation. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 3–8 November 2019; IEEE: Piscataway, NJ, USA, 2019. [Google Scholar]
- Choy, C.; Gwak, J.; Savarese, S. 4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 16–20 June 2019; IEEE: Piscataway, NJ, USA, 2019. [Google Scholar]
- Lai, X.; Chen, Y.; Lu, F.; Liu, J.; Jia, J. Spherical Transformer for LiDAR-Based 3D Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; IEEE: Piscataway, NJ, USA, 2023. [Google Scholar]
- Tang, H.; Liu, Z.; Zhao, S.; Lin, Y.; Lin, J.; Wang, H.; Han, S. Searching Efficient 3D Architectures with Sparse Point-Voxel Convolution. In Proceedings of the 16th European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020. [Google Scholar]
- Ruan, Y.; Wang, D.; Yuan, Y.; Jiang, S.; Yang, X. SKPNet: Snake KAN Perceive Bridge Cracks through Semantic Segmentation. Intell. Robot. 2025, 5, 105–118. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Xu, X.; Liu, Z.; Yuan, S.; Cao, M.; Xie, L. AEOS: Active Environment-Aware Optimal Scanning Control for UAV LiDAR-Inertial Odometry in Complex Scenes. ISPRS J. Photogramm. Remote Sens. 2026, 232, 476–491. [Google Scholar] [CrossRef] [Scilit]
- Qi, C.R.; Yi, L.; Su, H.; Guibas, L.J. PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS), Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Hu, Q.; Yang, B.; Xie, L.; Rosa, S.; Guo, Y.; Wang, Z.; Trigoni, N.; Markham, A. RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; IEEE: Piscataway, NJ, USA, 2020. [Google Scholar]
- Wu, X.; Jiang, L.; Wang, P.-S.; Liu, Z.; Liu, X.; Qiao, Y.; Ouyang, W.; He, T.; Zhao, H. Point Transformer V3: Simpler, Faster, Stronger. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; IEEE: Piscataway, NJ, USA, 2024. [Google Scholar]
- Yang, W.; Lu, X.; Chen, B.; Lin, C.; Bao, X.; Liu, W.; Zang, Y.; Xu, J.; Wang, C. DeLA: An Extremely Faster Network with Decoupled Local Aggregation for Large Scale Point Cloud Learning. Int. J. Appl. Earth Obs. Geoinf. 2024, 135, 104255. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.


