Next Article in Journal
Data Science in the Management of Healthcare Organizations
Previous Article in Journal
Generation of Sparse Antennas and Scatterers Based on Optimal Current Grid Approximation
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

MPVF: Multi-Modal 3D Object Detection Algorithm with Pointwise and Voxelwise Fusion

1
School of Mechanical and Automotive Engineering, Anhui Polytechnic University, Wuhu 241000, China
2
Polytechnic Institute, Zhejiang University, Hangzhou 310015, China
*
Author to whom correspondence should be addressed.
Algorithms 2025, 18(3), 172; https://doi.org/10.3390/a18030172
Submission received: 23 February 2025 / Revised: 12 March 2025 / Accepted: 18 March 2025 / Published: 19 March 2025
(This article belongs to the Section Algorithms for Multidisciplinary Applications)

Abstract

3D object detection plays a pivotal role in achieving accurate environmental perception, particularly in complex traffic scenarios where single-modal detection methods often fail to meet precision requirements. This highlights the necessity of multi-modal fusion approaches to enhance detection performance. However, existing camera-LiDAR intermediate fusion methods suffer from insufficient interaction between local and global features and limited fine-grained feature extraction capabilities, which results in inadequate small object detection and unstable performance in complex scenes. To address these issues, the multi-modal 3D object detection algorithm with pointwise and voxelwise fusion (MPVF) is proposed, which enhances multi-modal feature interaction and optimizes feature extraction strategies to improve detection precision and robustness. First, the pointwise and voxelwise fusion (PVWF) module is proposed to combine local features from the pointwise fusion (PWF) module with global features from the voxelwise fusion (VWF) module, enhancing the interaction between features across modalities, improving small object detection capabilities, and boosting model performance in complex scenes. Second, an expressive feature extraction module, improved ResNet-101 and feature pyramid (IRFP), is developed, comprising the improved ResNet-101 (IR) and feature pyramid (FP) modules. The IR module uses a group convolution strategy to inject high-level semantic features into the PWF and VWF modules, improving extraction efficiency. The FP module, placed at an intermediate stage, captures fine-grained features at various resolutions, enhancing the model’s precision and robustness. Finally, evaluation on the KITTI dataset demonstrates a mean Average Precision (mAP) of 69.24%, a 2.75% improvement over GraphAlign++. Detection accuracy for cars, pedestrians, and cyclists reaches 85.12%, 48.61%, and 70.12%, respectively, with the proposed method excelling in pedestrian and cyclist detection.
Keywords: 3D object detection; autonomous driving; image; point cloud; multi-modal fusion 3D object detection; autonomous driving; image; point cloud; multi-modal fusion

Share and Cite

MDPI and ACS Style

Shi, P.; Wu, W.; Yang, A. MPVF: Multi-Modal 3D Object Detection Algorithm with Pointwise and Voxelwise Fusion. Algorithms 2025, 18, 172. https://doi.org/10.3390/a18030172

AMA Style

Shi P, Wu W, Yang A. MPVF: Multi-Modal 3D Object Detection Algorithm with Pointwise and Voxelwise Fusion. Algorithms. 2025; 18(3):172. https://doi.org/10.3390/a18030172

Chicago/Turabian Style

Shi, Peicheng, Wenchao Wu, and Aixi Yang. 2025. "MPVF: Multi-Modal 3D Object Detection Algorithm with Pointwise and Voxelwise Fusion" Algorithms 18, no. 3: 172. https://doi.org/10.3390/a18030172

APA Style

Shi, P., Wu, W., & Yang, A. (2025). MPVF: Multi-Modal 3D Object Detection Algorithm with Pointwise and Voxelwise Fusion. Algorithms, 18(3), 172. https://doi.org/10.3390/a18030172

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop