electronics-logo

Journal Browser

Journal Browser

Advances in 2D/3D Object Detection Techniques and Systems

A special issue of Electronics (ISSN 2079-9292). This special issue belongs to the section "Computer Science & Engineering".

Deadline for manuscript submissions: 15 October 2026 | Viewed by 2358

Editors


E-Mail Website
Guest Editor
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China
Interests: computer vision; robot multimodal perception; robot skill learning and development
Special Issues, Collections and Topics in MDPI journals
College of Information Science and Technology, Beijing University of Chemical Technology, Beijing 100013, China
Interests: power electronics; power systems; fault detection; object detection in industry; interdisciplinary research combined energy with remote sensing

Special Issue Information

Dear Colleagues,

The rapid advancement of 2D and 3D sensing technologies is fundamentally transforming intelligent systems—from robotics and autonomous driving to augmented reality and industrial automation. While 2D object detection provides mature, efficient, and semantically rich scene understanding, 3D detection adds crucial geometric and spatial awareness, enabling machines to interact with the world in a truly physical sense. The integration and co-design of 2D/3D perception are now key to building robust, reliable, and context-aware autonomous systems.

Despite remarkable progress, significant challenges remain. In 2D detection, issues such as occlusion, scale variation, and domain adaptation persist. In 3D detection, challenges include sparse and irregular point cloud data, high computational cost, and sensitivity to sensor viewpoints. Moreover, effectively fusing 2D and 3D modalities to leverage their complementary strengths, such as marrying the rich texture from images with precise geometry from point clouds, presents a central, open research problem. Achieving real-time performance, robustness in diverse environments, and generalizability across applications further compounds these challenges.

This Special Issue, titled “Advances in 2D/3D Object Detection Techniques and Systems,” will capture the latest breakthroughs and innovative solutions across the entire spectrum of object perception. We welcome contributions that advance the state of the art in either 2D or 3D detection, as well as pioneering research on their synergistic fusion. Our goal is to foster a cross-disciplinary dialogue that accelerates the development of next-generation perception systems.

We invite submissions on a broad range of topics, including but not limited to, the following:

  • Novel Architectures for 2D and 3D Detection: Transformers, efficient CNNs, point-based networks, and hybrid models for image and point cloud processing.
  • Multi-Modal Fusion and Cross-Modal Learning: Innovative methods to integrate RGB images, LiDAR, radar, depth maps, and IMU data for enhanced perception.
  • Learning with Limited Supervision: Self-supervised, semi-supervised, and weakly supervised techniques for 2D/3D detection to reduce annotation dependency.
  • Efficiency and Deployment: Model compression, neural architecture search, and optimization for edge devices, drones, and mobile robots.
  • Robustness and Generalization: Domain adaptation, test-time augmentation, and uncertainty estimation for real-world conditions (e.g., weather, lighting).
  • Holistic Scene Understanding: Context-aware detection, panoptic segmentation, and leveraging temporal or spatial relationships in complex scenes.
  • Datasets, Simulation and Benchmarking: The creation of large-scale datasets, realistic simulators, and standardized evaluation protocols for 2D/3D tasks.
  • Application-Driven Systems: Case studies in autonomous driving, robotic manipulation, smart manufacturing, industrial application, AR/VR, and healthcare, highlighting system integration and practical insights.

Conclusions

We invite researchers and practitioners from academia and industry to submit original research articles, comprehensive reviews, and insightful case studies. By bridging the realms of 2D and 3D perception, this Special Issue will chart a course towards more intelligent, adaptive, and capable perception systems for the future.

Dr. Yanfeng Lu
Dr. Yi Li
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Electronics is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2400 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • 2D/3D object detection
  • cross-modal 3D perception
  • intelligent perception systems
  • robotics perception

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (4 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

18 pages, 17266 KB  
Article
Efficient 3D Semantic Occupancy Prediction via Integrated 2D-3D Feature Fusion
by Sang-Min Park and Jong-Eun Ha
Electronics 2026, 15(16), 3530; https://doi.org/10.3390/electronics15163530 - 8 Aug 2026
Viewed by 333
Abstract
This paper presents a novel approach to 3D semantic occupancy prediction that leverages the integration of 2D and 3D features. Traditional 3D voxel representations, while detailed, are computationally intensive. Our method addresses this challenge by encoding 3D voxel features into a Bird’s-Eye View [...] Read more.
This paper presents a novel approach to 3D semantic occupancy prediction that leverages the integration of 2D and 3D features. Traditional 3D voxel representations, while detailed, are computationally intensive. Our method addresses this challenge by encoding 3D voxel features into a Bird’s-Eye View (BEV) representation, then decoding them back into voxels using a multi-layer perceptron (MLP). This fusion approach reduces computational resources compared to voxel-only methods while maintaining state-of-the-art accuracy. By leveraging multi-scale features and deformable attention mechanisms, our network achieves a mean intersection-over-union (mIoU) of 41.26% on the Occ3D-nuScenes dataset, outperforming recent methods. By introducing a 3D backbone and multi-scale voxel-to-BEV-to-voxel feature transformation, the model achieves improved mIoU with moderate computational overhead. Our results demonstrate the effectiveness of integrating 2D and 3D information for accurate 3D semantic occupancy prediction, providing a potentially useful scene representation for downstream autonomous driving tasks. Full article
(This article belongs to the Special Issue Advances in 2D/3D Object Detection Techniques and Systems)
Show Figures

Figure 1

25 pages, 1722 KB  
Article
OPT-Net: An Orientation-Preserving Transformer for End-to-End Oriented Object Detection in Remote Sensing Images
by Jiaxin Xu, Hua Huo, Aokun Mei and Chen Zhang
Electronics 2026, 15(13), 2819; https://doi.org/10.3390/electronics15132819 - 26 Jun 2026
Viewed by 522
Abstract
The objects in high-resolution remote sensing images usually exhibit arbitrary orientations, multi-scale variations, dense distributions, and complex background interference, posing significant challenges to oriented object detection. Although existing DETR-style end-to-end detectors eliminate the need for anchor design and non-maximum suppression, they still suffer [...] Read more.
The objects in high-resolution remote sensing images usually exhibit arbitrary orientations, multi-scale variations, dense distributions, and complex background interference, posing significant challenges to oriented object detection. Although existing DETR-style end-to-end detectors eliminate the need for anchor design and non-maximum suppression, they still suffer from insufficient orientation priors in object queries, limited orientation consistency in decoder feature interaction, and unstable set matching for oriented bounding boxes. To address these issues, this paper proposes an end-to-end Transformer framework, termed OPT-Net (Orientation-Preserving Transformer Network), for oriented object detection in remote sensing images. OPT-Net treats orientation information as a structured geometric prior and propagates it through query initialization, feature interaction, and matching optimization. Specifically, an Orientation-Aware Query Initialization (OAQI) module is designed to generate initial queries using center confidence and orientation priors. An Orientation-Consistent Cross-Attention (OCCA) mechanism is proposed to perform orientation-conditioned modulation on Value features while keeping the standard Query–Key attention computation unchanged. Furthermore, an Uncertainty-aware Matching Loss (UML) is introduced to incorporate instance-level geometric uncertainty into Hungarian matching and regression optimization. Experimental results on the DOTA-v1.0 and HRSC2016 datasets show that OPT-Net achieves 76.83% and 90.58% mAP, respectively, demonstrating competitive detection accuracy and adaptability to complex remote sensing scenarios. Ablation studies and visualization results further validate the effectiveness of each proposed module. Full article
(This article belongs to the Special Issue Advances in 2D/3D Object Detection Techniques and Systems)
Show Figures

Figure 1

19 pages, 4235 KB  
Article
MV3-YOLO: A MobileNetV3-Based Lightweight Variant of YOLO for Efficient Object Detection
by Bojun Liu and Yanfeng Lu
Electronics 2026, 15(12), 2741; https://doi.org/10.3390/electronics15122741 - 22 Jun 2026
Viewed by 459
Abstract
Efficient object detection is needed in automated driving and edge perception. In these scenarios, a detector must work under limits on latency, power, and memory. YOLOv8 is a strong real-time baseline, but its computation can still be high for compact deployment. This paper [...] Read more.
Efficient object detection is needed in automated driving and edge perception. In these scenarios, a detector must work under limits on latency, power, and memory. YOLOv8 is a strong real-time baseline, but its computation can still be high for compact deployment. This paper proposes MV3-YOLO, a lightweight YOLOv8 variant with a stage-wise hybrid backbone. The early Conv/C2f stages are kept to retain low-level spatial details. Lightweight modules are placed in deeper stages, where feature maps are smaller and redundant computation is more common. C2fMixed is used at the stride-16 stage to balance feature capacity and cost. C2fGhostis used at the deepest stage to generate high-level features with fewer parameters. The YOLOv8 neck and head are kept unchanged for stable multi-scale fusion. On the KITTI validation set, MV3-YOLO reaches mAP@0.5 = 0.859 and mAP@0.5:0.95 = 0.610 with only 2.53 M parameters and 6.6 GFLOPs. Compared with YOLOv8n, it reduces parameters by 19.7% and GFLOPs by 25.0% while improving mAP@0.5 by 1.66% and mAP@0.5:0.95 by 1.50%. On COCO val2017, MV3-YOLO obtains 38.4 mAP@0.5:0.95, which is higher than the YOLOv8n reference result and close to YOLOv10n. These results show that MV3-YOLO reduces deployment cost while keeping competitive detection accuracy. Full article
(This article belongs to the Special Issue Advances in 2D/3D Object Detection Techniques and Systems)
Show Figures

Figure 1

21 pages, 3457 KB  
Article
Hardware-Accelerated 3D LiDAR-Based Object Detection with BEV Spatial Mapping on Embedded FPGA Platforms
by Güner Tatar and Mahmud Esad Arar
Electronics 2026, 15(11), 2296; https://doi.org/10.3390/electronics15112296 - 25 May 2026
Viewed by 627
Abstract
This paper introduces a hardware/software co-designed 3D object detection pipeline based on the PointPillars architecture for low-power embedded MPSoC deployment. The proposed system accelerates the computationally intensive stages in programmable logic (PL), including ROI filtering, coordinate transformation, pillarization, centroid extraction, and INT8 neural [...] Read more.
This paper introduces a hardware/software co-designed 3D object detection pipeline based on the PointPillars architecture for low-power embedded MPSoC deployment. The proposed system accelerates the computationally intensive stages in programmable logic (PL), including ROI filtering, coordinate transformation, pillarization, centroid extraction, and INT8 neural inference, using Vitis high-level synthesis (HLS) and an integrated Deep Learning Processing Unit (DPU). Control-oriented and irregular operations, such as data acquisition, Direct Memory Access (DMA) control, lightweight Non-Maximum Suppression (NMS), visualization, and logging, remain on the processing system (PS). The design targets the AMD Kria KV260 platform and achieves an accelerated core pipeline latency of 11.4 ms per frame at 300 MHz, corresponding to 87.4 Hz throughput, with 6.842 W board-level power consumption. Including PS-side NMS, the practical end-to-end latency is approximately 12.2 ms for typical KITTI scenes. Compared with existing Field-Programmable Gate Array (FPGA)-based implementations implementations, the proposed design reduces latency by up to 33×. It achieves a 202× improvement in on-chip BRAM efficiency across HLS optimization versions through FIFO streaming, dataflow execution, and array partitioning. Experimental validation on physical hardware confirms that the proposed PL-accelerated hardware/software co-design provides a practical and cost-effective solution for real-time 3D LiDAR perception on embedded FPGA platforms. Full article
(This article belongs to the Special Issue Advances in 2D/3D Object Detection Techniques and Systems)
Show Figures

Figure 1

Back to TopTop