SparseDet: A Simple and Effective Framework for Fully Sparse LiDAR-Based 3-D Object Detection

被引:1
|
作者
Liu, Lin [1 ]
Song, Ziying [1 ]
Xia, Qiming [2 ]
Jia, Feiyang [1 ]
Jia, Caiyan [1 ]
Yang, Lei [3 ,4 ]
Gong, Yan [5 ]
Pan, Hongyu [6 ]
机构
[1] Beijing Jiaotong Univ, Sch Comp Sci & Technol, Beijing Key Lab Traff Data Anal & Min, Beijing 100044, Peoples R China
[2] Xiamen Univ, Fujian Key Lab Sensing & Comp Smart Cities, Xiamen 361005, Fujian, Peoples R China
[3] Tsinghua Univ, State Key Lab Intelligent Green Vehicle & Mobil, Beijing 100084, Peoples R China
[4] Tsinghua Univ, Sch Vehicle & Mobil, Beijing 100084, Peoples R China
[5] JD Logist, Autonomous Driving Dept X Div, Beijing 101111, Peoples R China
[6] Horizon Robot, Beijing 100190, Peoples R China
关键词
Feature extraction; Three-dimensional displays; Point cloud compression; Detectors; Aggregates; Object detection; Computational efficiency; 3-D object detection; feature aggregation; sparse detectors;
D O I
10.1109/TGRS.2024.3468394
中图分类号
P3 [地球物理学]; P59 [地球化学];
学科分类号
0708 ; 070902 ;
摘要
LiDAR-based sparse 3-D object detection plays a crucial role in autonomous driving applications due to its computational efficiency advantages. Existing methods either use the features of a single central voxel as an object proxy or treat an aggregated cluster of foreground points as an object proxy. However, the former cannot aggregate contextual information, resulting in insufficient information expression in object proxies. The latter relies on multistage pipelines and auxiliary tasks, which reduce the inference speed. To maintain the efficiency of the sparse framework while fully aggregating contextual information, in this work, we propose SparseDet that designs sparse queries as object proxies. It introduces two key modules: the local multiscale feature aggregation (LMFA) module and the global feature aggregation (GFA) module, aiming to fully capture the contextual information, thereby enhancing the ability of the proxies to represent objects. The LMFA module achieves feature fusion across different scales for sparse key voxels via coordinate transformations and using nearest neighbor relationships to capture object-level details and local contextual information, whereas the GFA module uses self-attention mechanisms to selectively aggregate the features of the key voxels across the entire scene for capturing scene-level contextual information. Experiments on nuScenes and KITTI demonstrate the effectiveness of our method. Specifically, SparseDet surpasses the previous best sparse detector VoxelNeXt (a typical method using voxels as object proxies) by 2.2% mean average precision (mAP) with 13.5 frames/s on nuScenes and outperforms VoxelNeXt by 1.12% AP(3-D) on hard level tasks with 17.9 frames/s on KITTI. What is more, not only the mAP of SparseDet exceeds that of FSDV2 (a classical method using clusters of foreground points as object proxies) but also its inference speed is 1.3 times faster than FSDV2 on the nuScenes test set. The code has been released in https://github.com/liulin813/SparseDet.git.
引用
收藏
页数:14
相关论文
共 50 条
  • [1] Diversity Knowledge Distillation for LiDAR-Based 3-D Object Detection
    Ning, Kanglin
    Liu, Yanfei
    Su, Yanzhao
    Jiang, Ke
    IEEE SENSORS JOURNAL, 2023, 23 (11) : 11181 - 11193
  • [2] SAFDNet: A Simple and Effective Network for Fully Sparse 3D Object Detection
    Zhang, Gang
    Chen, Junnan
    Gao, Guohuan
    Li, Jianmin
    Liu, Si
    Hu, Xiaolin
    2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2024, : 14477 - 14486
  • [3] OpenSight: A Simple Open-Vocabulary Framework for LiDAR-Based Object Detection
    Zhang, Hu
    Ku, Jianhua
    Tang, Tao
    Sun, Haiyang
    Huang, Xin
    Huang, Zi
    Yu, Kaicheng
    COMPUTER VISION - ECCV 2024, PT LXXXIV, 2025, 15142 : 1 - 19
  • [4] LiDAR-Based Symmetrical Guidance for 3D Object Detection
    Chu, Huazhen
    Ma, Huimin
    Liu, Haizhuang
    Wang, Rongquan
    PATTERN RECOGNITION AND COMPUTER VISION, PT IV, 2021, 13022 : 472 - 483
  • [5] LiDAR-based 3D Object Detection for Autonomous Driving
    Li, Zirui
    2022 INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, COMPUTER VISION AND MACHINE LEARNING (ICICML), 2022, : 507 - 512
  • [6] RTL3D: real-time LIDAR-based 3D object detection with sparse CNN
    Yan, Lin
    Liu, Kai
    Belyaev, Evgeny
    Duan, Meiyu
    IET COMPUTER VISION, 2020, 14 (05) : 224 - 232
  • [7] Gradual Batch Alternation for Effective Domain Adaptation in LiDAR-Based 3D Object Detection
    Rochan, Mrigank
    Chen, Xingxin
    Grandhi, Alaap
    Corral-Soto, Eduardo R.
    Liu, Bingbing
    2024 35TH IEEE INTELLIGENT VEHICLES SYMPOSIUM, IEEE IV 2024, 2024, : 2213 - 2219
  • [8] Multi-modal information fusion for LiDAR-based 3D object detection framework
    Ma, Ruixin
    Yin, Yong
    Chen, Jing
    Chang, Rihao
    MULTIMEDIA TOOLS AND APPLICATIONS, 2024, 83 (03) : 7995 - 8012
  • [9] Exploring Context Information in Camera and Lidar-Based 3-D Object Detection Using Graphs and Transformers
    Sun, Chang
    Li, Yuehua
    Xing, Yan
    Zhang, Weidong
    Ai, Yibo
    Wang, Sheng
    IEEE TRANSACTIONS ON INSTRUMENTATION AND MEASUREMENT, 2025, 74
  • [10] Multi-modal information fusion for LiDAR-based 3D object detection framework
    Ruixin Ma
    Yong Yin
    Jing Chen
    Rihao Chang
    Multimedia Tools and Applications, 2024, 83 : 7995 - 8012