CenterFormer: Center-Based Transformer for 3D Object Detection

被引:61
|
作者
Zhou, Zixiang [1 ,2 ]
Zhao, Xiangchen [1 ]
Wang, Yu [1 ]
Wang, Panqu [1 ]
Foroosh, Hassan [2 ]
机构
[1] TuSimple, San Diego, CA 92122 USA
[2] Univ Cent Florida, Computat Imaging Lab, Orlando, FL 32816 USA
来源
关键词
LiDAR point cloud; 3D object detection; Transformer; Multi-frame fusion;
D O I
10.1007/978-3-031-19839-7_29
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Query-based transformer has shown great potential in constructing long-range attention in many image-domain tasks, but has rarely been considered in LiDAR-based 3D object detection due to the overwhelming size of the point cloud data. In this paper, we propose CenterFormer, a center-based transformer network for 3D object detection. CenterFormer first uses a center heatmap to select center candidates on top of a standard voxel-based point cloud encoder. It then uses the feature of the center candidate as the query embedding in the transformer. To further aggregate features from multiple frames, we design an approach to fuse features through cross-attention. Lastly, regression heads are added to predict the bounding box on the output center feature representation. Our design reduces the convergence difficulty and computational complexity of the transformer structure. The results show significant improvements over the strong baseline of anchor-free object detection networks. CenterFormer achieves state-of-the-art performance for a single model on the Waymo Open Dataset, with 73.7% mAPH on the validation set and 75.6% mAPH on the test set, significantly outperforming all previously published CNN and transformer-based methods. Our code is publicly available at https://github.com/TuSimple/centerformer
引用
收藏
页码:496 / 513
页数:18
相关论文
共 50 条
  • [31] Transformer3D-Det: Improving 3D Object Detection by Vote Refinement
    Zhao, Lichen
    Guo, Jinyang
    Xu, Dong
    Sheng, Lu
    IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2021, 31 (12) : 4735 - 4746
  • [32] F-Transformer: Point Cloud Fusion Transformer for Cooperative 3D Object Detection
    Wang, Jie
    Luo, Guiyang
    Yuan, Quan
    Li, Jinglin
    ARTIFICIAL NEURAL NETWORKS AND MACHINE LEARNING - ICANN 2022, PT I, 2022, 13529 : 171 - 182
  • [33] 3DPPE: 3D Point Positional Encoding for Transformer-based Multi-Camera 3D Object Detection
    Shu, Changyong
    Deng, Jiajun
    Yu, Fisher
    Liu, Yifan
    2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION, ICCV, 2023, : 3557 - 3566
  • [34] ReAGFormer: Reaggregation Transformer with Affine Group Features for 3D Object Detection
    Lu, Chenguang
    Yue, Kang
    Liu, Yue
    COMPUTER VISION - ACCV 2022, PT I, 2023, 13841 : 262 - 279
  • [35] PVTransformer: Point-to-Voxel Transformer for Scalable 3D Object Detection
    Leng, Zhaoqi
    Sun, Pei
    He, Tong
    Anguelov, Dragomir
    Tan, Mingxing
    2024 IEEE INTERNATIONAL CONFERENCE ON ROBOTICS AND AUTOMATION, ICRA 2024, 2024, : 4238 - 4244
  • [36] BEV transformer for visual 3D object detection applied with retentive mechanism
    Pan, Jincheng
    Huang, Xiaoci
    Luo, Suyun
    Ma, Fang
    TRANSACTIONS OF THE INSTITUTE OF MEASUREMENT AND CONTROL, 2025,
  • [37] MVTr: multi-feature voxel transformer for 3D object detection
    Ai, Lingmei
    Xie, Zhuoyu
    Yao, Ruoxia
    Yang, Mengyao
    VISUAL COMPUTER, 2024, 40 (03): : 1453 - 1466
  • [38] Radar-camera fusion for 3D object detection with aggregation transformer
    Li, Jun
    Zhang, Han
    Wu, Zizhang
    Xu, Tianhao
    APPLIED INTELLIGENCE, 2024, 54 (21) : 10627 - 10639
  • [39] MVTr: multi-feature voxel transformer for 3D object detection
    Lingmei Ai
    Zhuoyu Xie
    Ruoxia Yao
    Mengyao Yang
    The Visual Computer, 2024, 40 : 1453 - 1466
  • [40] MonoDTR: Monocular 3D Object Detection with Depth-Aware Transformer
    Huang, Kuan-Chih
    Wu, Tsung-Han
    Su, Hung-Ting
    Hsu, Winston H.
    2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, : 4002 - 4011