Category-Level Pose Estimation and Iterative Refinement for Monocular RGB-D Image

被引:0
|
作者
Bao, Yongtang [1 ]
Qi, Yutong [2 ]
Su, Chunjian [1 ]
Geng, Yanbing [3 ]
Li, Haojie [1 ]
机构
[1] Shandong Univ Sci & Technol, Coll Comp Sci & Engn, Qingdao, Peoples R China
[2] Univ Toronto, Dept Comp & Math Sci, Scarborough, ON, Canada
[3] North Univ China, Sch Data Sci & Technol, Taiyuan, Peoples R China
基金
中国国家自然科学基金;
关键词
Deep learning; category-level pose estimation; scene understanding; transformer; TRANSFORMER;
D O I
10.1145/3695877
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Category-level pose estimation is proposed to predict the 6D pose of objects under a specific category and has wide applications in fields such as robotics, virtual reality, and autonomous driving. With the development of VR/AR technology, pose estimation has gradually become a research hotspot in 3D scene understanding. However, most methods fail to fully utilize geometric and color information to solve intra-class shape variations, which leads to inaccurate prediction results. To solve the above problems, we propose a novel pose estimation and iterative refinement network, use an attention mechanism to fuse multi-modal information to obtain color features after a coordinate transformation, and design iterative modules to ensure the accuracy of object geometric features. Specifically, we use an encoder-decoder architecture to implicitly generate a coarse-grained initial pose and refine it through an iterative refinement module. In addition, due to the differences between rotation and position estimation, we design a multi-head pose decoder that utilizes the local geometry and global features. Finally, we design a transformer-based coordinate transformation attention module to extract pose-sensitive features from RGB images and supervise color information by correlating point cloud features in different coordinate systems. We train and test our network on the synthetic dataset CAMERA25 and the real dataset REAL275. Experimental results show that our method achieves state-of-the-art performance on multiple evaluation metrics.
引用
收藏
页数:20
相关论文
共 50 条
  • [1] Attention-guided RGB-D Fusion Network for Category-level 6D Object Pose Estimation
    Wang, Hao
    Li, Weiming
    Kim, Jiyeon
    Wang, Qiang
    2022 IEEE/RSJ INTERNATIONAL CONFERENCE ON INTELLIGENT ROBOTS AND SYSTEMS (IROS), 2022, : 10651 - 10658
  • [2] Bi-directional attention based RGB-D fusion for category-level object pose and shape estimation
    Tang, Kaifeng
    Xu, Chi
    Chen, Ming
    MULTIMEDIA TOOLS AND APPLICATIONS, 2023, 83 (17) : 53043 - 53063
  • [3] Bi-directional attention based RGB-D fusion for category-level object pose and shape estimation
    Kaifeng Tang
    Chi Xu
    Ming Chen
    Multimedia Tools and Applications, 2024, 83 : 53043 - 53063
  • [4] iCaps: Iterative Category-Level Object Pose and Shape Estimation
    Deng, Xinke
    Geng, Junyi
    Bretl, Timothy
    Xiang, Yu
    Fox, Dieter
    IEEE ROBOTICS AND AUTOMATION LETTERS, 2022, 7 (02): : 1784 - 1791
  • [5] CATRE: Iterative Point Clouds Alignment for Category-Level Object Pose Refinement
    Liu, Xingyu
    Wang, Gu
    Li, Yi
    Ji, Xiangyang
    COMPUTER VISION - ECCV 2022, PT II, 2022, 13662 : 499 - 516
  • [6] An RGB-D Refinement Solution for Accurate Object Pose Estimation
    Saadi, Lounes
    Besbes, Bassem
    Kramm, Sebastien
    Bensrhair, Abdelaziz
    2021 IEEE INTERNATIONAL SYMPOSIUM ON MIXED AND AUGMENTED REALITY ADJUNCT PROCEEDINGS (ISMAR-ADJUNCT 2021), 2021, : 189 - 194
  • [7] Object Level Depth Reconstruction for Category Level 6D Object Pose Estimation from Monocular RGB Image
    Fan, Zhaoxin
    Song, Zhenbo
    Xu, Jian
    Wang, Zhicheng
    Wu, Kejian
    Liu, Hongyan
    He, Jun
    COMPUTER VISION - ECCV 2022, PT II, 2022, 13662 : 220 - 236
  • [8] Single-Stage Keypoint-Based Category-Level Object Pose Estimation from an RGB Image
    Lin, Yunzhi
    Tremblay, Jonathan
    Tyree, Stephen
    Vela, Patricio A.
    Birchfield, Stan
    2022 IEEE INTERNATIONAL CONFERENCE ON ROBOTICS AND AUTOMATION (ICRA 2022), 2022, : 1547 - 1553
  • [9] Category-Level Articulated Object Pose Estimation
    Li, Xiaolong
    Wang, He
    Yi, Li
    Guibas, Leonidas
    Abbott, A. Lynn
    Song, Shuran
    2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2020, : 3703 - 3712
  • [10] Hand Pose Estimation from a Single RGB-D Image
    Kuznetsova, Alina
    Rosenhahn, Bodo
    ADVANCES IN VISUAL COMPUTING, PT II, 2013, 8034 : 592 - 602