Category-Level Pose Estimation and Iterative Refinement for Monocular RGB-D Image

被引：0

作者：

Bao, Yongtang ^{[1
]}

Qi, Yutong ^{[2
]}

Su, Chunjian ^{[1
]}

Geng, Yanbing ^{[3
]}

Li, Haojie ^{[1
]}

机构：

[1] Shandong Univ Sci & Technol, Coll Comp Sci & Engn, Qingdao, Peoples R China

[2] Univ Toronto, Dept Comp & Math Sci, Scarborough, ON, Canada

[3] North Univ China, Sch Data Sci & Technol, Taiyuan, Peoples R China

来源：

ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS | 2024年 / 20卷 / 12期

基金：

中国国家自然科学基金;

关键词：

Deep learning; category-level pose estimation; scene understanding; transformer; TRANSFORMER;

D O I：

10.1145/3695877

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Category-level pose estimation is proposed to predict the 6D pose of objects under a specific category and has wide applications in fields such as robotics, virtual reality, and autonomous driving. With the development of VR/AR technology, pose estimation has gradually become a research hotspot in 3D scene understanding. However, most methods fail to fully utilize geometric and color information to solve intra-class shape variations, which leads to inaccurate prediction results. To solve the above problems, we propose a novel pose estimation and iterative refinement network, use an attention mechanism to fuse multi-modal information to obtain color features after a coordinate transformation, and design iterative modules to ensure the accuracy of object geometric features. Specifically, we use an encoder-decoder architecture to implicitly generate a coarse-grained initial pose and refine it through an iterative refinement module. In addition, due to the differences between rotation and position estimation, we design a multi-head pose decoder that utilizes the local geometry and global features. Finally, we design a transformer-based coordinate transformation attention module to extract pose-sensitive features from RGB images and supervise color information by correlating point cloud features in different coordinate systems. We train and test our network on the synthetic dataset CAMERA25 and the real dataset REAL275. Experimental results show that our method achieves state-of-the-art performance on multiple evaluation metrics.

引用

页数：20

共 50 条

[21] Zero-Shot Category-Level Object Pose Estimation
Goodwin, Walter
Vaze, Sagar
Havoutis, Ioannis
Posner, Ingmar
COMPUTER VISION, ECCV 2022, PT XXXIX, 2022, 13699 : 516 - 532
[22] Optimal Pose and Shape Estimation for Category-level 3D Object Perception
Shi, Jingnan
Yang, Heng
Carlone, Luca
ROBOTICS: SCIENCE AND SYSTEM XVII, 2021,
[23] Category-Level Metric Scale Object Shape and Pose Estimation
Lee, Taeyeop
Lee, Byeong-Uk
Kim, Myungchul
Kweon, I. S.
IEEE ROBOTICS AND AUTOMATION LETTERS, 2021, 6 (04) : 8575 - 8582
[24] CPPF: Towards Robust Category-Level 9D Pose Estimation in the Wild
You, Yang
Shi, Ruoxi
Wang, Weiming
Lu, Cewu
2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, : 6856 - 6865
[25] Generative Category-Level Shape and Pose Estimation with Semantic Primitives
Li, Guanglin
Li, Yifeng
Ye, Zhichao
Zhang, Qihang
Kong, Tao
Cui, Zhaopeng
Zhang, Guofeng
CONFERENCE ON ROBOT LEARNING, VOL 205, 2022, 205 : 1390 - 1400
[26] RGB-D Image-based Pose Estimation with Monte Carlo Localization
Li, Ming
Qin, Hao
Huang, May
Cao, Jian
Zhang, Xing
2017 3RD INTERNATIONAL CONFERENCE ON CONTROL, AUTOMATION AND ROBOTICS (ICCAR), 2017, : 109 - 114
[27] Textured/textureless object recognition and pose estimation using RGB-D image
Wei Wang
Lili Chen
Ziyuan Liu
Kolja Kühnlenz
Darius Burschka
Journal of Real-Time Image Processing, 2015, 10 : 667 - 682
[28] Textured/textureless object recognition and pose estimation using RGB-D image
Wang, Wei
Chen, Lili
Liu, Ziyuan
Kuehnlenz, Kolja
Burschka, Darius
JOURNAL OF REAL-TIME IMAGE PROCESSING, 2015, 10 (04) : 667 - 682
[29] RBP-Pose: Residual Bounding Box Projection for Category-Level Pose Estimation
Zhang, Ruida
Di, Yan
Lou, Zhiqiang
Manhardi, Fabian
Tombari, Federico
Ji, Xiangyang
COMPUTER VISION - ECCV 2022, PT I, 2022, 13661 : 655 - 672
[30] i2c-net: Using Instance-Level Neural Networks for Monocular Category-Level 6D Pose Estimation
Remus, Alberto
D'Avella, Salvatore
Di Felice, Francesco
Tripicchio, Paolo
Avizzano, Carlo Alberto
IEEE ROBOTICS AND AUTOMATION LETTERS, 2023, 8 (03) : 1515 - 1522

← 1 2 3 4 5 →