RQFormer: Rotated Query Transformer for end-to-end oriented object detection

被引:0
|
作者
Zhao, Jiaqi [1 ,2 ,3 ]
Ding, Zeyu [1 ,2 ]
Zhou, Yong [1 ,2 ]
Zhu, Hancheng [1 ,2 ]
Du, Wen-Liang [1 ,2 ]
Yao, Rui [1 ,2 ]
El Saddik, Abdulmotaleb [4 ]
机构
[1] China Univ Min & Technol, Sch Comp Sci & Technol, Xuzhou 221116, Peoples R China
[2] Minist Educ, Mine Digitizat Engn Res Ctr, Xuzhou 221116, Peoples R China
[3] Innovat Res Ctr Disaster Intelligent Prevent & Eme, Xuzhou 221116, Peoples R China
[4] Univ Ottawa, Sch Elect Engn & Comp Sci, Ottawa, ON K1N 6N5, Canada
基金
中国国家自然科学基金;
关键词
Oriented object detection; Transformer; End-to-end detectors; Attention; Query update;
D O I
10.1016/j.eswa.2024.126034
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Oriented object detection presents a challenging task due to the presence of object instances with multiple orientations, varying scales, and dense distributions. Recently, end-to-end detectors have made significant strides by employing attention mechanisms and refining a fixed number of queries through consecutive decoder layers. However, existing end-to-end oriented object detectors still face two primary challenges: (1) misalignment between positional queries and keys, leading to inconsistency between classification and localization; and (2) the presence of a large number of similar queries, which complicates one-to-one label assignments and optimization. To address these limitations, we propose an end-to-end oriented detector called the Rotated Query Transformer, which integrates two key technologies: Rotated RoI Attention (RRoI Attention) and Selective Distinct Queries (SDQ). First, RRoI Attention aligns positional queries and keys from oriented regions of interest through cross-attention. Second, SDQ collects queries from intermediate decoder layers and filters out similar ones to generate distinct queries, thereby facilitating the optimization of one-to-one label assignments. Finally, extensive experiments conducted on four remote sensing datasets and one scene text dataset demonstrate the effectiveness of our method. To further validate its generalization capability, we also extend our approach to horizontal object detection. The code is available at https://github.com/ wokaikaixinxin/RQFormer.
引用
收藏
页数:14
相关论文
共 50 条
  • [41] End-to-End Object Detection with Enhanced Positive Sample Filter
    Song, Xiaolin
    Chen, Binghui
    Li, Pengyu
    Wang, Biao
    Zhang, Honggang
    APPLIED SCIENCES-BASEL, 2023, 13 (03):
  • [42] Dynamic DETR: End-to-End Object Detection with Dynamic Attention
    Dai, Xiyang
    Chen, Yinpeng
    Yang, Jianwei
    Zhang, Pengchuan
    Yuan, Lu
    Zhang, Lei
    2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, : 2968 - 2977
  • [43] EENED: End-to-End Neural Epilepsy Detection based on Convolutional Transformer
    Liu, Chenyu
    Zhou, Xinliang
    Liu, Yang
    2023 IEEE CONFERENCE ON ARTIFICIAL INTELLIGENCE, CAI, 2023, : 368 - 371
  • [44] DefectTR: End-to-end defect detection for sewage networks using a transformer
    Dang, L. Minh
    Wang, Hanxiang
    Li, Yanfen
    Nguyen, Tan N.
    Moon, Hyeonjoon
    CONSTRUCTION AND BUILDING MATERIALS, 2022, 325
  • [45] ARCHITECTURAL MODEL FOR SONET END-TO-END MANAGEMENT WITH OBJECT-ORIENTED APPROACH
    NISHIDA, T
    HASEGAWA, S
    KANESAMA, A
    DALLAS GLOBECOM 89, VOLS 1-3: COMMUNICATIONS TECHNOLOGY FOR THE 1990S AND BEYOND, 1989, : 1500 - 1505
  • [46] Transformer Generates Conditional Convolution Kernels for End-to-End Lane Detection
    Zhuang, Long
    Jiang, Tiezhen
    Qiu, Meng
    Wang, Anqi
    Huang, Zhixiang
    IEEE SENSORS JOURNAL, 2024, 24 (17) : 28383 - 28396
  • [47] End-to-End Real-Time Vanishing Point Detection with Transformer
    Tong, Xin
    Peng, Shi
    Guo, Yufei
    Huang, Xuhui
    THIRTY-EIGHTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 38 NO 6, 2024, : 5243 - 5251
  • [48] DFDT: An End-to-End DeepFake Detection Framework Using Vision Transformer
    Khormali, Aminollah
    Yuan, Jiann-Shiun
    APPLIED SCIENCES-BASEL, 2022, 12 (06):
  • [49] Semantic-Aware and Goal-Oriented Communications for Object Detection in Wireless End-to-End Image Transmission
    Safaeipour, Fatemeh Zahra
    Hashemi, Morteza
    2024 INTERNATIONAL CONFERENCE ON COMPUTING, NETWORKING AND COMMUNICATIONS, ICNC, 2024, : 182 - 187
  • [50] Neural Network Based End-to-End Query by Example Spoken Term Detection
    Ram, Dhananjay
    Miculicich, Lesly
    Bourlard, Herve
    IEEE-ACM TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2020, 28 (28) : 1416 - 1427