Attentive Excitation and Aggregation for Bilingual Referring Image Segmentation

被引:4
|
作者
Zhou, Qianli [1 ]
Hui, Tianrui [2 ,5 ]
Wang, Rong [1 ]
Hu, Haimiao [3 ]
Liu, Si [4 ]
机构
[1] Peoples Publ Secur Univ China, 1 Muxidi Nanli, Beijing, Peoples R China
[2] Chinese Acad Sci, Inst Informat Engn, 89 Minzhuang Rd, Beijing, Peoples R China
[3] Beihang Univ, 37 Xueyuan Rd, Beijing, Peoples R China
[4] Beihang Univ, Inst Artificial Intelligence, 37 Xueyuan Rd, Beijing, Peoples R China
[5] Univ Chinese Acad Sci, Sch Cyber Secur, 19 Yuquan Rd, Beijing, Peoples R China
基金
北京市自然科学基金; 中国国家自然科学基金;
关键词
Bilingual referring segmentation; channel excitation; spatial aggregation;
D O I
10.1145/3446345
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
The goal of referring image segmentation is to identify the object matched with an input natural language expression. Previous methods only support English descriptions, whereas Chinese is also broadly used around theworld, which limits the potential application of this task. Therefore, we propose to extend existing datasets with Chinese descriptions and preprocessing tools for training and evaluating bilingual referring segmentation models. In addition, previous methods also lack the ability to collaboratively learn channel-wise and spatial-wise cross-modal attention to well align visual and linguistic modalities. To tackle these limitations, we propose a Linguistic Excitation module to excite image channels guided by language information and a Linguistic Aggregation module to aggregate multimodal information based on image-language relationships. Since different levels of features from the visual backbone encode rich visual information, we also propose a Cross-Level Attentive Fusion module to fuse multilevel features gated by language information. Extensive experiments on four English and Chinese benchmarks show that our bilingual referring image segmentation model outperforms previous methods.
引用
收藏
页数:17
相关论文
共 50 条
  • [31] Multi-Scale Attentive Aggregation for LiDAR Point Cloud Segmentation
    Geng, Xiaoxiao
    Ji, Shunping
    Lu, Meng
    Zhao, Lingli
    REMOTE SENSING, 2021, 13 (04) : 1 - 12
  • [32] Text-Vision Relationship Alignment for Referring Image Segmentation
    Pu, Mingxing
    Luo, Bing
    Zhang, Chao
    Xu, Li
    Xu, Fayou
    Kong, Mingming
    NEURAL PROCESSING LETTERS, 2024, 56 (02)
  • [33] Local-global coordination with transformers for referring image segmentation
    Liu, Fang
    Kong, Yuqiu
    Zhang, Lihe
    Feng, Guang
    Yin, Baocai
    NEUROCOMPUTING, 2023, 522 : 39 - 52
  • [34] Prompt-Driven Referring Image Segmentation with Instance Contrasting
    Shang, Chao
    Song, Zichen
    Qiu, Heqian
    Wang, Lanxiao
    Meng, Fanman
    Li, Hongliang
    2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2024, 2024, : 4124 - 4134
  • [35] Referring Image Segmentation via Language-Driven Attention
    Chen, Ding-Jie
    Hsieh, He-Yen
    Liu, Tyng-Luh
    2021 IEEE INTERNATIONAL CONFERENCE ON ROBOTICS AND AUTOMATION (ICRA 2021), 2021, : 13997 - 14003
  • [36] Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation
    Yan, Yichen
    He, Xingjian
    Chen, Sihan
    Liu, Jing
    PROCEEDINGS OF THE 4TH ANNUAL ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA RETRIEVAL, ICMR 2024, 2024, : 451 - 459
  • [37] Beyond One-to-One: Rethinking the Referring Image Segmentation
    Hu, Yutao
    Wang, Qixiong
    Shao, Wenqi
    Xie, Enze
    Li, Zhenguo
    Han, Jungong
    Luo, Ping
    2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION, ICCV, 2023, : 4044 - 4054
  • [38] Vision-Aware Language Reasoning for Referring Image Segmentation
    Xu, Fayou
    Luo, Bing
    Zhang, Chao
    Xu, Li
    Pu, Mingxing
    Li, Bo
    NEURAL PROCESSING LETTERS, 2023, 55 (08) : 11313 - 11331
  • [39] Global Selection and Local Attention Network for Referring Image Segmentation
    Ding, Haixin
    Zhang, Shengchuan
    Cao, Liujuan
    PATTERN RECOGNITION AND COMPUTER VISION, PRCV 2023, PT VII, 2024, 14431 : 284 - 295
  • [40] Bidirectional Relationship Inferring Network for Referring Image Localization and Segmentation
    Feng, Guang
    Hu, Zhiwei
    Zhang, Lihe
    Sun, Jiayu
    Lu, Huchuan
    IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2023, 34 (05) : 2246 - 2258