Attentive Excitation and Aggregation for Bilingual Referring Image Segmentation

被引:4
|
作者
Zhou, Qianli [1 ]
Hui, Tianrui [2 ,5 ]
Wang, Rong [1 ]
Hu, Haimiao [3 ]
Liu, Si [4 ]
机构
[1] Peoples Publ Secur Univ China, 1 Muxidi Nanli, Beijing, Peoples R China
[2] Chinese Acad Sci, Inst Informat Engn, 89 Minzhuang Rd, Beijing, Peoples R China
[3] Beihang Univ, 37 Xueyuan Rd, Beijing, Peoples R China
[4] Beihang Univ, Inst Artificial Intelligence, 37 Xueyuan Rd, Beijing, Peoples R China
[5] Univ Chinese Acad Sci, Sch Cyber Secur, 19 Yuquan Rd, Beijing, Peoples R China
基金
北京市自然科学基金; 中国国家自然科学基金;
关键词
Bilingual referring segmentation; channel excitation; spatial aggregation;
D O I
10.1145/3446345
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
The goal of referring image segmentation is to identify the object matched with an input natural language expression. Previous methods only support English descriptions, whereas Chinese is also broadly used around theworld, which limits the potential application of this task. Therefore, we propose to extend existing datasets with Chinese descriptions and preprocessing tools for training and evaluating bilingual referring segmentation models. In addition, previous methods also lack the ability to collaboratively learn channel-wise and spatial-wise cross-modal attention to well align visual and linguistic modalities. To tackle these limitations, we propose a Linguistic Excitation module to excite image channels guided by language information and a Linguistic Aggregation module to aggregate multimodal information based on image-language relationships. Since different levels of features from the visual backbone encode rich visual information, we also propose a Cross-Level Attentive Fusion module to fuse multilevel features gated by language information. Extensive experiments on four English and Chinese benchmarks show that our bilingual referring image segmentation model outperforms previous methods.
引用
收藏
页数:17
相关论文
共 50 条
  • [21] Locate then Segment: A Strong Pipeline for Referring Image Segmentation
    Jing, Ya
    Kong, Tao
    Wang, Wei
    Wang, Liang
    Li, Lei
    Tan, Tieniu
    2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 9853 - 9862
  • [22] Learning From Box Annotations for Referring Image Segmentation
    Feng, Guang
    Zhang, Lihe
    Hu, Zhiwei
    Lu, Huchuan
    IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2024, 35 (03) : 3927 - 3937
  • [23] CARIS: Context-Aware Referring Image Segmentation
    Liu, Sun-Ao
    Zhang, Yiheng
    Qiu, Zhaofan
    Xie, Hongtao
    Zhang, Yongdong
    Yao, Ting
    PROCEEDINGS OF THE 31ST ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2023, 2023, : 779 - 788
  • [24] Query Reconstruction Network for Referring Expression Image Segmentation
    Shi, Hengcan
    Li, Hongliang
    Wu, Qingbo
    Ngan, King Ngi
    IEEE TRANSACTIONS ON MULTIMEDIA, 2021, 23 : 995 - 1007
  • [25] PRNet: A Progressive Refinement Network for referring image segmentation
    Liu, Jing
    Jiang, Huajie
    Hu, Yongli
    Yin, Baocai
    NEUROCOMPUTING, 2025, 630
  • [26] A CONTEXT-BASED NETWORK FOR REFERRING IMAGE SEGMENTATION
    Li, Xinyu
    Liu, Yu
    Xu, Kaiping
    Zhao, Zhehuan
    Liu, Sipei
    2020 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2020, : 1436 - 1440
  • [27] Advancing Referring Expression Segmentation Beyond Single Image
    Wu, Yixuan
    Zhang, Zhao
    Xie, Chi
    Zhu, Feng
    Zhao, Rui
    2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION, ICCV, 2023, : 2628 - 2638
  • [28] Referring Image Segmentation via Recurrent Refinement Networks
    Li, Ruiyu
    Li, Kaican
    Kuo, Yi-Chun
    Shu, Michelle
    Qi, Xiaojuan
    Shen, Xiaoyong
    Jia, Jiaya
    2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, : 5745 - 5753
  • [29] Bilateral Knowledge Interaction Network for Referring Image Segmentation
    Ding, Haixin
    Zhang, Shengchuan
    Wu, Qiong
    Yu, Songlin
    Hu, Jie
    Cao, Liujuan
    Ji, Rongrong
    IEEE TRANSACTIONS ON MULTIMEDIA, 2024, 26 : 2966 - 2977
  • [30] Dual Context Perception Transformer for Referring Image Segmentation
    Kong, Yuqiu
    Liu, Junhua
    Yao, Cuili
    PATTERN RECOGNITION AND COMPUTER VISION, PT V, PRCV 2024, 2025, 15035 : 216 - 230