Distillation and Supplementation of Features for Referring Image Segmentation

被引：0

作者：

Tan, Zeyu ^{[1
]}

Xu, Dahong ^{[1
]}

Li, Xi ^{[1
]}

Liu, Hong ^{[1
]}

机构：

[1] Hunan Normal Univ, Coll Informat Sci & Engn, Changsha 410081, Peoples R China

来源：

IEEE ACCESS | 2024年 / 12卷

关键词：

Training; Image segmentation; Visualization; Feature extraction; Linguistics; Image reconstruction; Decoding; Filtering; Transformers; Accuracy; Referring image segmentation; multi-modal task; vision-language understanding;

D O I：

10.1109/ACCESS.2024.3482108

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Referring Image Segmentation (RIS) aims to accurately match specific instance objects in an input image with natural language expressions and generate corresponding pixel-level segmentation masks. Existing methods typically obtain multi-modal features by fusing linguistic features with visual features, which are fed into a mask decoder to generate segmentation masks. However, these methods ignore interfering noise in the multi-modal features that will adversely affect the generation of the target segmentation masks. In addition, the vast majority of current RIS models incorporate only a residual structure derived from a block within the Transformer model. The limitations of this information propagation approach hinder the stratification of the model structure, consequently affecting the training efficacy of the model. In this paper, we propose a RIS method called DSFRIS, which combines the knowledge of sparse reconstruction and employs a novel training mechanism in the process of training the decoder. Specifically, we propose a feature distillation mechanism for the multi-modal feature fusion stage and a feature supplementation mechanism for the mask decoder training process, which are two novel mechanisms for reducing the noise information in the multi-modal fusion features and enriching the feature information in the decoder training process, respectively. Through extensive experiments on three widely used RIS benchmark datasets, we demonstrate the state-of-the-art performance of our proposed method.

引用

页码：171269 / 171279

页数：11

共 50 条

[21] Dual Convolutional LSTM Network for Referring Image Segmentation
Ye, Linwei
Liu, Zhi
Wang, Yang
IEEE TRANSACTIONS ON MULTIMEDIA, 2020, 22 (12) : 3224 - 3235
[22] A survey of methods for addressing the challenges of referring image segmentation
Ji, Lixia
Du, Yunlong
Dang, Yiping
Gao, Wenzhao
Zhang, Han
NEUROCOMPUTING, 2024, 583
[23] Locate then Segment: A Strong Pipeline for Referring Image Segmentation
Jing, Ya
Kong, Tao
Wang, Wei
Wang, Liang
Li, Lei
Tan, Tieniu
2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 9853 - 9862
[24] Learning From Box Annotations for Referring Image Segmentation
Feng, Guang
Zhang, Lihe
Hu, Zhiwei
Lu, Huchuan
IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2024, 35 (03) : 3927 - 3937
[25] CARIS: Context-Aware Referring Image Segmentation
Liu, Sun-Ao
Zhang, Yiheng
Qiu, Zhaofan
Xie, Hongtao
Zhang, Yongdong
Yao, Ting
PROCEEDINGS OF THE 31ST ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2023, 2023, : 779 - 788
[26] Query Reconstruction Network for Referring Expression Image Segmentation
Shi, Hengcan
Li, Hongliang
Wu, Qingbo
Ngan, King Ngi
IEEE TRANSACTIONS ON MULTIMEDIA, 2021, 23 : 995 - 1007
[27] PRNet: A Progressive Refinement Network for referring image segmentation
Liu, Jing
Jiang, Huajie
Hu, Yongli
Yin, Baocai
NEUROCOMPUTING, 2025, 630
[28] A CONTEXT-BASED NETWORK FOR REFERRING IMAGE SEGMENTATION
Li, Xinyu
Liu, Yu
Xu, Kaiping
Zhao, Zhehuan
Liu, Sipei
2020 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2020, : 1436 - 1440
[29] Advancing Referring Expression Segmentation Beyond Single Image
Wu, Yixuan
Zhang, Zhao
Xie, Chi
Zhu, Feng
Zhao, Rui
2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION, ICCV, 2023, : 2628 - 2638
[30] Referring Image Segmentation via Recurrent Refinement Networks
Li, Ruiyu
Li, Kaican
Kuo, Yi-Chun
Shu, Michelle
Qi, Xiaojuan
Shen, Xiaoyong
Jia, Jiaya
2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, : 5745 - 5753

← 1 2 3 4 5 →