Distillation and Supplementation of Features for Referring Image Segmentation

被引：0

作者：

Tan, Zeyu ^{[1
]}

Xu, Dahong ^{[1
]}

Li, Xi ^{[1
]}

Liu, Hong ^{[1
]}

机构：

[1] Hunan Normal Univ, Coll Informat Sci & Engn, Changsha 410081, Peoples R China

来源：

IEEE ACCESS | 2024年 / 12卷

关键词：

Training; Image segmentation; Visualization; Feature extraction; Linguistics; Image reconstruction; Decoding; Filtering; Transformers; Accuracy; Referring image segmentation; multi-modal task; vision-language understanding;

D O I：

10.1109/ACCESS.2024.3482108

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Referring Image Segmentation (RIS) aims to accurately match specific instance objects in an input image with natural language expressions and generate corresponding pixel-level segmentation masks. Existing methods typically obtain multi-modal features by fusing linguistic features with visual features, which are fed into a mask decoder to generate segmentation masks. However, these methods ignore interfering noise in the multi-modal features that will adversely affect the generation of the target segmentation masks. In addition, the vast majority of current RIS models incorporate only a residual structure derived from a block within the Transformer model. The limitations of this information propagation approach hinder the stratification of the model structure, consequently affecting the training efficacy of the model. In this paper, we propose a RIS method called DSFRIS, which combines the knowledge of sparse reconstruction and employs a novel training mechanism in the process of training the decoder. Specifically, we propose a feature distillation mechanism for the multi-modal feature fusion stage and a feature supplementation mechanism for the mask decoder training process, which are two novel mechanisms for reducing the noise information in the multi-modal fusion features and enriching the feature information in the decoder training process, respectively. Through extensive experiments on three widely used RIS benchmark datasets, we demonstrate the state-of-the-art performance of our proposed method.

引用

页码：171269 / 171279

页数：11

共 50 条

[41] Global Selection and Local Attention Network for Referring Image Segmentation
Ding, Haixin
Zhang, Shengchuan
Cao, Liujuan
PATTERN RECOGNITION AND COMPUTER VISION, PRCV 2023, PT VII, 2024, 14431 : 284 - 295
[42] Bidirectional Relationship Inferring Network for Referring Image Localization and Segmentation
Feng, Guang
Hu, Zhiwei
Zhang, Lihe
Sun, Jiayu
Lu, Huchuan
IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2023, 34 (05) : 2246 - 2258
[43] Global and Local Interactive Perception Network for Referring Image Segmentation
Liu, Jing
Tan, Hongchen
Hu, Yongli
Sun, Yanfeng
Wang, Huasheng
Yin, Baocai
IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2023, 35 (12) : 1 - 14
[44] Comprehensive Multi-Modal Interactions for Referring Image Segmentation
Jain, Kanishk
Gandhi, Vineet
FINDINGS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2022), 2022, : 3427 - 3435
[45] Vision-Aware Language Reasoning for Referring Image Segmentation
Fayou Xu
Bing Luo
Chao Zhang
Li Xu
Mingxing Pu
Bo Li
Neural Processing Letters, 2023, 55 : 11313 - 11331
[46] See-Through-Text Grouping for Referring Image Segmentation
Chen, Ding-Jie
Jia, Songhao
Lo, Yi-Chen
Chen, Hwann-Tzong
Liu, Tyng-Luh
2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, : 7453 - 7462
[47] Bottom-Up Shift and Reasoning for Referring Image Segmentation
Yang, Sibei
Xia, Meng
Li, Guanbin
Zhou, Hong-Yu
Yu, Yizhou
2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 11261 - 11270
[48] De-noising mask transformer for referring image segmentation
Wang, Yehui
Lei, Fang
Wang, Baoyan
Zhang, Qiang
Zhen, Xiantong
Zhang, Lei
IMAGE AND VISION COMPUTING, 2025, 154
[49] Text-Vision Relationship Alignment for Referring Image Segmentation
Mingxing Pu
Bing Luo
Chao Zhang
Li Xu
Fayou Xu
Mingming Kong
Neural Processing Letters, 56
[50] Shatter and Gather: Learning Referring Image Segmentation with Text Supervision
Kim, Dongwon
Kim, Namyup
Lan, Cuiling
Kwak, Suha
2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, : 15501 - 15511

← 1 2 3 4 5 →