Distillation and Supplementation of Features for Referring Image Segmentation

被引：0

作者：

Tan, Zeyu ^{[1
]}

Xu, Dahong ^{[1
]}

Li, Xi ^{[1
]}

Liu, Hong ^{[1
]}

机构：

[1] Hunan Normal Univ, Coll Informat Sci & Engn, Changsha 410081, Peoples R China

来源：

IEEE ACCESS | 2024年 / 12卷

关键词：

Training; Image segmentation; Visualization; Feature extraction; Linguistics; Image reconstruction; Decoding; Filtering; Transformers; Accuracy; Referring image segmentation; multi-modal task; vision-language understanding;

D O I：

10.1109/ACCESS.2024.3482108

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Referring Image Segmentation (RIS) aims to accurately match specific instance objects in an input image with natural language expressions and generate corresponding pixel-level segmentation masks. Existing methods typically obtain multi-modal features by fusing linguistic features with visual features, which are fed into a mask decoder to generate segmentation masks. However, these methods ignore interfering noise in the multi-modal features that will adversely affect the generation of the target segmentation masks. In addition, the vast majority of current RIS models incorporate only a residual structure derived from a block within the Transformer model. The limitations of this information propagation approach hinder the stratification of the model structure, consequently affecting the training efficacy of the model. In this paper, we propose a RIS method called DSFRIS, which combines the knowledge of sparse reconstruction and employs a novel training mechanism in the process of training the decoder. Specifically, we propose a feature distillation mechanism for the multi-modal feature fusion stage and a feature supplementation mechanism for the mask decoder training process, which are two novel mechanisms for reducing the noise information in the multi-modal fusion features and enriching the feature information in the decoder training process, respectively. Through extensive experiments on three widely used RIS benchmark datasets, we demonstrate the state-of-the-art performance of our proposed method.

引用

页码：171269 / 171279

页数：11

共 50 条

[31] Bilateral Knowledge Interaction Network for Referring Image Segmentation
Ding, Haixin
Zhang, Shengchuan
Wu, Qiong
Yu, Songlin
Hu, Jie
Cao, Liujuan
Ji, Rongrong
IEEE TRANSACTIONS ON MULTIMEDIA, 2024, 26 : 2966 - 2977
[32] Dual Context Perception Transformer for Referring Image Segmentation
Kong, Yuqiu
Liu, Junhua
Yao, Cuili
PATTERN RECOGNITION AND COMPUTER VISION, PT V, PRCV 2024, 2025, 15035 : 216 - 230
[33] Medical image segmentation based on multi-layer features and spatial information distillation
Zheng Y.
Hao P.
Wu D.
Bai C.
Beijing Hangkong Hangtian Daxue Xuebao/Journal of Beijing University of Aeronautics and Astronautics, 2022, 48 (08): : 1409 - 1417
[34] Text-Vision Relationship Alignment for Referring Image Segmentation
Pu, Mingxing
Luo, Bing
Zhang, Chao
Xu, Li
Xu, Fayou
Kong, Mingming
NEURAL PROCESSING LETTERS, 2024, 56 (02)
[35] Local-global coordination with transformers for referring image segmentation
Liu, Fang
Kong, Yuqiu
Zhang, Lihe
Feng, Guang
Yin, Baocai
NEUROCOMPUTING, 2023, 522 : 39 - 52
[36] Prompt-Driven Referring Image Segmentation with Instance Contrasting
Shang, Chao
Song, Zichen
Qiu, Heqian
Wang, Lanxiao
Meng, Fanman
Li, Hongliang
2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2024, 2024, : 4124 - 4134
[37] Referring Image Segmentation via Language-Driven Attention
Chen, Ding-Jie
Hsieh, He-Yen
Liu, Tyng-Luh
2021 IEEE INTERNATIONAL CONFERENCE ON ROBOTICS AND AUTOMATION (ICRA 2021), 2021, : 13997 - 14003
[38] Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation
Yan, Yichen
He, Xingjian
Chen, Sihan
Liu, Jing
PROCEEDINGS OF THE 4TH ANNUAL ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA RETRIEVAL, ICMR 2024, 2024, : 451 - 459
[39] Beyond One-to-One: Rethinking the Referring Image Segmentation
Hu, Yutao
Wang, Qixiong
Shao, Wenqi
Xie, Enze
Li, Zhenguo
Han, Jungong
Luo, Ping
2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION, ICCV, 2023, : 4044 - 4054
[40] Vision-Aware Language Reasoning for Referring Image Segmentation
Xu, Fayou
Luo, Bing
Zhang, Chao
Xu, Li
Pu, Mingxing
Li, Bo
NEURAL PROCESSING LETTERS, 2023, 55 (08) : 11313 - 11331

← 1 2 3 4 5 →