Cross-coupled prompt learning for few-shot image recognition☆

被引：0

作者：

Zhang, Fangyuan ^{[1
]}

Wei, Rukai ^{[2
]}

Xie, Yanzhao ^{[1
]}

Wang, Yangtao ^{[1
]}

Tan, Xin ^{[3
]}

Ma, Lizhuang ^{[4
]}

Tang, Maobin ^{[4
]}

Fan, Lisheng ^{[1
]}

机构：

[1] Guangzhou Univ, Guangzhou Higher Educ Mega Ctr, Sch Civil Engn, 230 Wai Huan Xi Rd, Guangzhou 510006, Peoples R China

[2] Huazhong Univ Sci & Technol, Wuhan Natl Lab Optoelect, Luoyu Rd 1037, Wuhan 430074, Peoples R China

[3] East China Normal Univ, 3663 North Zhongshan Rd, Shanghai 200062, Peoples R China

[4] Shanghai Jiao Tong Univ, 800 Dongchuan RD, Shanghai 200240, Peoples R China

来源：

DISPLAYS | 2024年 / 85卷

基金：

中国国家自然科学基金;

关键词：

Prompt learning; Image recognition; Few-shot; Cross-attention;

D O I：

10.1016/j.displa.2024.102862

中图分类号：

TP3 [计算技术、计算机技术];

学科分类号：

0812 ;

摘要：

Prompt learning based on large models shows great potential to reduce training time and resource costs, which has been progressively applied to visual tasks such as image recognition. Nevertheless, the existing prompt learning schemes suffer from either inadequate prompt information from a single modality or insufficient prompt interaction between multiple modalities, resulting in low efficiency and performance. To address these limitations, we propose a Cross-Coupled Prompt Learning (CCPL) architecture, which is designed with two novel components (i.e., Cross-Coupled Prompt Generator (CCPG) module and Cross-Modal Fusion (CMF) module) to achieve efficient interaction between visual and textual prompts. Specifically, the CCPG module incorporates a cross-attention mechanism to automatically generate visual and textual prompts, each of which will be adaptively updated using the self-attention mechanism in their respective image and text encoders. Furthermore, the CMF module implements a deep fusion to reinforce the cross-modal feature interaction from the output layer with the Image-Text Matching (ITM) loss function. We conduct extensive experiments on 8 image datasets. The experimental results verify that our proposed CCPL outperforms the SOTA methods on few- shot image recognition tasks. The source code of this project is released at: https://github.com/elegantTechie/ CCPL.

引用

页数：11

共 50 条

[1] Semantic Prompt for Few-Shot Image Recognition
Chen, Wentao
Si, Chenyang
Zhang, Zhang
Wang, Liang
Wang, Zilei
Tan, Tieniu
2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, : 23581 - 23591
[2] Self-Prompt Mechanism for Few-Shot Image Recognition
Song, Mingchen
Wang, Huiqiang
Zhong, Guoqiang
THIRTY-EIGHTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 38 NO 5, 2024, : 4934 - 4942
[3] Few-Shot Composition Learning for Image Retrieval with Prompt Tuning
Wu, Junda
Wang, Rui
Zhao, Handong
Zhang, Ruiyi
Lu, Chaochao
Li, Shuai
Henao, Ricardo
THIRTY-SEVENTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 37 NO 4, 2023, : 4729 - 4737
[4] Few-Shot Image Classification Method Based on Visual Language Prompt Learning
Li B.
Wang X.
Teng S.
Lyu X.
Beijing Youdian Daxue Xuebao/Journal of Beijing University of Posts and Telecommunications, 2024, 47 (02): : 11 - 17
[5] Few-shot learning for ear recognition
Zhang, Jie
Yu, Wen
Yang, Xudong
Deng, Fang
PROCEEDINGS OF 2019 INTERNATIONAL CONFERENCE ON IMAGE, VIDEO AND SIGNAL PROCESSING (IVSP 2019), 2019, : 50 - 54
[6] Efficient Cross-Task Prompt Tuning for Few-Shot Conversational Emotion Recognition
Xu, Yige
Zeng, Zhiwei
Shen, Zhiqi
FINDINGS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (EMNLP 2023), 2023, : 11654 - 11666
[7] Few-Shot Learning for Image Denoising
Jiang, Bo
Lu, Yao
Zhang, Bob
Lu, Guangming
IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2023, 33 (09) : 4741 - 4753
[8] Decomposed Two-Stage Prompt Learning for Few-Shot Named Entity Recognition
Ye, Feiyang
Huang, Liang
Liang, Senjie
Chi, KaiKai
INFORMATION, 2023, 14 (05)
[9] Few-Shot Image Recognition with Knowledge Transfer
Peng, Zhimao
Li, Zechao
Zhang, Junge
Li, Yan
Qi, Guo-Jun
Tang, Jinhui
2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, : 441 - 449
[10] A Two-Stage Approach to Few-Shot Learning for Image Recognition
Das, Debasmit
Lee, C. S. George
IEEE TRANSACTIONS ON IMAGE PROCESSING, 2020, 29 : 3336 - 3350

← 1 2 3 4 5 →