Decomposed Soft Prompt Guided Fusion Enhancing for Compositional Zero-Shot Learning

被引:12
|
作者
Lu, Xiaocheng [1 ]
Guo, Song [1 ,2 ]
Liu, Ziming [1 ]
Guo, Jingcai [1 ,2 ]
机构
[1] Hong Kong Polytech Univ, Dept Comp, Hong Kong, Peoples R China
[2] Hong Kong Polytech Univ, Shenzhen Res Inst, Hong Kong, Peoples R China
基金
中国国家自然科学基金;
关键词
D O I
10.1109/CVPR52729.2023.02256
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Compositional Zero-Shot Learning (CZSL) aims to recognize novel concepts formed by known states and objects during training. Existing methods either learn the combined state-object representation, challenging the generalization of unseen compositions, or design two classifiers to identify state and object separately from image features, ignoring the intrinsic relationship between them. To jointly eliminate the above issues and construct a more robust CZSL system, we propose a novel framework termed Decomposed Fusion with Soft Prompt (DFSP)1, by involving vision-language models (VLMs) for unseen composition recognition. Specifically, DFSP constructs a vector combination of learnable soft prompts with state and object to establish the joint representation of them. In addition, a cross-modal decomposed fusion module is designed between the language and image branches, which decomposes state and object among language features instead of image features. Notably, being fused with the decomposed features, the image features can be more expressive for learning the relationship with states and objects, respectively, to improve the response of unseen compositions in the pair space, hence narrowing the domain gap between seen and unseen sets. Experimental results on three challenging benchmarks demonstrate that our approach significantly outperforms other state-of-the-art methods by large margins.
引用
收藏
页码:23560 / 23569
页数:10
相关论文
共 50 条
  • [1] Hierarchical Prompt Learning for Compositional Zero-Shot Recognition
    Wang, Henan
    Yang, Muli
    Wei, Kun
    Deng, Cheng
    PROCEEDINGS OF THE THIRTY-SECOND INTERNATIONAL JOINT CONFERENCE ON ARTIFICIAL INTELLIGENCE, IJCAI 2023, 2023, : 1470 - 1478
  • [2] Adaptive Fusion Learning for Compositional Zero-Shot Recognition
    Min, Lingtong
    Fan, Ziman
    Wang, Shunzhou
    Dou, Feiyang
    Li, Xin
    Wang, Binglu
    IEEE TRANSACTIONS ON MULTIMEDIA, 2025, 27 : 1193 - 1204
  • [3] Enhancing Zero-Shot Stance Detection with Contrastive and Prompt Learning
    Yao, Zhenyin
    Yang, Wenzhong
    Wei, Fuyuan
    ENTROPY, 2024, 26 (04)
  • [4] Knowledge Guided Transformer Network for Compositional Zero-Shot Learning
    Panda, Aditya
    Prasad, Dipti
    ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS, 2024, 20 (11)
  • [5] Zero-Shot Compositional Concept Learning
    Xu, Guangyue
    Kordjamshidi, Parisa
    Chai, Joyce Y.
    1ST WORKSHOP ON META LEARNING AND ITS APPLICATIONS TO NATURAL LANGUAGE PROCESSING (METANLP 2021), 2021, : 19 - 27
  • [6] Open World Compositional Zero-Shot Learning
    Mancini, Massimiliano
    Naeem, Muhammad Ferjad
    Xian, Yongqin
    Akata, Zeynep
    2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 5218 - 5226
  • [7] Learning the Compositional Domains for Generalized Zero-shot Learning
    Dong, Hanze
    Fu, Yanwei
    Hwang, Sung Ju
    Sigal, Leonid
    Xue, Xiangyang
    COMPUTER VISION AND IMAGE UNDERSTANDING, 2022, 221
  • [8] Learning Attention Propagation for Compositional Zero-Shot Learning
    Khan, Muhammad Gul Zain Ali
    Naeem, Muhammad Ferjad
    Van Gool, Luc
    Pagani, A.
    Stricker, Didier
    Afzal, Muhammad Zeshan
    2023 IEEE/CVF WINTER CONFERENCE ON APPLICATIONS OF COMPUTER VISION (WACV), 2023, : 3817 - 3826
  • [9] Learning Attention as Disentangler for Compositional Zero-shot Learning
    Hao, Shaozhe
    Han, Kai
    Wong, Kwan-Yee K.
    2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, : 15315 - 15324
  • [10] Learning Graph Embeddings for Compositional Zero-shot Learning
    Naeem, Muhammad Ferjad
    Xian, Yongqin
    Tombari, Federico
    Akata, Zeynep
    2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 953 - 962