On the Imaginary Wings: Text-Assisted Complex-Valued Fusion Network for Fine-Grained Visual Classification

被引:10
|
作者
Guan, Xiang [1 ]
Yang, Yang [1 ]
Li, Jingjing [1 ]
Zhu, Xiaofeng [1 ]
Song, Jingkuan [1 ]
Shen, Heng Tao [1 ,2 ]
机构
[1] Univ Elect Sci & Technol China, Ctr Future Media, Chengdu 611731, Peoples R China
[2] Peng Cheng Lab, Shenzhen 518066, Peoples R China
基金
中国国家自然科学基金;
关键词
Complex values; fine-grained visual classification (FGVC); graph convolutional networks (GCNs); multimodal;
D O I
10.1109/TNNLS.2021.3126046
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Fine-grained visual classification (FGVC) is challenging due to the interclass similarity and intraclass variation in datasets. In this work, we explore the great merit of complex values in introducing an imaginary part for modeling data uncertainty (e.g., different points on the complex plane can describe the same state) and graph convolutional networks (GCNs) in learning interdependently among classes to simultaneously tackle the above two major challenges. To the end, we propose a novel approach, termed text-assisted complex-valued fusion network (TA-CFN). Specifically, we expand each feature from 1-D real values to 2-D complex value by disassembling feature maps, thereby enabling the extension of traditional deep convolutional neural networks over the complex domain. Then, we fuse the real and imaginary parts of complex features through complex projection and modulus operation. Finally, we build an undirected graph over the object labels with the assistance of a text corpus, and a GCN is learned to map this graph into a set of classifiers. The benefits are in two folds: 1) complex features allow for a richer algebraic structure to better model the large variation within the same category and 2) leveraging the interclass dependencies brought by the GCN to capture key factors of the slight variation among different categories. We conduct extensive experiments to verify that our proposed model can achieve the state-of-the-art performance on two widely used FGVC datasets.
引用
收藏
页码:5112 / 5121
页数:10
相关论文
共 50 条
  • [1] Multiscale Progressive Complementary Fusion Network for Fine-Grained Visual Classification
    Lei, Jingsheng
    Yang, Xinqi
    Yang, Shengying
    IEEE ACCESS, 2022, 10 : 62800 - 62810
  • [2] Fine-Grained Visual Classification Network Based on Fusion Pooling and Attention Enhancement
    Xiao B.
    Guo J.
    Zhang X.
    Wang M.
    Moshi Shibie yu Rengong Zhineng/Pattern Recognition and Artificial Intelligence, 2023, 36 (07): : 661 - 670
  • [3] Visual Analytics for Fine-grained Text Classification Models and Datasets
    Battogtokh, M.
    Xing, Y.
    Davidescu, C.
    Abdul-Rahman, A.
    Luck, M.
    Borgo, R.
    COMPUTER GRAPHICS FORUM, 2024, 43 (03)
  • [4] SemLa: A Visual Analysis System for Fine-Grained Text Classification
    Battogtokh, Munkhtulga
    Davidescu, Cosmin
    Luck, Michael
    Borgo, Rita
    THIRTY-EIGTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 38 NO 21, 2024, : 23772 - 23774
  • [5] Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification
    Bai, Xiang
    Yang, Mingkun
    Lyu, Pengyuan
    Xu, Yongchao
    Luo, Jiebo
    IEEE ACCESS, 2018, 6 : 66322 - 66335
  • [6] Convolutionally Enhanced Feature Fusion Visual Transformer for Fine-Grained Visual Classification
    Huang, Min
    Zhu, Saixing
    Wang, Zehua
    Qu, Shuanghong
    2024 16TH INTERNATIONAL CONFERENCE ON MACHINE LEARNING AND COMPUTING, ICMLC 2024, 2024, : 447 - 452
  • [7] A collaborative gated attention network for fine-grained visual classification
    Zhu, Qiangxi
    Kuang, Wenlan
    Li, Zhixin
    DISPLAYS, 2023, 79
  • [8] WEB-SUPERVISED NETWORK FOR FINE-GRAINED VISUAL CLASSIFICATION
    Zhang, Chuanyi
    Ya, Yazhou
    Zhang, Jiachao
    Chen, Jiaxin
    Huang, Pu
    Zhang, Jian
    Tang, Zhenmin
    2020 IEEE INTERNATIONAL CONFERENCE ON MULTIMEDIA AND EXPO (ICME), 2020,
  • [9] PFNet: a novel part fusion network for fine-grained visual categorization
    Jingyun Liang
    Jinlin Guo
    Yanming Guo
    Songyang Lao
    Multimedia Tools and Applications, 2020, 79 : 33397 - 33416
  • [10] PFNet: a novel part fusion network for fine-grained visual categorization
    Liang, Jingyun
    Guo, Jinlin
    Guo, Yanming
    Lao, Songyang
    MULTIMEDIA TOOLS AND APPLICATIONS, 2020, 79 (45-46) : 33397 - 33416