Semantic-Aware Data Augmentation for Text-to-Image Synthesis

被引:0
|
作者
Tan, Zhaorui [1 ,2 ]
Yang, Xi [1 ]
Huang, Kaizhu [3 ]
机构
[1] Xian Jiaotong Liverpool Univ, Dept Intelligent Sci, Suzhou, Peoples R China
[2] Univ Liverpool, Dept Comp Sci, Liverpool, Merseyside, England
[3] Duke Kunshan Univ, Data Sci Res Ctr, Suzhou, Peoples R China
基金
中国国家自然科学基金;
关键词
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Data augmentation has been recently leveraged as an effective regularizer in various vision-language deep neural networks. However, in text-to-image synthesis (T2Isyn), current augmentation wisdom still suffers from the semantic mismatch between augmented paired data. Even worse, semantic collapse may occur when generated images are less semantically constrained. In this paper, we develop a novel Semantic-aware Data Augmentation (SADA) framework dedicated to T2Isyn. In particular, we propose to augment texts in the semantic space via an Implicit Textual Semantic Preserving Augmentation, in conjunction with a specifically designed Image Semantic Regularization Loss as Generated Image Semantic Conservation, to cope well with semantic mismatch and collapse. As one major contribution, we theoretically show that Implicit Textual Semantic Preserving Augmentation can certify better text-image consistency while Image Semantic Regularization Loss regularizing the semantics of generated images would avoid semantic collapse and enhance image quality. Extensive experiments validate that SADA enhances text-image consistency and improves image quality significantly in T2Isyn models across various backbones. Especially, incorporating SADA during the tuning process of Stable Diffusion models also yields performance improvements.
引用
收藏
页码:5098 / 5107
页数:10
相关论文
共 50 条
  • [21] Semantic-Aware Visual Decomposition for Image Coding
    Chang, Jianhui
    Zhang, Jian
    Li, Jiguo
    Wang, Shiqi
    Mao, Qi
    Jia, Chuanmin
    Ma, Siwei
    Gao, Wen
    INTERNATIONAL JOURNAL OF COMPUTER VISION, 2023, 131 (09) : 2333 - 2355
  • [22] Text-to-Image Generation with Multiscale Semantic Context-Aware Generative Adversarial Networks
    Dong, Pei
    Wu, Lei
    Meng, Lei
    Meng, Xiangxu
    ADVANCED INTELLIGENT COMPUTING TECHNOLOGY AND APPLICATIONS, PT XII, ICIC 2024, 2024, 14873 : 192 - 203
  • [23] Generative Adversarial Networks with Adaptive Semantic Normalization for text-to-image synthesis
    Huang, Siyue
    Chen, Ying
    DIGITAL SIGNAL PROCESSING, 2022, 120
  • [24] Survey of text-to-image synthesis
    Cao Y.
    Qin J.
    Ma Q.
    Sun H.
    Yan K.
    Wang L.
    Ren J.
    Zhejiang Daxue Xuebao (Gongxue Ban)/Journal of Zhejiang University (Engineering Science), 2024, 58 (02): : 219 - 238
  • [25] Unsupervised text-to-image synthesis
    Dong, Yanlong
    Zhang, Ying
    Ma, Lin
    Wang, Zhi
    Luo, Jiebo
    Pattern Recognition, 2021, 110
  • [26] Unsupervised text-to-image synthesis
    Dong, Yanlong
    Zhang, Ying
    Ma, Lin
    Wang, Zhi
    Luo, Jiebo
    PATTERN RECOGNITION, 2021, 110
  • [27] Semantic-Aware Fingerprints of Symbolic Research Data
    Graebe, Hans-Gert
    MATHEMATICAL SOFTWARE, ICMS 2016, 2016, 9725 : 411 - 418
  • [28] A semantic-aware data generator for ETL workflows
    Du, Naiqiao
    Ye, Xiaojun
    Wang, Jianmin
    CONCURRENCY AND COMPUTATION-PRACTICE & EXPERIENCE, 2016, 28 (04): : 1016 - 1040
  • [29] Text-to-Image Generation Method Based on Image-Text Semantic Consistency
    Xue Z.
    Xu Z.
    Lang C.
    Feng S.
    Wang T.
    Li Y.
    Jisuanji Yanjiu yu Fazhan/Computer Research and Development, 2023, 60 (09): : 2180 - 2190
  • [30] Semantic-aware Responsive Listener Head Synthesis
    Zhao, Wei
    Xiao, Peng
    Zhang, Rongju
    Wang, Yijun
    Lin, Jianxin
    PROCEEDINGS OF THE 30TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2022, 2022, : 7065 - 7069