Semantic-Aware Data Augmentation for Text-to-Image Synthesis

被引:0
|
作者
Tan, Zhaorui [1 ,2 ]
Yang, Xi [1 ]
Huang, Kaizhu [3 ]
机构
[1] Xian Jiaotong Liverpool Univ, Dept Intelligent Sci, Suzhou, Peoples R China
[2] Univ Liverpool, Dept Comp Sci, Liverpool, Merseyside, England
[3] Duke Kunshan Univ, Data Sci Res Ctr, Suzhou, Peoples R China
基金
中国国家自然科学基金;
关键词
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Data augmentation has been recently leveraged as an effective regularizer in various vision-language deep neural networks. However, in text-to-image synthesis (T2Isyn), current augmentation wisdom still suffers from the semantic mismatch between augmented paired data. Even worse, semantic collapse may occur when generated images are less semantically constrained. In this paper, we develop a novel Semantic-aware Data Augmentation (SADA) framework dedicated to T2Isyn. In particular, we propose to augment texts in the semantic space via an Implicit Textual Semantic Preserving Augmentation, in conjunction with a specifically designed Image Semantic Regularization Loss as Generated Image Semantic Conservation, to cope well with semantic mismatch and collapse. As one major contribution, we theoretically show that Implicit Textual Semantic Preserving Augmentation can certify better text-image consistency while Image Semantic Regularization Loss regularizing the semantics of generated images would avoid semantic collapse and enhance image quality. Extensive experiments validate that SADA enhances text-image consistency and improves image quality significantly in T2Isyn models across various backbones. Especially, incorporating SADA during the tuning process of Stable Diffusion models also yields performance improvements.
引用
收藏
页码:5098 / 5107
页数:10
相关论文
共 50 条
  • [1] Semantic-aware data quality assessment for image big data
    Liu, Yu
    Wang, Yangtao
    Zhou, Ke
    Yang, Yujuan
    Liu, Yifei
    FUTURE GENERATION COMPUTER SYSTEMS-THE INTERNATIONAL JOURNAL OF ESCIENCE, 2020, 102 : 53 - 65
  • [2] Semantic-Aware Video Text Detection
    Feng, Wei
    Yin, Fei
    Zhang, Xu-Yao
    Liu, Cheng-Lin
    2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 1695 - 1705
  • [3] Inferring Semantic Layout for Hierarchical Text-to-Image Synthesis
    Hong, Seunghoon
    Yang, Dingdong
    Choi, Jongwook
    Lee, Honglak
    2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, : 7986 - 7994
  • [4] Semantic Object Accuracy for Generative Text-to-Image Synthesis
    Hinz, Tobias
    Heinrich, Stefan
    Wermter, Stefan
    IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2022, 44 (03) : 1552 - 1565
  • [5] Semantic Distance Adversarial Learning for Text-to-Image Synthesis
    Yuan, Bowen
    Sheng, Yefei
    Bao, Bing-Kun
    Chen, Yi-Ping Phoebe
    Xu, Changsheng
    IEEE TRANSACTIONS ON MULTIMEDIA, 2024, 26 : 1255 - 1266
  • [6] Bash comment generation via data augmentation and semantic-aware CodeBERT
    Shen, Yiheng
    Ju, Xiaolin
    Chen, Xiang
    Yang, Guang
    AUTOMATED SOFTWARE ENGINEERING, 2024, 31 (01)
  • [7] Bash comment generation via data augmentation and semantic-aware CodeBERT
    Yiheng Shen
    Xiaolin Ju
    Xiang Chen
    Guang Yang
    Automated Software Engineering, 2024, 31
  • [8] SEMANTIC-AWARE NETWORK FOR AERIAL-TO-GROUND IMAGE SYNTHESIS
    Jang, Jinhyun
    Song, Taeyong
    Sohn, Kwanghoon
    2021 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2021, : 3862 - 3866
  • [9] Learning to Generate Semantic Layouts for Higher Text-Image Correspondence in Text-to-Image Synthesis
    Park, Minho
    Yun, Jooyeol
    Choi, Seunghwan
    Choo, Jaegul
    2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION, ICCV, 2023, : 7557 - 7566
  • [10] Semantic-aware blind image quality assessment
    Siahaan, Ernestasia
    Hanjalic, Alan
    Redi, Judith A.
    SIGNAL PROCESSING-IMAGE COMMUNICATION, 2018, 60 : 237 - 252