Human-like Controllable Image Captioning with Verb-specific Semantic Roles

被引:42
|
作者
Chen, Long [2 ,3 ]
Jiang, Zhihong [1 ]
Xiao, Jun [1 ]
Liu, Wei [4 ]
机构
[1] Zhejiang Univ, Hangzhou, Peoples R China
[2] Tencent AI Lab, Bellevue, WA USA
[3] Columbia Univ, New York, NY 10027 USA
[4] Tencent Data Platform, New York, NY USA
基金
浙江省自然科学基金; 中国国家自然科学基金;
关键词
D O I
10.1109/CVPR46437.2021.01657
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Controllable Image Captioning (CIC) - generating image descriptions following designated control signals- has received unprecedented attention over the last few years. To emulate the human ability in controlling caption generation, current CIC studies focus exclusively on control signals concerning objective properties, such as contents of interest or descriptive patterns. However, we argue that almost all existing objective control signals have overlooked two indispensable characteristics of an ideal control signal: 1) Event-compatible: all visual contents referred to in a single sentence should be compatible with the described activity. 2) Sample-suitable: the control signals should be suitable for a specific image sample. To this end, we propose a new control signal for CIC: Verb-specific Semantic Roles (VSR). VSR consists of a verb and some semantic roles, which represents a targeted activity and the roles of entities involved in this activity. Given a designated VSR, we first train a grounded semantic role labeling (GSRL) model to identify and ground all entities for each role. Then, we propose a semantic structure planner (SSP) to learn human-like descriptive semantic structures. Lastly, we use a roleshift captioning model to generate the captions. Extensive experiments and ablations demonstrate that our framework can achieve better controllability than several strong baselines on two challenging CIC benchmarks. Besides, we can generate multi-level diverse captions easily.
引用
收藏
页码:16841 / 16851
页数:11
相关论文
共 40 条
  • [1] Thematic roles as verb-specific concepts
    McRae, K
    Ferretti, TR
    Amyote, L
    LANGUAGE AND COGNITIVE PROCESSES, 1997, 12 (2-3): : 137 - 176
  • [2] Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive Style
    Ge, Hongwei
    Yan, Zehang
    Zhang, Kai
    Zhao, Mingde
    Sun, Liang
    2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, : 1754 - 1763
  • [3] Distributional Learning in English: The Effect of Verb-Specific Biases and Verb-General Semantic Mappings on Sentence Production
    Thothathiri, Malathi
    Braiuca, Maria C.
    JOURNAL OF EXPERIMENTAL PSYCHOLOGY-LEARNING MEMORY AND COGNITION, 2021, 47 (01) : 113 - 128
  • [4] Say in Human-Like Way: Hierarchical Cross-modal Information Abstraction and Summarization for Controllable Captioning
    Wang, Xiaoyi
    Huang, Jun
    ARTIFICIAL NEURAL NETWORKS AND MACHINE LEARNING - ICANN 2021, PT I, 2021, 12891 : 217 - 228
  • [5] CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
    Basioti, Kalliopi
    Abdelsalam, Mohamed A.
    Fancellu, Federico
    Pavlovic, Vladimir
    Fazly, Afsaneh
    COMPUTER VISION - ECCV 2024, PT LXVI, 2025, 15124 : 444 - 461
  • [6] Accepting Human-like Avatars in Social and Professional Roles
    Sharma, Medha
    Vemuri, Kavita
    ACM TRANSACTIONS ON HUMAN-ROBOT INTERACTION, 2022, 11 (03)
  • [7] Semantic Similarity Measures Applied to an Ontology for Human-Like Interaction
    Albacete, Esperanza
    Calle, Javier
    Castro, Elena
    Cuadra, Dolores
    JOURNAL OF ARTIFICIAL INTELLIGENCE RESEARCH, 2012, 44 : 397 - 421
  • [8] Learning Accurate and Human-Like Driving using Semantic Maps and Attention
    Hecker, Simon
    Dai, Dengxin
    Liniger, Alexander
    Hahner, Martin
    Van Gool, Luc
    2020 IEEE/RSJ INTERNATIONAL CONFERENCE ON INTELLIGENT ROBOTS AND SYSTEMS (IROS), 2020, : 2346 - 2353
  • [9] Semantic Segmentation Optimization in Power Systems: Enhancing Human-Like Switching Operations
    Hua, Jin
    Zhao, Yue
    Zhang, Huijun
    Zhao, Haiming
    Wang, Lei
    TRAITEMENT DU SIGNAL, 2023, 40 (04) : 1401 - 1412
  • [10] A Human-Like Semantic Cognition Network for Aspect-Level Sentiment Classification
    Lei, Zeyang
    Yang, Yujiu
    Yang, Min
    Zhao, Wei
    Guo, Jun
    Liu, Yi
    THIRTY-THIRD AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE / THIRTY-FIRST INNOVATIVE APPLICATIONS OF ARTIFICIAL INTELLIGENCE CONFERENCE / NINTH AAAI SYMPOSIUM ON EDUCATIONAL ADVANCES IN ARTIFICIAL INTELLIGENCE, 2019, : 6650 - 6657