EmoComicNet: A multi-task model for comic emotion recognition

被引：4

作者：

Dutta, Arpita ^{[1
,2
]}

Biswas, Samit ^{[1
]}

Das, Amit Kumar ^{[1
]}

机构：

[1] Indian Inst Engn Science&Technol, Dept Comp Science&Technol, Howrah 711103, West Bengal, India

[2] Techno Main, Artificial Intelligence & Machine Learning, Dept Comp Sci & Engn, Kolkata 700091, West Bengal, India

来源：

PATTERN RECOGNITION | 2024年 / 150卷

关键词：

Comic analysis; Multi-modal emotion recognition; Document image processing; Deep learning; Multi-task learning;

D O I：

10.1016/j.patcog.2024.110261

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

The emotion and sentiment associated with comic scenes can provide potential information for inferring the context of comic stories, which is an essential pre -requisite for developing comics' automatic content understanding tools. Here, we address this open area of comic research by exploiting the multi -modal nature of comics. The general assumptions for multi -modal sentiment analysis methods are that both image and text modalities are always present at the test phase. However, this assumption is not always satisfied for comics since comic characters' facial expressions, gestures, etc., are not always clearly visible. Also, the dialogues between comic characters are often challenging to comprehend the underlying context. To deal with these constraints of comic emotion analysis, we propose a multi -task -based framework, namely EmoComicNet, to fuse multi -modal information (i.e., both image and text) if it is available. However, the proposed EmoComicNet is designed to perform even when any modality is weak or completely missing. The proposed method potentially improves the overall performance. Besides, EmoComicNet can also deal with the problem of weak or absent modality during the training phase.

引用

页数：11

共 50 条

[1] Speech Emotion Recognition with Multi-task Learning
Cai, Xingyu
Yuan, Jiahong
Zheng, Renjie
Huang, Liang
Church, Kenneth
INTERSPEECH 2021, 2021, : 4508 - 4512
[2] Multi-task Learning for Speech Emotion and Emotion Intensity Recognition
Yue, Pengcheng
Qu, Leyuan
Zheng, Shukai
Li, Taihao
PROCEEDINGS OF 2022 ASIA-PACIFIC SIGNAL AND INFORMATION PROCESSING ASSOCIATION ANNUAL SUMMIT AND CONFERENCE (APSIPA ASC), 2022, : 1232 - 1237
[3] Multi-Task Emotion Recognition Based on Dimensional Model and Category Label
Huo, Yi
Ge, Yun
IEEE ACCESS, 2024, 12 : 75169 - 75179
[4] Multi-task Model for Comic Book Image Analysis
Nhu-Van Nguyen
Rigaud, Christophe
Burie, Jean-Christophe
MULTIMEDIA MODELING, MMM 2019, PT II, 2019, 11296 : 637 - 649
[5] Meta Multi-task Learning for Speech Emotion Recognition
Cai, Ruichu
Guo, Kaibin
Xu, Boyan
Yang, Xiaoyan
Zhang, Zhenjie
INTERSPEECH 2020, 2020, : 3336 - 3340
[6] Emotion Recognition With Sequential Multi-task Learning Technique
Phan Tran Dac Thinh
Hoang Manh Hung
Yang, Hyung-Jeong
Kim, Soo-Hyung
Lee, Guee-Sang
2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION WORKSHOPS (ICCVW 2021), 2021, : 3586 - 3589
[7] Speech Emotion Recognition based on Multi-Task Learning
Zhao, Huijuan
Han Zhijie
Wang, Ruchuan
2019 IEEE 5TH INTL CONFERENCE ON BIG DATA SECURITY ON CLOUD (BIGDATASECURITY) / IEEE INTL CONFERENCE ON HIGH PERFORMANCE AND SMART COMPUTING (HPSC) / IEEE INTL CONFERENCE ON INTELLIGENT DATA AND SECURITY (IDS), 2019, : 186 - 188
[8] Facial Emotion Recognition with Noisy Multi-task Annotations
Zhang, Siwei
Huang, Zhiwu
Paudel, Danda Pani
Van Gool, Luc
2021 IEEE WINTER CONFERENCE ON APPLICATIONS OF COMPUTER VISION (WACV 2021), 2021, : 21 - 31
[9] MMER: Multimodal Multi-task Learning for Speech Emotion Recognition
Ghosh, Sreyan
Tyagi, Utkarsh
Ramaneswaran, S.
Srivastava, Harshvardhan
Manocha, Dinesh
INTERSPEECH 2023, 2023, : 1209 - 1213
[10] Multi-Task and Attention Collaborative Network for Facial Emotion Recognition
Wang, Xiaohua
Yu, Cong
Gu, Yu
Hu, Min
Ren, Fuji
IEEJ TRANSACTIONS ON ELECTRICAL AND ELECTRONIC ENGINEERING, 2021, 16 (04) : 568 - 576

← 1 2 3 4 5 →