Learning Spatial and Temporal Cues for Multi-label Facial Action Unit Detection

被引:85
|
作者
Chu, Wen-Sheng [1 ]
De la Torre, Fernando [1 ]
Cohn, Jeffrey F. [1 ,2 ]
机构
[1] Carnegie Mellon Univ, Robot Inst, Pittsburgh, PA 15213 USA
[2] Univ Pittsburgh, Dept Psychol, Pittsburgh, PA 15260 USA
基金
美国国家卫生研究院;
关键词
EXPRESSIONS; EMOTION;
D O I
10.1109/FG.2017.13
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Facial action units (AU) are the fundamental units to decode human facial expressions. At least three aspects affect performance of automated AU detection: spatial representation, temporal modeling, and AU correlation. Unlike most studies that tackle these aspects separately, we propose a hybrid network architecture to jointly model them. Specifically, spatial representations are extracted by a Convolutional Neural Network (CNN), which, as analyzed in this paper, is able to reduce person-specific biases caused by hand-crafted descriptors (e.g., HOG and Gabor). To model temporal dependencies, Long Short-Term Memory (LSTMs) are stacked on top of these representations, regardless of the lengths of input videos. The outputs of CNNs and LSTMs are further aggregated into a fusion network to produce per-frame prediction of 12 AUs. Our network naturally addresses the three issues together, and yields superior performance compared to existing methods that consider these issues independently. Extensive experiments were conducted on two large spontaneous datasets, GFT and BP4D, with more than 400,000 frames coded with 12 AUs. On both datasets, we report improvements over a standard multi-label CNN and feature-based state-of-the-art. Finally, we provide visualization of the learned AU models, which, to our best knowledge, reveal how machines see AUs for the first time.
引用
收藏
页码:25 / 32
页数:8
相关论文
共 50 条
  • [1] Deep Region and Multi-label Learning for Facial Action Unit Detection
    Zhao, Kaili
    Chu, Wen-Sheng
    Zhang, Honggang
    2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, : 3391 - 3399
  • [2] Joint Patch and Multi-label Learning for Facial Action Unit Detection
    Zhao, Kaili
    Chu, Wen-Sheng
    De la Torre, Fernando
    Cohn, Jeffrey F.
    Zhang, Honggang
    2015 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2015, : 2207 - 2216
  • [3] Learning facial action units with spatiotemporal cues and multi-label sampling
    Chu, Wen-Sheng
    De la Torre, Fernando
    Cohn, Jeffrey F.
    IMAGE AND VISION COMPUTING, 2019, 81 : 1 - 14
  • [4] Discriminant Multi-Label Manifold Embedding for Facial Action Unit Detection
    Yuce, Anil
    Gao, Hua
    Thiran, Jean-Philippe
    2015 11TH IEEE INTERNATIONAL CONFERENCE AND WORKSHOPS ON AUTOMATIC FACE AND GESTURE RECOGNITION (FG), VOL. 6, 2015,
  • [5] Region and Temporal Dependency Fusion for Multi-label Action Unit Detection
    Mei, Chuanneng
    Jiang, Fei
    Shen, Ruimin
    Hu, Qiaoping
    2018 24TH INTERNATIONAL CONFERENCE ON PATTERN RECOGNITION (ICPR), 2018, : 848 - 853
  • [6] Facial Action Unit Detection with Multilayer Fused Multi-Task and Multi-Label Deep Learning Network
    He, Jun
    Li, Dongliang
    Bo, Sun
    Yu, Lejun
    KSII TRANSACTIONS ON INTERNET AND INFORMATION SYSTEMS, 2019, 13 (11) : 5546 - 5559
  • [7] An Attention-based Method for Multi-label Facial Action Unit Detection
    Le Hoai, Duy
    Lim, Eunchae
    Choi, Eunbin
    Kim, Sieun
    Pant, Sudarshan
    Lee, Guee-Sang
    Kim, Soo-Huyng
    Yang, Hyung-Jeong
    2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION WORKSHOPS, CVPRW 2022, 2022, : 2453 - 2458
  • [8] Multi-label learning with missing labels for image annotation and facial action unit recognition
    Wu, Baoyuan
    Lyu, Siwei
    Hu, Bao-Gang
    Ji, Qiang
    PATTERN RECOGNITION, 2015, 48 (07) : 2279 - 2289
  • [9] Joint Patch and Multi-label Learning for Facial Action Unit and Holistic Expression Recognition
    Zhao, Kaili
    Chu, Wen-Sheng
    De la Torre, Fernando
    Cohn, Jeffrey F.
    Zhang, Honggang
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2016, 25 (08) : 3931 - 3946
  • [10] Dual DETRs for Multi-Label Temporal Action Detection
    Zhu, Yuhan
    Zhang, Guozhen
    Tan, Jing
    Wu, Gangshan
    Wang, Limin
    2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2024, : 18559 - 18569