TEST: Temporal-spatial separated transformer for temporal action localization

被引:0
|
作者
Wan, Herun [1 ,2 ,3 ]
Luo, Minnan [1 ,2 ,3 ]
Li, Zhihui [4 ]
Wang, Yang [5 ]
机构
[1] Xi An Jiao Tong Univ, Sch Comp Sci & Technol, Xian 710049, Peoples R China
[2] Xi An Jiao Tong Univ, Minist Educ, Key Lab Intelligent Networks & Network Secur, Xian 710049, Peoples R China
[3] Xi An Jiao Tong Univ, Shaanxi Prov Key Lab Big Data Knowledge Engn, Xian 710049, Peoples R China
[4] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230026, Peoples R China
[5] Xi An Jiao Tong Univ, Sch Continuing Educ, Xian 710049, Peoples R China
基金
中国国家自然科学基金;
关键词
Video transformer; Temporal action localization; High efficiency; NETWORK;
D O I
10.1016/j.neucom.2024.128688
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Temporal action localization is a fundamental task in video understanding. Existing methods fall into three categories: anchor-based, actionness-guided, and anchor-free. Anchor-based and actionness-guided models need huge computation resources to process redundant proposals or enumerate every possible proposal. Anchor- free models with lighter parameters become a more attractive option as temporal actions become more complex. However, they typically struggle to achieve high performance due to the need to aggregate global temporal-spatial features at every time step. To overcome this limitation, we design three efficient transformer- based architectures, bringing two advantages: (i) the global receptive field of transformers enables models to aggregate spatial and temporal at each time step, and (ii) the transformers could capture the moment-level feature, enhancing localization performance. Our designed architectures are adapted to any framework, thus we propose a simple but effective anchor-free framework named TEST. Compared to strong baselines, TEST achieves 0.96% to 3.20% improvement on two real-world datasets. Meanwhile, it improves time efficiency by 1.36 times and space efficiency by 1.08 times. Further experiments prove the effectiveness of TEST's modules. Implementation of our work is available at https://github.com/whr000001/TeST.
引用
收藏
页数:9
相关论文
共 50 条
  • [1] A Multitemporal Scale and Spatial-Temporal Transformer Network for Temporal Action Localization
    Gao, Zan
    Cui, Xinglei
    Zhuo, Tao
    Cheng, Zhiyong
    Liu, An-An
    Wang, Meng
    Chen, Shenyong
    IEEE TRANSACTIONS ON HUMAN-MACHINE SYSTEMS, 2023, 53 (03) : 569 - 580
  • [2] Temporal-Spatial Mapping for Action Recognition
    Song, Xiaolin
    Lan, Cuiling
    Zeng, Wenjun
    Xing, Junliang
    Sun, Xiaoyan
    Yang, Jingyu
    IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2020, 30 (03) : 748 - 759
  • [3] EEG temporal-spatial transformer for person identification
    Du, Yang
    Xu, Yongling
    Wang, Xiaoan
    Liu, Li
    Ma, Pengcheng
    SCIENTIFIC REPORTS, 2022, 12 (01)
  • [4] Temporal Deformable Transformer for Action Localization
    Wang, Haoying
    Wei, Ping
    Liu, Meiqin
    Zheng, Nanning
    ARTIFICIAL NEURAL NETWORKS AND MACHINE LEARNING, ICANN 2023, PT VI, 2023, 14259 : 563 - 575
  • [5] Design on temporal-spatial Transformer model for air target intention recognition
    Wang, Ke
    Li, Chenghai
    Song, Yafei
    Wang, Peng
    Li, Lemin
    Xibei Gongye Daxue Xuebao/Journal of Northwestern Polytechnical University, 2024, 42 (04): : 753 - 763
  • [6] Temporal-spatial unpredictable auditory information modulates temporal-spatial coincident audiovisual integrationa
    Li, Qi
    Yang, Jingjing
    Wu, Jinglong
    2013 ICME INTERNATIONAL CONFERENCE ON COMPLEX MEDICAL ENGINEERING (CME), 2013, : 31 - 34
  • [7] An Adaptive Dual Selective Transformer for Temporal Action Localization
    Li, Qiang
    Zu, Guang
    Xu, Hui
    Kong, Jun
    Zhang, Yanni
    Wang, Jianzhong
    IEEE TRANSACTIONS ON MULTIMEDIA, 2024, 26 : 7398 - 7412
  • [8] EFFICIENT TEMPORAL-SPATIAL FEATURE GROUPING FOR VIDEO ACTION RECOGNITION
    Qiu, Zhikang
    Zhao, Xu
    Hu, Zhilan
    2020 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2020, : 2176 - 2180
  • [9] System with temporal-spatial noise
    Li, JH
    PHYSICAL REVIEW E, 2003, 67 (06):
  • [10] Action recognition and localization with spatial and temporal contexts
    Xu, Wanru
    Miao, Zhenjiang
    Yu, Jian
    Ji, Qiang
    NEUROCOMPUTING, 2019, 333 : 351 - 363