TEST: Temporal-spatial separated transformer for temporal action localization

被引:0
|
作者
Wan, Herun [1 ,2 ,3 ]
Luo, Minnan [1 ,2 ,3 ]
Li, Zhihui [4 ]
Wang, Yang [5 ]
机构
[1] Xi An Jiao Tong Univ, Sch Comp Sci & Technol, Xian 710049, Peoples R China
[2] Xi An Jiao Tong Univ, Minist Educ, Key Lab Intelligent Networks & Network Secur, Xian 710049, Peoples R China
[3] Xi An Jiao Tong Univ, Shaanxi Prov Key Lab Big Data Knowledge Engn, Xian 710049, Peoples R China
[4] Univ Sci & Technol China, Sch Informat Sci & Technol, Hefei 230026, Peoples R China
[5] Xi An Jiao Tong Univ, Sch Continuing Educ, Xian 710049, Peoples R China
基金
中国国家自然科学基金;
关键词
Video transformer; Temporal action localization; High efficiency; NETWORK;
D O I
10.1016/j.neucom.2024.128688
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Temporal action localization is a fundamental task in video understanding. Existing methods fall into three categories: anchor-based, actionness-guided, and anchor-free. Anchor-based and actionness-guided models need huge computation resources to process redundant proposals or enumerate every possible proposal. Anchor- free models with lighter parameters become a more attractive option as temporal actions become more complex. However, they typically struggle to achieve high performance due to the need to aggregate global temporal-spatial features at every time step. To overcome this limitation, we design three efficient transformer- based architectures, bringing two advantages: (i) the global receptive field of transformers enables models to aggregate spatial and temporal at each time step, and (ii) the transformers could capture the moment-level feature, enhancing localization performance. Our designed architectures are adapted to any framework, thus we propose a simple but effective anchor-free framework named TEST. Compared to strong baselines, TEST achieves 0.96% to 3.20% improvement on two real-world datasets. Meanwhile, it improves time efficiency by 1.36 times and space efficiency by 1.08 times. Further experiments prove the effectiveness of TEST's modules. Implementation of our work is available at https://github.com/whr000001/TeST.
引用
收藏
页数:9
相关论文
共 50 条
  • [41] Temporal-spatial and pedagogical flexibility in distance education
    Hrastinski, Stefan
    Paul, Enni
    Akerfeldt, Anna
    DISTANCE EDUCATION, 2024,
  • [42] Analysis on Temporal-Spatial Variations of Australian TEC
    Ouyang, Gary
    Wang, Jian
    Wang, Jinling
    Cole, David
    OBSERVING OUR CHANGING EARTH, 2009, 133 : 751 - +
  • [43] Route recommendation based on temporal-spatial metric
    Liang, Feng
    Chen, Honglong
    Lin, Kai
    Li, Junjian
    Li, Zhe
    Xue, Huansheng
    Shakhov, Vladimir
    Bin Liaqat, Hannan
    COMPUTERS & ELECTRICAL ENGINEERING, 2022, 97
  • [44] HoMS: A Breakthrough Development with Temporal-Spatial Ordering
    Lu, G. Q.
    CHEM, 2020, 6 (06): : 1215 - 1216
  • [45] TEMPORAL-SPATIAL DISTRIBUTION OF LEUKEMIA AND LYMPHOMA IN CONNECTICUT
    EDERER, F
    MYERS, MH
    EISENBERG, H
    CAMPBELL, PC
    JNCI-JOURNAL OF THE NATIONAL CANCER INSTITUTE, 1965, 35 (04): : 625 - +
  • [46] TEMPORAL-SPATIAL DISORIENTATION - .2. TIME
    MARCHAIS, P
    ANNALES MEDICO-PSYCHOLOGIQUES, 1977, 135 (05): : 898 - 907
  • [47] ON A TEMPORAL-SPATIAL RADIATION FUNCTIONAL AND ITS MEASUREMENT
    GASE, R
    PONATH, HE
    SCHUBERT, M
    ANNALEN DER PHYSIK, 1986, 43 (6-8) : 487 - 498
  • [48] Nonlinear temporal-spatial surface plasmon polaritons
    Ablowitz, M. J.
    Butler, S. T. J.
    OPTICS COMMUNICATIONS, 2014, 330 : 49 - 55
  • [49] Modeling Temporal-Spatial Correlations for Crime Prediction
    Zhao, Xiangyu
    Tang, Jiliang
    CIKM'17: PROCEEDINGS OF THE 2017 ACM CONFERENCE ON INFORMATION AND KNOWLEDGE MANAGEMENT, 2017, : 497 - 506
  • [50] Frequency controlled terahertz temporal-spatial modulator
    Zhang, Yan
    Wang, Guocui
    2022 IEEE MTT-S INTERNATIONAL MICROWAVE WORKSHOP SERIES ON ADVANCED MATERIALS AND PROCESSES FOR RF AND THZ APPLICATIONS, IMWS-AMP, 2022,