FRAMEWORK FOR EVALUATION OF SOUND EVENT DETECTION IN WEB VIDEOS

被引:0
|
作者
Badlani, Rohan [2 ]
Shah, Ankit [1 ]
Elizalde, Benjamin [1 ]
Kumar, Anurag [1 ]
Raj, Bhiksha [1 ]
机构
[1] Carnegie Mellon Univ, Language Technol Inst, Pittsburgh, PA 15213 USA
[2] BITS Pilani, Dept Comp Sci, Hyderabad, Telangana, India
关键词
Sound Event Detection; Convolutional Neural Network; Large-Scale audio event detection; Video Content Analysis;
D O I
暂无
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
The largest source of sound events is web videos. Most videos lack sound event labels at segment level, however, a significant number of them do respond to text queries, from a match found using metadata by search engines. In this paper we explore the extent to which a search query can be used as the true label for detection of sound events in videos. We present a framework for large-scale sound event recognition on web videos. The framework crawls videos using search queries corresponding to 78 sound event labels drawn from three datasets. The datasets are used to train three classifiers, and we obtain a prediction on 3.7 million web video segments. We evaluated performance using the search query as true label and compare it with human labeling. Both types of ground truth exhibited close performance, to within 10%, and similar performance trend with increasing number of evaluated segments. Hence, our experiments show potential for using search query as a preliminary true label for sound event recognition in web videos.
引用
收藏
页码:3096 / 3100
页数:5
相关论文
共 50 条
  • [1] A FRAMEWORK FOR THE ROBUST EVALUATION OF SOUND EVENT DETECTION
    Bilen, Cagdas
    Ferroni, Giacomo
    Tuveri, Francesco
    Azcarreta, Juan
    Krstulovic, Sacha
    2020 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, 2020, : 61 - 65
  • [2] Multimodal Feature Fusion for Robust Event Detection in Web Videos
    Natarajan, Pradeep
    Wu, Shuang
    Vitaladevuni, Shiv
    Zhuang, Xiaodan
    Tsakalidis, Stavros
    Park, Unsang
    Prasad, Rohit
    Natarajan, Premkumar
    2012 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2012, : 1298 - 1305
  • [3] How Unlabeled Web Videos Help Complex Event Detection?
    Liu, Huan
    Zheng, Qinghua
    Luo, Minnan
    Zhang, Dingwen
    Chang, Xiaojun
    Deng, Cheng
    PROCEEDINGS OF THE TWENTY-SIXTH INTERNATIONAL JOINT CONFERENCE ON ARTIFICIAL INTELLIGENCE, 2017, : 4040 - 4046
  • [4] Hierarchic ConvNets Framework for Rare Sound Event Detection
    Vesperini, Fabio
    Droghini, Diego
    Principi, Emanuele
    Gabrielli, Leonardo
    Squartini, Stefano
    2018 26TH EUROPEAN SIGNAL PROCESSING CONFERENCE (EUSIPCO), 2018, : 1497 - 1501
  • [5] How Related Exemplars Help Complex Event Detection in Web Videos?
    Yang, Yi
    Ma, Zhigang
    Xu, Zhongwen
    Yan, Shuicheng
    Hauptmann, Alexander G.
    2013 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2013, : 2104 - 2111
  • [6] VideoSkip: event detection in social web videos with an implicit user heuristic
    Chrysoula Gkonela
    Konstantinos Chorianopoulos
    Multimedia Tools and Applications, 2014, 69 : 383 - 396
  • [7] VideoSkip: event detection in social web videos with an implicit user heuristic
    Gkonela, Chrysoula
    Chorianopoulos, Konstantinos
    MULTIMEDIA TOOLS AND APPLICATIONS, 2014, 69 (02) : 383 - 396
  • [8] MULTIMODAL EVALUATION METHOD FOR SOUND EVENT DETECTION
    Modaresi, Seyed M. R.
    Osmani, Aomar
    Razzazi, Mohammadreza
    Chibani, Abdelghani
    2022 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), 2022, : 1026 - 1030
  • [9] Single Modality-Based Event Detection Framework for Complex Videos
    Arif, Sheeraz
    Siddiqui, Adnan Ahmed
    Kumar, Rajesh
    Maheshwari, Avinash
    Maheshwari, Komal
    Saeed, Muhammad Imran
    INTERNATIONAL JOURNAL OF ADVANCED COMPUTER SCIENCE AND APPLICATIONS, 2020, 11 (11) : 86 - 93
  • [10] Neural network based framework for goal event detection in soccer videos
    Wickramaratna, K
    Chen, M
    Chen, SC
    Shyu, ML
    ISM 2005: SEVENTH IEEE INTERNATIONAL SYMPOSIUM ON MULTIMEDIA, PROCEEDINGS, 2005, : 21 - 28