Transforming Scene Text Detection and Recognition: A Multi-Scale End-to-End Approach With Transformer Framework

被引:0
|
作者
Geng, Tianyu [1 ]
机构
[1] Nanjing Tech Univ, Coll Artificial Intelligence, Coll Comp & Informat Engn, Nanjing 211816, Jiangsu, Peoples R China
关键词
Text recognition; text recognition; transformer; end-to-end; multi-scale;
D O I
10.1109/ACCESS.2024.3375497
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Text is an essential means for humans to acquire information and engage in social communication. Accurate text extraction from images is crucial for various tasks in real-life scenarios and scene understanding. However, text detection and recognition in natural scenes are challenged by noise in the images, irregular distribution of text fonts, and degradation of image quality under complex acquisition conditions. These factors severely impact the accuracy of text recognition. Issues such as poor image quality, diverse text formats, and complex image backgrounds significantly affect the accuracy of the recognition, and these challenges remain urgent to be addressed in the field. To address these challenges, this paper proposes a transformer-based scene image text detection and recognition algorithm within a multi-scale end-to-end framework. Firstly, by integrating detection and recognition stages into an end-to-end framework, the process is simplified, reducing computation and errors. Subsequently, multi-scale characteristics are incorporated to effectively capture text information at various scales, enhancing recognition accuracy and robustness through feature fusion and anti-interference capability. Lastly, leveraging the transformer framework, the algorithm efficiently handles text information of different scales and positions, improving generalization ability. The self-attention mechanism, multi-layer stacking structure, and positional encoding in the transformer framework contribute to its effectiveness in processing diverse text information. Through validation, the proposed method demonstrates improved efficiency in scene text detection and recognition.
引用
收藏
页码:40582 / 40596
页数:15
相关论文
共 50 条
  • [41] JOINT VERIFICATION-IDENTIFICATION IN END-TO-END MULTI-SCALE CNN FRAMEWORK FOR TOPIC IDENTIFICATION
    Pappagari, Raghavendra
    Villalba, Jesus
    Dehak, Najim
    2018 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), 2018, : 6199 - 6203
  • [42] A Multi-level Acoustic Feature Extraction Framework for Transformer Based End-to-End Speech Recognition
    Li, Jin
    Su, Rongfeng
    Xie, Xurong
    Yan, Nan
    Wang, Lan
    INTERSPEECH 2022, 2022, : 3173 - 3177
  • [43] Soft set-based MSER end-to-end system for occluded scene text detection, recognition and prediction
    Das, Alloy
    Palaiahnakote, Shivakumara
    Banerjee, Ayan
    Antonacopoulos, Apostolos
    Pal, Umapada
    KNOWLEDGE-BASED SYSTEMS, 2024, 305
  • [44] DFDT: An End-to-End DeepFake Detection Framework Using Vision Transformer
    Khormali, Aminollah
    Yuan, Jiann-Shiun
    APPLIED SCIENCES-BASEL, 2022, 12 (06):
  • [45] Cursive-Text: A Comprehensive Dataset for End-to-End Urdu Text Recognition in Natural Scene Images
    Chandio, Asghar Ali
    Asikuzzamana, Md.
    Pickering, Mark
    Leghari, Mehwish
    DATA IN BRIEF, 2020, 31
  • [46] An End-to-End Scene Text Detector with Dynamic Attention
    Lin, Jingyu
    Yan, Yan
    Wang, Hanzi
    PROCEEDINGS OF THE 4TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA IN ASIA, MMASIA 2022, 2022,
  • [47] A Multi-Level Optimization Framework for End-to-End Text Augmentation
    Somayajula, Sai Ashish
    Song, Linfeng
    Xie, Pengtao
    TRANSACTIONS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, 2022, 10 : 343 - 358
  • [48] Variable Scale Pruning for Transformer Model Compression in End-to-End Speech Recognition
    Ben Letaifa, Leila
    Rouas, Jean-Luc
    ALGORITHMS, 2023, 16 (09)
  • [49] Multi-Scale End-to-End Speaker Recognition System Based on Improved Res2Net
    Deng, Lihong
    Deng, Fei
    Zhang, Gexiang
    Yang, Qiang
    Computer Engineering and Applications, 2023, 59 (24) : 110 - 120
  • [50] MULTI-SCALE END-TO-END LEARNING FOR POINT CLOUD GEOMETRY COMPRESSION
    Xu, Yiqun
    Yin, Qian
    Wang, Shanshe
    Zhang, Xinfeng
    Ma, Siwei
    Gao, Wen
    2022 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, ICIP, 2022, : 2107 - 2111