Transforming Scene Text Detection and Recognition: A Multi-Scale End-to-End Approach With Transformer Framework

被引:0
|
作者
Geng, Tianyu [1 ]
机构
[1] Nanjing Tech Univ, Coll Artificial Intelligence, Coll Comp & Informat Engn, Nanjing 211816, Jiangsu, Peoples R China
关键词
Text recognition; text recognition; transformer; end-to-end; multi-scale;
D O I
10.1109/ACCESS.2024.3375497
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Text is an essential means for humans to acquire information and engage in social communication. Accurate text extraction from images is crucial for various tasks in real-life scenarios and scene understanding. However, text detection and recognition in natural scenes are challenged by noise in the images, irregular distribution of text fonts, and degradation of image quality under complex acquisition conditions. These factors severely impact the accuracy of text recognition. Issues such as poor image quality, diverse text formats, and complex image backgrounds significantly affect the accuracy of the recognition, and these challenges remain urgent to be addressed in the field. To address these challenges, this paper proposes a transformer-based scene image text detection and recognition algorithm within a multi-scale end-to-end framework. Firstly, by integrating detection and recognition stages into an end-to-end framework, the process is simplified, reducing computation and errors. Subsequently, multi-scale characteristics are incorporated to effectively capture text information at various scales, enhancing recognition accuracy and robustness through feature fusion and anti-interference capability. Lastly, leveraging the transformer framework, the algorithm efficiently handles text information of different scales and positions, improving generalization ability. The self-attention mechanism, multi-layer stacking structure, and positional encoding in the transformer framework contribute to its effectiveness in processing diverse text information. Through validation, the proposed method demonstrates improved efficiency in scene text detection and recognition.
引用
收藏
页码:40582 / 40596
页数:15
相关论文
共 50 条
  • [21] An adaptive n-gram transformer for multi-scale scene text recognition
    Yan, Xueming
    Fang, Zhihang
    Jin, Yaochu
    KNOWLEDGE-BASED SYSTEMS, 2023, 280
  • [22] Speech-and-Text Transformer: Exploiting Unpaired Text for End-to-End Speech Recognition
    Wang, Qinyi
    Zhou, Xinyuan
    Li, Haizhou
    APSIPA TRANSACTIONS ON SIGNAL AND INFORMATION PROCESSING, 2023, 12 (01)
  • [23] END-TO-END MULTI-CHANNEL TRANSFORMER FOR SPEECH RECOGNITION
    Chang, Feng-Ju
    Radfar, Martin
    Mouchtaris, Athanasios
    King, Brian
    Kunzmann, Siegfried
    2021 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP 2021), 2021, : 5884 - 5888
  • [24] END-TO-END MULTI-SPEAKER SPEECH RECOGNITION WITH TRANSFORMER
    Chang, Xuankai
    Zhang, Wangyou
    Qian, Yanmin
    Le Roux, Jonathan
    Watanabe, Shinji
    2020 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, 2020, : 6134 - 6138
  • [25] MAQT: multi-scale attention and query-optimized transformer for end-to-end pose estimation
    Liang, Hong
    Wang, Cuiping
    Shao, Mingwen
    Zhang, Qian
    JOURNAL OF SUPERCOMPUTING, 2025, 81 (02):
  • [26] Towards End-to-End Unified Scene Text Detection and Layout Analysis
    Long, Shangbang
    Qin, Siyang
    Panteleev, Dmitry
    Bissacco, Alessandro
    Fujii, Yasuhisa
    Raptis, Michalis
    2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, : 1039 - 1049
  • [27] An end-to-end multi-scale airway segmentation framework based on pulmonary CT image
    Yuan, Ye
    Tan, Wenjun
    Xu, Lisheng
    Bao, Nan
    Zhu, Quan
    Wang, Zhe
    Wang, Ruoyu
    PHYSICS IN MEDICINE AND BIOLOGY, 2024, 69 (11):
  • [28] Feature Fusion Pyramid Network for End-to-End Scene Text Detection
    Wu, Yirui
    Zhang, Lilai
    Li, Hao
    Zhang, Yunfei
    Wan, Shaohua
    ACM TRANSACTIONS ON ASIAN AND LOW-RESOURCE LANGUAGE INFORMATION PROCESSING, 2024, 23 (11)
  • [29] Scene text spotting based on end-to-end
    Wei G.
    Rong W.
    Liang Y.
    Xiao X.
    Liu X.
    Journal of Intelligent and Fuzzy Systems, 2021, 40 (05): : 8871 - 8881
  • [30] Visual place recognition from end-to-end semantic scene text features
    Raisi, Zobeir
    Zelek, John
    FRONTIERS IN ROBOTICS AND AI, 2024, 11