Convolutional recurrent neural networks with hidden Markov model bootstrap for scene text recognition

被引:12
|
作者
Wang, Fenglei [1 ]
Guo, Qiang [1 ]
Lei, Jun [1 ]
Zhang, Jun [1 ]
机构
[1] Natl Univ Def Technol, Dept Informat Syst & Management, Changsha, Hunan, Peoples R China
基金
中国国家自然科学基金;
关键词
recurrent neural nets; text detection; hidden Markov models; convolutional recurrent neural networks; scene text recognition; RNN; CNN; Gaussian mixture model-hidden Markov model; lexicon-free text; lexicon-based text; HANDWRITING RECOGNITION;
D O I
10.1049/iet-cvi.2016.0417
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Text recognition in natural scene remains a challenging problem due to the highly variable appearance in unconstrained condition. The authors develop a system that directly transcribes scene text images to text without character segmentation. They formulate the problem as sequence labelling. They build a convolutional recurrent neural network (RNN) by using deep convolutional neural networks (CNN) for modelling text appearance and RNNs for sequence dynamics. The two models are complementary in modelling capabilities and so integrated together to form the segmentation free system. They train a Gaussian mixture model-hidden Markov model to supervise the training of the CNN model. The system is data driven and needs no hand labelled training data. Their method has several appealing properties: (i) It can recognise arbitrary length text images. (ii) The recognition process does not involve sophisticated character segmentation. (iii) It is trained on scene text images with only word-level transcriptions. (iv) It can recognise both the lexicon-based or lexicon-free text. The proposed system achieves competitive performance comparison with the state of the art on several public scene text datasets, including both lexicon-based and non-lexicon ones.
引用
收藏
页码:497 / 504
页数:8
相关论文
共 50 条
  • [1] An Attention-Based Convolutional Recurrent Neural Networks for Scene Text Recognition
    Alshawi, Adil Abdullah Abdulhussein
    Tanha, Jafar
    Balafar, Mohammad Ali
    IEEE ACCESS, 2024, 12 : 8123 - 8134
  • [2] SCENE TEXT RECOGNITION WITH DEEPER CONVOLUTIONAL NEURAL NETWORKS
    Zhang, Yuqi
    Wang, Wei
    Wang, Liang
    Wang, Liuan
    2015 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2015, : 2384 - 2388
  • [3] Scene Text Script Identification with Convolutional Recurrent Neural Networks
    Mei, Jieru
    Dai, Luo
    Shi, Baoguang
    Bai, Xiang
    2016 23RD INTERNATIONAL CONFERENCE ON PATTERN RECOGNITION (ICPR), 2016, : 4053 - 4058
  • [4] Scene text recognition using residual convolutional recurrent neural network
    Lei, Zhengchao
    Zhao, Sanyuan
    Song, Hongmei
    Shen, Jianbing
    MACHINE VISION AND APPLICATIONS, 2018, 29 (05) : 861 - 871
  • [5] Scene text recognition using residual convolutional recurrent neural network
    Zhengchao Lei
    Sanyuan Zhao
    Hongmei Song
    Jianbing Shen
    Machine Vision and Applications, 2018, 29 : 861 - 871
  • [6] Deep Convolutional Neural Network Based Hidden Markov Model for Offline Handwritten Chinese Text Recognition
    Wang, Zi-Rui
    Du, Jun
    Hu, Jin-Shui
    Hu, Yu-Long
    PROCEEDINGS 2017 4TH IAPR ASIAN CONFERENCE ON PATTERN RECOGNITION (ACPR), 2017, : 816 - 821
  • [7] Convolutional Attention Networks for Scene Text Recognition
    Xie, Hongtao
    Fang, Shancheng
    Zha, Zheng-Jun
    Yang, Yating
    Li, Yan
    Zhang, Yongdong
    ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS, 2019, 15 (01)
  • [8] Facial Recognition Using Hidden Markov Model and Convolutional Neural Network
    Bilal, Muhammad
    Razzaq, Saqlain
    Bhowmike, Nirman
    Farooq, Azib
    Zahid, Muhammad
    Shoaib, Sultan
    AI, 2024, 5 (03) : 1633 - 1647
  • [9] Recurrent Convolutional Neural Networks for Scene Labeling
    Pinheiro, Pedro O.
    Collobert, Ronan
    INTERNATIONAL CONFERENCE ON MACHINE LEARNING, VOL 32 (CYCLE 1), 2014, 32
  • [10] Automatic Speech Recognition: Comparisons Between Convolutional Neural Networks, Hidden Markov Model and Hybrid Architecture
    Santos, Lyndaines
    Moreira, Nicolas de Araujo
    Sampaio, Robson
    Lima, Raizielle
    Oliveira, Francisco Carlos Mattos Brito
    EXPERT SYSTEMS, 2025, 42 (05)