Hybrid explainable image caption generation using image processing and natural language processing

被引:0
|
作者
Mishra, Atul [1 ]
Agrawal, Anubhav [1 ]
Bhasker, Shailendra [2 ]
机构
[1] BML Munjal Univ, Gurgaon, India
[2] Harcourt Butler Tech Univ, Kanpur, India
关键词
NLP; Image caption generation; CNN; LSTM; InceptionV3; HYPERPARAMETER OPTIMIZATION;
D O I
10.1007/s13198-024-02495-5
中图分类号
T [工业技术];
学科分类号
08 ;
摘要
Image caption generation is among the most rapidly growing research areas that combine image processing methodologies with natural language processing (NLP) technique(s). The effectiveness of the combination of image processing and NLP techniques can revolutionaries the areas of content creation, media analysis, and accessibility. The study proposed a novel model to generate automatic image captions by consuming visual and linguistic features. Visual image features are extracted by applying Convolutional Neural Network and linguistic features by Long Short-Term Memory (LSTM) to generate text. Microsoft Common Objects in Context dataset with over 330,000 images having corresponding captions is used to train the proposed model. A comprehensive evaluation of various models, including VGGNet + LSTM, ResNet + LSTM, GoogleNet + LSTM, VGGNet + RNN, AlexNet + RNN, and AlexNet + LSTM, was conducted based on different batch sizes and learning rates. The assessment was performed using metrics such as BLEU-2 Score, METEOR Score, ROUGE-L Score, and CIDEr. The proposed method demonstrated competitive performance, suggesting its potential for further exploration and refinement. These findings underscore the importance of careful parameter tuning and model selection in image captioning tasks.
引用
收藏
页码:4874 / 4884
页数:11
相关论文
共 50 条
  • [41] Objective Type Question Generation using Natural Language Processing
    Deena, G.
    Raja, K.
    INTERNATIONAL JOURNAL OF ADVANCED COMPUTER SCIENCE AND APPLICATIONS, 2022, 13 (02) : 539 - 548
  • [42] Underwater Image Enhancement Using Image Processing
    Nagamma, V.
    Halse, S. V.
    THIRD INTERNATIONAL CONFERENCE ON IMAGE PROCESSING AND CAPSULE NETWORKS (ICIPCN 2022), 2022, 514 : 13 - 22
  • [43] Image Caption Generation With Adaptive Transformer
    Zhang, Wei
    Nie, Wenbo
    Li, Xinle
    Yu, Yao
    2019 34RD YOUTH ACADEMIC ANNUAL CONFERENCE OF CHINESE ASSOCIATION OF AUTOMATION (YAC), 2019, : 521 - 526
  • [44] The Accurate Guidance for Image Caption Generation
    Qi, Xinyuan
    Cao, Zhiguo
    Xiao, Yang
    Wang, Jian
    Zhang, Chao
    PATTERN RECOGNITION AND COMPUTER VISION, PT III, 2018, 11258 : 15 - 26
  • [45] An Overview of Image Caption Generation Methods
    Wang, Haoran
    Zhang, Yue
    Yu, Xiaosheng
    COMPUTATIONAL INTELLIGENCE AND NEUROSCIENCE, 2020, 2020
  • [46] A survey on automatic image caption generation
    Bai, Shuang
    An, Shan
    NEUROCOMPUTING, 2018, 311 : 291 - 304
  • [47] Image caption generation using transformer learning methods: a case study on instagram image
    Dittakan, Kwankamon
    Prompitak, Kamontorn
    Thungklang, Phutphisit
    Wongwattanakit, Chatchawan
    MULTIMEDIA TOOLS AND APPLICATIONS, 2023, 83 (15) : 46397 - 46417
  • [48] Combat COVID-19 infodemic using explainable natural language processing models
    Ayoub, Jackie
    Yang, X. Jessie
    Zhou, Feng
    INFORMATION PROCESSING & MANAGEMENT, 2021, 58 (04)
  • [49] Image caption generation using transformer learning methods: a case study on instagram image
    Kwankamon Dittakan
    Kamontorn Prompitak
    Phutphisit Thungklang
    Chatchawan Wongwattanakit
    Multimedia Tools and Applications, 2024, 83 : 46397 - 46417
  • [50] Hybrid perceptual image processing using new interpolating wavelets
    Shi, Z
    Wang, HX
    Zhang, D
    Kouri, DJ
    Hoffman, DK
    2000 INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, VOL III, PROCEEDINGS, 2000, : 817 - 820