Scene-Text Oriented Referring Expression Comprehension

被引:1
|
作者
Bu, Yuqi [1 ,2 ]
Li, Liuwu [1 ,2 ]
Xie, Jiayuan [1 ,2 ]
Liu, Qiong [1 ,2 ]
Cai, Yi [1 ,2 ]
Huang, Qingbao [3 ,4 ]
Li, Qing [5 ]
机构
[1] South China Univ Technol, Sch Software Engn, Guangzhou 510006, Peoples R China
[2] Key Lab Big Data & Intelligent Robot SCUT, MOE China, Guangzhou 510006, Peoples R China
[3] Guangxi Univ, Inst Artificial Intelligence, Sch Elect Engn, Nanning 530004, Guangxi, Peoples R China
[4] Guangxi Key Lab Intelligent Control & Maintenance, Nanning 530004, Peoples R China
[5] Hong Kong Polytech Univ, Dept Comp, Hong Kong 999077, Peoples R China
基金
中国国家自然科学基金;
关键词
Visualization; Task analysis; Semantics; Text recognition; Pain; Natural languages; Image recognition; Referring expression comprehension; scene text representation; multimodal alignment; LOCALIZATION;
D O I
10.1109/TMM.2022.3219642
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Referring expression comprehension (REC) aims to identify and locate a specific object in visual scenes referred to by a natural language expression. Existing studies of REC only focus on basic visual attributes and neglect scene text. Since scene text has the functions of object identification and disambiguation, it is naturally and frequently used to refer to objects. However, existing methods do not explicitly recognize text in images and fail to align scene text mentioned in expressions with the text shown in images, resulting in object localization errors. This article takes the first step toward addressing these limitations. First, we introduce a new task called scene-text oriented referring expression comprehension, which aims to align visual cues and textual semantics of scene text with referring expressions and visual contents. Second, we propose a scene text awareness network that can bridge the gap between texts from two modalities by grounding visual representations of expression-correlated scene texts. Specifically, we propose a correlated text extraction module to solve the problem of lacking semantic understanding, and a correlated region activation module to address the fixed alignment problem and absent alignment problem. These modules ensure that the proposed method focuses on local regions that are most relevant to scene text, thus mitigating the misalignment of scene text with irrelevant regions. Third, to conduct quantitative evaluations, we establish a new benchmark dataset called RefText. Experimental results demonstrate that the proposed method can effectively comprehend scene-text oriented referring expressions and achieves excellent performance.
引用
收藏
页码:7208 / 7221
页数:14
相关论文
共 50 条
  • [1] Knowledge Mining of Scene Text for Referring Expression Comprehension
    Gao, Chenyang
    Yang, Biao
    Yu, Wenwen
    Liu, Yuliang
    Bai, Xiang
    DOCUMENT ANALYSIS AND RECOGNITION-ICDAR 2024, PT V, 2024, 14808 : 245 - 262
  • [2] Unambiguous Scene Text Segmentation With Referring Expression Comprehension
    Rong, Xuejian
    Yi, Chucai
    Tian, Yingli
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2020, 29 (29) : 591 - 601
  • [3] Bridging the Gap between Expression and Scene Text for Referring Expression Comprehension (Student Abstract)
    Bu, Yuqi
    Xie, Jiayuan
    Li, Liuwu
    Liu, Qiong
    Cai, Yi
    THIRTY-SIXTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE / THIRTY-FOURTH CONFERENCE ON INNOVATIVE APPLICATIONS OF ARTIFICIAL INTELLIGENCE / TWELVETH SYMPOSIUM ON EDUCATIONAL ADVANCES IN ARTIFICIAL INTELLIGENCE, 2022, : 12921 - 12922
  • [4] Scene-text Oriented Visual Entailment: Task, Dataset and Solution
    Li, Nan
    Li, Pijian
    Xu, Dongsheng
    Zhao, Wenye
    Cai, Yi
    Huang, Qingbao
    PROCEEDINGS OF THE 31ST ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2023, 2023, : 5562 - 5571
  • [5] Review of Text Extraction Algorithms for Scene-text and Document Images
    Sahare, Parul
    Dhok, Sanjay B.
    IETE TECHNICAL REVIEW, 2017, 34 (02) : 144 - 164
  • [6] Scene Graph Enhanced Pseudo-Labeling for Referring Expression Comprehension
    Wu, Cantao
    Cai, Yi
    Li, Liuwu
    Wang, Jiexin
    FINDINGS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (EMNLP 2023), 2023, : 11978 - 11990
  • [7] GLASS: Global to Local Attention for Scene-Text Spotting
    Ronen, Roi
    Tsiper, Shahar
    Anschel, Oron
    Lavi, Inbal
    Markovitz, Amir
    Manmatha, R.
    COMPUTER VISION - ECCV 2022, PT XXVIII, 2022, 13688 : 249 - 266
  • [8] Adaptive scene-text binarisation on images captured by smartphones
    Belhedi, Amira
    Marcotegui, Beatriz
    IET IMAGE PROCESSING, 2016, 10 (07) : 515 - 523
  • [9] PreSTU: Pre-Training for Scene-Text Understanding
    Kil, Jihyung
    Changpinyo, Soravit
    Chen, Xi
    Hu, Hexiang
    Goodman, Sebastian
    Chao, Wei-Lun
    Soricut, Radu
    2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, : 15224 - 15234
  • [10] HyText - A Scene-Text Extraction Method for Video Retrieval
    Theus, Alexander
    Rossetto, Luca
    Bernstein, Abraham
    MULTIMEDIA MODELING, MMM 2022, PT II, 2022, 13142 : 182 - 193