Conference Articles

Scene Text Recognition with Image-Text Matching-Guided Dictionary

Jiajun Wei, East China Normal University
Hongjian Zhan, East China Normal University
Xiao Tu, East China Normal University
Yue Lu, East China Normal University
Umapada Pal, Indian Statistical Institute, Kolkata

Document Type

Conference Article

Publication Title

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)

Abstract

Employing a dictionary can efficiently rectify the deviation between the visual prediction and the ground truth in scene text recognition methods. However, the independence of the dictionary on the visual features may lead to incorrect rectification of accurate visual predictions. In this paper, we propose a new dictionary language model leveraging the Scene Image-Text Matching(SITM) network, which avoids the drawbacks of the explicit dictionary language model: 1) the independence of the visual features; 2) noisy choice in candidates etc. The SITM network accomplishes this by using Image-Text Contrastive (ITC) Learning to match an image with its corresponding text among candidates in the inference stage. ITC is widely used in vision-language learning to pull the positive image-text pair closer in feature space. Inspired by ITC, the SITM network combines the visual features and the text features of all candidates to identify the candidate with the minimum distance in the feature space. Our lexicon method achieves better results(93.8% accuracy) than the ordinary method results(92.1% accuracy) on six mainstream benchmarks. Additionally, we integrate our method with ABINet and establish new state-of-the-art results on several benchmarks.

First Page

Last Page

DOI

10.1007/978-3-031-41731-3_4

Publication Date

1-1-2023

Comments

Open Access, Green

Recommended Citation

Wei, Jiajun; Zhan, Hongjian; Tu, Xiao; Lu, Yue; and Pal, Umapada, "Scene Text Recognition with Image-Text Matching-Guided Dictionary" (2023). Conference Articles. 566.
https://digitalcommons.isical.ac.in/conf-articles/566

This document is currently not available here.

COinS

Conference Articles

Scene Text Recognition with Image-Text Matching-Guided Dictionary

Document Type

Publication Title

Abstract

First Page

Last Page

DOI

Publication Date

Comments

Recommended Citation

Browse

Search

Author Corner

Links

Conference Articles

Scene Text Recognition with Image-Text Matching-Guided Dictionary

Authors

Document Type

Publication Title

Abstract

First Page

Last Page

DOI

Publication Date

Comments

Recommended Citation

Share

Browse

Search

Author Corner

Links