MOoSE: Multi-Orientation Sharing Experts for Open-set Scene Text Recognition

Liu, Chang; Corbillé, Simon; Smith, Elisa H Barney

Computer Science > Computer Vision and Pattern Recognition

arXiv:2407.18616 (cs)

[Submitted on 26 Jul 2024]

Title:MOoSE: Multi-Orientation Sharing Experts for Open-set Scene Text Recognition

Authors:Chang Liu, Simon Corbillé, Elisa H Barney Smith

View PDF HTML (experimental)

Abstract:Open-set text recognition, which aims to address both novel characters and previously seen ones, is one of the rising subtopics in the text recognition field. However, the current open-set text recognition solutions only focuses on horizontal text, which fail to model the real-life challenges posed by the variety of writing directions in real-world scene text. Multi-orientation text recognition, in general, faces challenges from the diverse image aspect ratios, significant imbalance in data amount, and domain gaps between orientations. In this work, we first propose a Multi-Oriented Open-Set Text Recognition task (MOOSTR) to model the challenges of both novel characters and writing direction variety. We then propose a Multi-Orientation Sharing Experts (MOoSE) framework as a strong baseline solution. MOoSE uses a mixture-of-experts scheme to alleviate the domain gaps between orientations, while exploiting common structural knowledge among experts to alleviate the data scarcity that some experts face. The proposed MOoSE framework is validated by ablative experiments, and also tested for feasibility on the existing open-set benchmark. Code, models, and documents are available at: this https URL

Comments:	Accepted in ICDAR2024
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2407.18616 [cs.CV]
	(or arXiv:2407.18616v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2407.18616

Submission history

From: Chang Liu [view email]
[v1] Fri, 26 Jul 2024 09:20:29 UTC (4,254 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:MOoSE: Multi-Orientation Sharing Experts for Open-set Scene Text Recognition

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:MOoSE: Multi-Orientation Sharing Experts for Open-set Scene Text Recognition

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators