Optical Character Recognition of 19th Century Classical Commentaries: the Current State of Affairs

Romanello, Matteo; Najem-Meyer, Sven; Robertson, Bruce

Computer Science > Digital Libraries

arXiv:2110.06817 (cs)

[Submitted on 13 Oct 2021]

Title:Optical Character Recognition of 19th Century Classical Commentaries: the Current State of Affairs

Authors:Matteo Romanello, Sven Najem-Meyer, Bruce Robertson

View PDF

Abstract:Together with critical editions and translations, commentaries are one of the main genres of publication in literary and textual scholarship, and have a century-long tradition. Yet, the exploitation of thousands of digitized historical commentaries was hitherto hindered by the poor quality of Optical Character Recognition (OCR), especially on commentaries to Greek texts. In this paper, we evaluate the performances of two pipelines suitable for the OCR of historical classical commentaries. Our results show that Kraken + Ciaconna reaches a substantially lower character error rate (CER) than Tesseract/OCR-D on commentary sections with high density of polytonic Greek text (average CER 7% vs. 13%), while Tesseract/OCR-D is slightly more accurate than Kraken + Ciaconna on text sections written predominantly in Latin script (average CER 8.2% vs. 8.4%). As part of this paper, we also release GT4HistComment, a small dataset with OCR ground truth for 19th classical commentaries and Pogretra, a large collection of training data and pre-trained models for a wide variety of ancient Greek typefaces.

Subjects:	Digital Libraries (cs.DL); Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2110.06817 [cs.DL]
	(or arXiv:2110.06817v1 [cs.DL] for this version)
	https://doi.org/10.48550/arXiv.2110.06817

Submission history

From: Matteo Romanello [view email]
[v1] Wed, 13 Oct 2021 16:01:16 UTC (1,644 KB)

Computer Science > Digital Libraries

Title:Optical Character Recognition of 19th Century Classical Commentaries: the Current State of Affairs

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Digital Libraries

Title:Optical Character Recognition of 19th Century Classical Commentaries: the Current State of Affairs

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators