Vol. 14 | No. 26-27, 2026


CHALLENGES IN DIGITIZATION AND OCR FOR METADATA EXTRACTION FROM THE JOURNAL ARCHIVE FOR ALBANIAN ANTIQUITY, LANGUAGE AND ETHNOLOGY

Aleksandra ANDRIĆ, Marija HRKALOVIĆ

Abstract

The paper presents the extraction, digitization and metadata analysis of the journal Archive for Arbanas Antiquity, Language and Ethnology , which represents a significant source for the study of Albanian cultural, linguistic, and ethnological heritage. The original issues of the journal from 1923-1926 were used as an initial resource. Python tools and libraries were used for text extraction, including PyMuPDF and Tesseract OCR, with appropriate language models for multiple languages and scripts. Metadata extraction received special care, covering title, author, year of publication, section, volume and pagination. The automatically obtained metadata was then compared with a manually created metadata database, where the similarity of titles by books was analyzed. The results showed that the OCR procedure can significantly speed up the processing of materials, but that the greatest challenges arise with multilingual titles, diacritical marks, different fonts, and longer bibliographic units. Due to the above-mentioned differences, significant deviations occur in the automatically extracted data. Bearing in mind that the journal is more than one hundred years old, recalling its importance and digitizing its basic metadata can contribute to the preservation and future reuse of the information contained in these rare issues.

Pages: 215 - 224

DOI: https://doi.org/10.62792/ut.filologjia.v14.i26-27.p3296