Change search
Link to record
Permanent link

Direct link
Megyesi, Beáta, ProfessorORCID iD iconorcid.org/0000-0002-4838-6518
Alternative names
Publications (10 of 38) Show all publications
Yousuf, O., Djibril Diagne, E., Høgel, C., Megyesi, B. & Nivre, J. (2026). A Dataset of Wolof Ajami Manuscripts for HTR and OCR. In: Stelios Piperidis; Núria Bel; Henk van den Heuvel; Nancy Ide; Simon Krek; Antonio Toral (Ed.), Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026): . Paper presented at The Fifteenth Language Resources and Evaluation Conference (LREC 2026), Palma, Spain, May 11-16, 2026 (pp. 3234-3239).
Open this publication in new window or tab >>A Dataset of Wolof Ajami Manuscripts for HTR and OCR
Show others...
2026 (English)In: Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) / [ed] Stelios Piperidis; Núria Bel; Henk van den Heuvel; Nancy Ide; Simon Krek; Antonio Toral, 2026, p. 3234-3239Conference paper, Published paper (Refereed)
Abstract [en]

We present the first ever dataset of manually segmented and transcribed Ajami manuscripts written in Wolof. The term Ajami refers to modified Arabic-script orthographies used to transcribe African languages. Handwritten text recognition (HTR) and optical character recognition (OCR) models for Arabic-script languages perform poorly on African languages written in Ajami orthographies because these languages are not represented in the pre-training data of the models. This leads to recognition models being unable to extract unique Arabic-script letters and ubiquitous diacritics used in African languages, and struggling to adapt to various calligraphy styles used across Africa. We release the following as an open-source dataset: an ALTO formatting of high-quality images of handwritten and printed, 20th–century Wolof manuscripts; manual segmentation (region and line); and manual transcriptions. We extend our contribution by evaluating several Arabic-script recognition models intended for historical manuscripts and find they produce character error rates (CER) of 61–81%. Transcriptions produced by the evaluated recognition models, as well as a keyboard to transcribe Wolof Ajami manuscripts, are released as well. The digitally transcribed text in the dataset can also be utilized for various natural language processing (NLP) and historical linguistic tasks.

National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-257063 (URN)10.63317/4pz98ojeeqpw (DOI)978-2-493814-49-4 (ISBN)
Conference
The Fifteenth Language Resources and Evaluation Conference (LREC 2026), Palma, Spain, May 11-16, 2026
Funder
Riksbankens Jubileumsfond, M24-0028
Available from: 2026-06-19 Created: 2026-06-19 Last updated: 2026-06-22Bibliographically approved
Bruton, M., Beloucif, M. & Megyesi, B. (2026). Bridging the Low Resource Gap in Historical Cryptology: A Multilingual Diachronic Synthetic Dataset for Reproducible Cryptanalysis. In: Felix Morger; Nikolai Ilinykh; Barbara Scalvini; Simon Dobnik; Dana Dannélls (Ed.), Proceedings of the Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL). LREC 2026: Workshop Proceedings. Paper presented at Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL). LREC 2026 (pp. 13-24). ELRA Language Resources Association (ELRA)
Open this publication in new window or tab >>Bridging the Low Resource Gap in Historical Cryptology: A Multilingual Diachronic Synthetic Dataset for Reproducible Cryptanalysis
2026 (English)In: Proceedings of the Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL). LREC 2026: Workshop Proceedings / [ed] Felix Morger; Nikolai Ilinykh; Barbara Scalvini; Simon Dobnik; Dana Dannélls, ELRA Language Resources Association (ELRA) , 2026, p. 13-24Conference paper, Published paper (Refereed)
Abstract [en]

Many NLP tasks suffer from limited aligned supervision in the target domain. Historical cipher decryption represents an extreme case: aligned plaintext–ciphertext pairs are scarce, access to decrypted archives is restricted, and prior work often relies on synthetic data that is neither released nor evaluated for realism. This limits reproducibility and obscures whether models trained on synthetic benchmarks transfer to archival conditions. We introduce HistCiph,  the first publicly available multilingual collection of historically grounded plaintext–ciphertext datasets for classical ciphers. Spanning ten languages (Czech, Dutch, English, French, Hungarian, Icelandic, Italian, Polish, Spanish, Swedish) and multiple centuries, the collection combines diachronically balanced historical plaintext with independently generated homophonic substitution keys and controlled transcription noise. Synthetic generation is explicitly constrained by documented properties of historical ciphers, including multi-homophone allocation and variable-length codes. We validate the datasets using information-theoretic diagnostics—entropy, redundancy, frequency masking, and unicity distance—showing that ciphertext distributions approach theoretical bounds while preserving cross-linguistic variation. HistCiph provides a reproducible benchmark for neural decryption and alignment, and illustrates a principled framework for empirically grounded synthetic data generation in low-resource NLP.

Place, publisher, year, edition, pages
ELRA Language Resources Association (ELRA), 2026
Keywords
synthetic data, low-resource NLP, historical text, substitution ciphers, entropy, unicity distance
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-256377 (URN)978-2-493814-94-4 (ISBN)
Conference
Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL). LREC 2026
Funder
Riksbankens Jubileumsfond, M24-0028
Available from: 2026-06-06 Created: 2026-06-06 Last updated: 2026-06-08Bibliographically approved
Megyesi, B., Rattenborg, R., Láng, B., Waldispühl, M. & Héder, M. (2026). Building a Corpus and Database for Rare and Undeciphered Scripts. In: Marco Passarotti; Rachele Sprugnoli (Ed.), Fourth Workshop on Language Technologies forHistorical and Ancient Languages (LT4HALA 2026) @LREC 2026: Workshop Proceedings. Paper presented at Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA), Language Resources and Evaluation (LREC),  (pp. 184-196). European Language Resources Association
Open this publication in new window or tab >>Building a Corpus and Database for Rare and Undeciphered Scripts
Show others...
2026 (English)In: Fourth Workshop on Language Technologies forHistorical and Ancient Languages (LT4HALA 2026) @LREC 2026: Workshop Proceedings / [ed] Marco Passarotti; Rachele Sprugnoli, European Language Resources Association , 2026, p. 184-196Conference paper, Published paper (Refereed)
Abstract [en]

Historical sources written in rare or undeciphered scripts represent an immense but underexploited part of the world’s cultural and linguistic heritage. Their study is often hindered by fragmentary preservation, non-standard symbol systems, and the absence of interoperable digital resources. While recent advances in imaging, transcription, and computational analysis have improved access to historical texts, most tools rely on large quantities of labeled data and standardized encodings, requirements that are rarely met for rare or unknown writing systems. This paper presents the design and methodology of a new corpus and database dedicated to rare and undeciphered scripts worldwide. The resource integrates high-quality images, transliterations, transcriptions, linguistic annotations, and metadata within a unified data model tailored for low-resource and non-standard scripts. By adhering to FAIR principles and existing standards for linguistic and cultural heritage data, the database enables reproducible, interdisciplinary research across philology, linguistics, cryptology, and computer science. The paper outlines the data collection and digitization workflow, describes the metadata and database architecture, and demonstrates applications in analysis and decipherment.

Place, publisher, year, edition, pages
European Language Resources Association, 2026
Keywords
decipherment, historical writing systems, rare scripts
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-256378 (URN)978-2-493814-58-6 (ISBN)
Conference
Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA), Language Resources and Evaluation (LREC), 
Funder
Riksbankens Jubileumsfond, M24-0028
Available from: 2026-06-06 Created: 2026-06-06 Last updated: 2026-06-08Bibliographically approved
Heil, R., Fornés, A., Láng, B. & Megyesi, B. (2026). Establishing a Document Layout Analysis Baseline for Historical Cipher Keys. In: Camille Desenclos; Cécile Pierrot (Ed.), Proceedings of the 9th International Conference on Historical Cryptology: HistoCrypt 2026. Paper presented at The 9th International Conference on Historical Cryptology (HistoCrypt 2026), Amiens, France, June 22-24, 2026 (pp. 100-112). Tartu: Tartu University Library
Open this publication in new window or tab >>Establishing a Document Layout Analysis Baseline for Historical Cipher Keys
2026 (English)In: Proceedings of the 9th International Conference on Historical Cryptology: HistoCrypt 2026 / [ed] Camille Desenclos; Cécile Pierrot, Tartu: Tartu University Library , 2026, p. 100-112Conference paper, Published paper (Refereed)
Abstract [en]

Historical cipher keys encode mappings between plaintext elements and cipher symbols and are characterized by complex, heterogeneous handwritten layouts. This paper establishes a baseline for document layout analysis (DLA) of historical cipher keys using a newly annotated dataset of 350 images from European archives dating from ca. 1300 to 1850 CE. We evaluate four YOLO-based architectures under three conditions: training from scratch, cross-domain transfer from models pre-trained on DocLayNet and CATMuS in a class-agnostic setting, and fine-tuning of these pre-trained models on cipher key data. Results show that training from scratch is limited by data scarcity and unstable convergence, while direct transfer across DLA domains performs poorly. In contrast, fine-tuning consistently improves performance across all architectures, demonstrating the feasibil- ity of adapting existing DLA models to cipher keys and supporting downstream tasks such as key extraction and comparative cryptographic analysis.

Place, publisher, year, edition, pages
Tartu: Tartu University Library, 2026
Series
NEALT Proceedings Series, ISSN 1736-8197, E-ISSN 1736- 6305 ; 61
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-257057 (URN)9789908539997 (ISBN)
Conference
The 9th International Conference on Historical Cryptology (HistoCrypt 2026), Amiens, France, June 22-24, 2026
Projects
DESCRYPT
Funder
Riksbankens Jubileumsfond, M24-0028
Available from: 2026-06-19 Created: 2026-06-19 Last updated: 2026-06-22Bibliographically approved
Reineres, A., Fornés, A., De Gregorio, G. & Megyesi, B. (2026). Exploring the Automatic Alphabet Identification of Images of Handwritten Ciphers. In: Camille Desenclos; Cécile Pierrot (Ed.), Proceedings of the 9th International Conference on Historical Cryptology: Histocrypt 2026. Paper presented at The 9th International Conference on Historical Cryptology (HistoCrypt 2026), Amiens, France, June 22-24, 2026 (pp. 208-213). Tartu, Estonia: Tartu University Press
Open this publication in new window or tab >>Exploring the Automatic Alphabet Identification of Images of Handwritten Ciphers
2026 (English)In: Proceedings of the 9th International Conference on Historical Cryptology: Histocrypt 2026 / [ed] Camille Desenclos; Cécile Pierrot, Tartu, Estonia: Tartu University Press, 2026, p. 208-213Conference paper, Published paper (Refereed)
Abstract [en]

Historical encrypted manuscripts often use invented or heterogeneous alphabets, making alphabet identification a necessary but traditionally manual first step prior to transcription and decryption. This work explores the use of unsupervised computer vision methods to automate this task without requiring labeled data. We propose a pipeline that segments characters from cipher manuscripts, groups them into clusters of visually similar symbols using unsupervised methods, and compares those clusters against a reference database of known alphabet symbols to identify the most likely underlying writing system. Experiments show that the method can correctly identify the alphabet when a handwritten alphabet is available, but performance degrades when handwritten symbols are compared against printed alphabets, with handwriting style dominating shape similarity. These results high- light the importance of realistic handwritten reference alphabets.

Place, publisher, year, edition, pages
Tartu, Estonia: Tartu University Press, 2026
Series
NEALT Proceedings Series, ISSN 1736-8197, E-ISSN 1736- 6305 ; 61
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-257061 (URN)9789908539997 (ISBN)
Conference
The 9th International Conference on Historical Cryptology (HistoCrypt 2026), Amiens, France, June 22-24, 2026
Funder
Riksbankens Jubileumsfond, M24-0028
Available from: 2026-06-19 Created: 2026-06-19 Last updated: 2026-06-22Bibliographically approved
Oliveros-Blanco, M., Kang, L., Fornés, A. & Megyesi, B. (2026). Joint Transcription and Decryption of Images of Ciphered Handwritten Documents: A Comparison with the Traditional Pipeline. In: Camille Desenclos; Cécile Pierrot (Ed.), Proceedings of the 9th International Conference on Historical Cryptology: HistoCrypt 2026. Paper presented at The 9th International Conference on Historical Cryptology (HistoCrypt 2026), Amiens, France, June 22-24, 2026 (pp. 1-11). Tartu, Estonia: Tartu University Press
Open this publication in new window or tab >>Joint Transcription and Decryption of Images of Ciphered Handwritten Documents: A Comparison with the Traditional Pipeline
2026 (English)In: Proceedings of the 9th International Conference on Historical Cryptology: HistoCrypt 2026 / [ed] Camille Desenclos; Cécile Pierrot, Tartu, Estonia: Tartu University Press, 2026, p. 1-11Conference paper, Published paper (Refereed)
Abstract [en]

Historical encrypted manuscripts present a challenging problem at the intersection of cryptology, linguistics, paleography, and computer vision. Current automatic decipherment approaches usually rely on a two-stage pipeline: transcription of cipher symbols from manuscript images, followed by decryption into plaintext. However, this design is sensitive to transcription errors, which propagate to the final output. We present Direct Image Decryption, an end-to-end approach that directly maps encrypted manuscript images to plaintext, bypassing the intermediate transcription stage. Using the Copiale cipher as a case study, we build a synthetic data generation pipeline to create large-scale cipher-like training data and compare the traditional pipeline with the proposed joint architecture. Results show that joint image-to-plaintext modeling is a promising alternative to traditional transcription-based pipelines.

Place, publisher, year, edition, pages
Tartu, Estonia: Tartu University Press, 2026
Series
NEALT Proceedings Series, ISSN 1736-8197, E-ISSN 1736- 6305 ; 61
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-257059 (URN)9789908539997 (ISBN)
Conference
The 9th International Conference on Historical Cryptology (HistoCrypt 2026), Amiens, France, June 22-24, 2026
Funder
Riksbankens Jubileumsfond, M24-0028
Available from: 2026-06-19 Created: 2026-06-19 Last updated: 2026-06-22Bibliographically approved
Waldispühl, M. & Megyesi, B. (2026). Language Choice in Eighteenth-Century Diplomatic Ciphers from Europe. In: Gleb Kazakov, Vladislav Rjéoutski (Ed.), Languages of Diplomacy in the Early Modern World: Europe and the USA (pp. 204-228). Routledge
Open this publication in new window or tab >>Language Choice in Eighteenth-Century Diplomatic Ciphers from Europe
2026 (English)In: Languages of Diplomacy in the Early Modern World: Europe and the USA / [ed] Gleb Kazakov, Vladislav Rjéoutski, Routledge, 2026, p. 204-228Chapter in book (Refereed)
Abstract [en]

This paper investigates the language choices made in 766 entirely or partly encrypted manuscripts—ciphertexts and keys—used in diplomatic correspondence in the eighteenth century. The encrypted sources originate from archives in Austria, Hungary, the Netherlands, Saxony, and the Vatican. The comparison of language choices reveals the prominence of French in the documents from all the regions except the Vatican. However, the local languages, such as German in Saxony, Dutch in the Netherlands, and Hungarian in Hungary, were used more frequently than French in these regions. As case studies show, the choice for a non-local language, often French, depended mainly on an external addressee. In the few multilingual documents present in our dataset, the languages used have different functions.

Place, publisher, year, edition, pages
Routledge, 2026
National Category
Natural Language Processing
Research subject
Linguistics
Identifiers
urn:nbn:se:su:diva-256379 (URN)10.4324/9781003698753-12 (DOI)2-s2.0-105039406954 (Scopus ID)9781003698753 (ISBN)
Funder
Riksbankens Jubileumsfond, M24-0028
Available from: 2026-06-06 Created: 2026-06-06 Last updated: 2026-06-08Bibliographically approved
Bruton, M., Tudor, C. M., Sinneave, W., Yousuf, O., Rirdance, S., Heil, R. & Megyesi, B. (2026). Language Similarity and Cross-Lingual Transfer in Historical HTR: Evidence from Swedish, Norwegian, and Medieval Latin. In: Proceedings of the 8th International Workshop on Historical Document Imaging and Processing (HIP’26): . Paper presented at The 8th International Workshop on Historical Document Imaging and Processing (HIP’26).
Open this publication in new window or tab >>Language Similarity and Cross-Lingual Transfer in Historical HTR: Evidence from Swedish, Norwegian, and Medieval Latin
Show others...
2026 (English)In: Proceedings of the 8th International Workshop on Historical Document Imaging and Processing (HIP’26), 2026Conference paper, Published paper (Refereed)
Abstract [en]

Handwritten text recognition (HTR) for historical documents is often limited by the scarcity of annotated training data, especially for under-resourced languages and collections. This paper examines how linguistic distance affects cross-lingual transfer in historical HTR, focusing on zero-shot transfer, target-language fine-tuning, and multilingual training under low-resource conditions.

We evaluate three historical line-level HTR datasets: Riksarkivet for Swedish, NorHand v3 for Norwegian Bokmål, and HOME-Alcar for Medieval Latin and Old French. Using a fully crossed experimental design, models trained on each dataset are tested across all datasets and subsequently fine-tuned on target-language data. The same data configurations are replicated in two HTR pipelines to assess whether transfer patterns are robust across implementations.

The results show that zero-shot transfer is substantially more effective between closely related languages than between linguistically distant ones, highlighting the importance of linguistic similarity for cross-lingual generalization. Fine-tuning consistently improves transfer performance and can approach monolingual baselines, while multilingual training mitigates degradation in severely low-resource settings. Qualitative analysis further reveals characteristic orthographic and graphemic transfer errors across different linguistic settings.

Keywords
HTR, multilingual fine-tuning, cross-lingual transfer learning, historical document processing
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-259279 (URN)
Conference
The 8th International Workshop on Historical Document Imaging and Processing (HIP’26)
Funder
Riksbankens Jubileumsfond, M24-0028
Available from: 2026-09-08 Created: 2026-09-08 Last updated: 2026-09-08Bibliographically approved
Kang, L., De Gregorio, G., Heil, R., Fornés, A. & Megyesi, B. (2026). Learning to Decipher from Pixels — A Case Study of Copiale. In: Camille Desenclos; Cécile Pierrot (Ed.), Proceedings of the 9th International Conference on Historical Cryptology: HistoCrypt 2026. Paper presented at The 9th International Conference on Historical Cryptology (HistoCrypt 2026), Amiens, France, June 22-24, 2026 (pp. 22-27). Tartu: Tartu University Press
Open this publication in new window or tab >>Learning to Decipher from Pixels — A Case Study of Copiale
Show others...
2026 (English)In: Proceedings of the 9th International Conference on Historical Cryptology: HistoCrypt 2026 / [ed] Camille Desenclos; Cécile Pierrot, Tartu: Tartu University Press, 2026, p. 22-27Conference paper, Published paper (Refereed)
Abstract [en]

Historical encrypted manuscripts require both paleographic interpretation of cipher symbols and cryptanalytic recovery of plaintext. Most existing computational workflows rely on a transcription-first paradigm, in which handwritten symbols are transcribed prior to decipherment. This intermediate step is labor-intensive, error-prone, and not always aligned with the goal of direct plaintext recovery. We propose an end-to-end, transcription-free approach that directly maps handwrit- ten cipher images to plaintext. Using the Copiale cipher as a case study, we introduce the first text-line-level dataset pairing cipher images with German plain- text. We show that pretraining on generic handwriting data followed by cipher-specific fine-tuning substantially improvesdecipherment accuracy. Our results demonstrate that transcription-free image- to-plaintext decipherment is both feasible and effective for historical substitution ciphers, offering a simplified and scalable alternative to traditional pipelines. https://github.com/leitro/Decipher-from-Pixels-Copiale.

Place, publisher, year, edition, pages
Tartu: Tartu University Press, 2026
Series
NEALT Proceedings Series Number, ISSN 1736-8197, E-ISSN 1736- 6305 ; 61
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-257058 (URN)9789908539997 (ISBN)
Conference
The 9th International Conference on Historical Cryptology (HistoCrypt 2026), Amiens, France, June 22-24, 2026
Funder
Riksbankens Jubileumsfond, M24-0028
Available from: 2026-06-19 Created: 2026-06-19 Last updated: 2026-06-22Bibliographically approved
Megyesi, B. & Ruan, R. (2026). SWEGRAM: Guidelines to Annotation and Analysis of English and Swedish Texts. Stockholm
Open this publication in new window or tab >>SWEGRAM: Guidelines to Annotation and Analysis of English and Swedish Texts
2026 (English)Report (Other academic)
Place, publisher, year, edition, pages
Stockholm: , 2026. p. 46
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-257054 (URN)
Available from: 2026-06-19 Created: 2026-06-19 Last updated: 2026-06-22Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-4838-6518

Search in DiVA

Show all publications