System update
On Tuesday, August 18th, between 12-1pm, a planned system update of DiVA will take place. During this time, DiVA will not be available.
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Towards better language representation in Natural Language Processing: A multilingual dataset for text-level Grammatical Error Correction
Show others and affiliations
Number of Authors: 302025 (English)In: International Journal of Learner Corpus Research, ISSN 2215-1478, E-ISSN 2215-1486, Vol. 11, no 2, p. 309-335Article in journal (Refereed) Published
Abstract [en]

This paper introduces MultiGEC, a dataset for multilingual Grammatical Error Correction (GEC) in twelve European languages: Czech, English, Estonian, German, Greek, Icelandic, Italian, Latvian, Russian, Slovene, Swedish and Ukrainian. MultiGEC distinguishes itself from previous GEC datasets in that it covers several underrepresented languages, which we argue should be included in resources used to train models for Natural Language Processing tasks which, as GEC itself, have implications for Learner Corpus Research and Second Language Acquisition. Aside from multilingualism, the novelty of the MultiGEC dataset is that it consists of full texts — typically learner essays — rather than individual sentences, making it possible to train systems that take a broader context into account. The dataset was built for MultiGEC-2025, the first shared task in multilingual text-level GEC, but it remains accessible after its competitive phase, serving as a resource to train new error correction systems and perform cross-lingual GEC studies.

Place, publisher, year, edition, pages
2025. Vol. 11, no 2, p. 309-335
Keywords [en]
grammatical error correction, learner corpora, Matthew effect, MultiGEC shared task, multilingual corpora
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:su:diva-243066DOI: 10.1075/ijlcr.24033.masISI: 001457603500001Scopus ID: 2-s2.0-105003035015OAI: oai:DiVA.org:su-243066DiVA, id: diva2:1957298
Available from: 2025-05-09 Created: 2025-05-09 Last updated: 2025-09-22Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Kurfalı, MurathanÖstling, Robert

Search in DiVA

By author/editor
Kurfalı, MurathanÖstling, Robert
By organisation
Department of Linguistics
In the same journal
International Journal of Learner Corpus Research
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 128 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf