Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
SweLL with pride: How to put a learner corpus to good use
Stockholm University, Faculty of Humanities, Department of Linguistics.ORCID iD: 0000-0002-4838-6518
Show others and affiliations
2025 (English)In: Huminfra handbook: Empowering digital and experimental humanities / [ed] Gerlof Bouma, Dana Dannélls Dimitrios Kokkinakis, Elena Volodina, Tartu: University of Tartu Library , 2025, p. 251-306Chapter in book (Refereed)
Abstract [en]

Second language (L2) learner corpora are collections of language samples that demonstrate learners’ abilities to perform some learning tasks, e.g. an ability to write essays, answer to reading comprehension questions, or talk on a given topic. Such corpora are necessary for both empirical-based research within Second Language Acquisition (SLA), and for development of methods for automatic processing of such data. L2 corpora are notoriously difficult to collect, and their value depends to a greater degree on the representativeness and balance of the sampled data, type of associated metadata and reliability of manual annotations. In this chapter we thoroughly describe the SweLL-gold corpus of L2 Swedish, its annotation, statistics and metadata, and showcase main types of its use, such as (1) in research on SLA through detailed instructions on how to perform corpus searches given SweLL-specific annotation, combined with guidelines for SVALA usage, a tool for correction annotation; and (2) in NLP research on problems such as grammatical error correction through guidelines on how to use the different available file formats that the SweLL-gold corpus is released in. Both cases are further supported by case studies and, where available, relevant scripts ready for reuse by researchers.

Place, publisher, year, edition, pages
Tartu: University of Tartu Library , 2025. p. 251-306
Series
NEALT Proceedings Series, ISSN 1736-8197, E-ISSN 1736-6305 ; 59
Keywords [en]
Learner language corpus, Swedish as a second language
Keywords [sv]
Inlärarkorpus, svenska som andraspråk
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
URN: urn:nbn:se:su:diva-249680DOI: 10.58009/aere-perennius0178ISBN: 978-99-0853-612-5 (electronic)ISBN: 978-91-531-7077-8 (print)OAI: oai:DiVA.org:su-249680DiVA, id: diva2:2014420
Funder
Swedish Research Council, 2021-00176Available from: 2025-11-17 Created: 2025-11-17 Last updated: 2025-12-01Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full text

Authority records

Megyesi, BeátaRudebeck, LisaSundberg, GunlögWirén, Mats

Search in DiVA

By author/editor
Megyesi, BeátaRudebeck, LisaSundberg, GunlögWirén, Mats
By organisation
Department of LinguisticsDepartment of Swedish Language and Multilingualism
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetric score

doi
isbn
urn-nbn
Total: 103 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf