Change search
Link to record
Permanent link

Direct link
Publications (10 of 32) Show all publications
Volodina, E., Masciolini, A., Megyesi, B., Prentice, J., Rudebeck, L., Sundberg, G. & Wirén, M. (2025). SweLL with pride: How to put a learner corpus to good use. In: Gerlof Bouma, Dana Dannélls Dimitrios Kokkinakis, Elena Volodina (Ed.), Huminfra handbook: Empowering digital and experimental humanities (pp. 251-306). Tartu: University of Tartu Library
Open this publication in new window or tab >>SweLL with pride: How to put a learner corpus to good use
Show others...
2025 (English)In: Huminfra handbook: Empowering digital and experimental humanities / [ed] Gerlof Bouma, Dana Dannélls Dimitrios Kokkinakis, Elena Volodina, Tartu: University of Tartu Library , 2025, p. 251-306Chapter in book (Refereed)
Abstract [en]

Second language (L2) learner corpora are collections of language samples that demonstrate learners’ abilities to perform some learning tasks, e.g. an ability to write essays, answer to reading comprehension questions, or talk on a given topic. Such corpora are necessary for both empirical-based research within Second Language Acquisition (SLA), and for development of methods for automatic processing of such data. L2 corpora are notoriously difficult to collect, and their value depends to a greater degree on the representativeness and balance of the sampled data, type of associated metadata and reliability of manual annotations. In this chapter we thoroughly describe the SweLL-gold corpus of L2 Swedish, its annotation, statistics and metadata, and showcase main types of its use, such as (1) in research on SLA through detailed instructions on how to perform corpus searches given SweLL-specific annotation, combined with guidelines for SVALA usage, a tool for correction annotation; and (2) in NLP research on problems such as grammatical error correction through guidelines on how to use the different available file formats that the SweLL-gold corpus is released in. Both cases are further supported by case studies and, where available, relevant scripts ready for reuse by researchers.

Place, publisher, year, edition, pages
Tartu: University of Tartu Library, 2025
Series
NEALT Proceedings Series, ISSN 1736-8197, E-ISSN 1736-6305 ; 59
Keywords
Learner language corpus, Swedish as a second language, Inlärarkorpus, svenska som andraspråk
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-249680 (URN)10.58009/aere-perennius0178 (DOI)978-99-0853-612-5 (ISBN)978-91-531-7077-8 (ISBN)
Funder
Swedish Research Council, 2021-00176
Available from: 2025-11-17 Created: 2025-11-17 Last updated: 2025-12-01Bibliographically approved
Sundberg, G. & Rudebeck, L. (2024). On the other side of the error tag: Predefined normalized texts as a basis for correction annotation. In: Katherine Ackerley; Erik Castello (Ed.), Continuing Learner Corpus Research: Challenges and Opportunities (pp. 123-152). Louvain-la-Neuve: Presses universitaires de Louvain
Open this publication in new window or tab >>On the other side of the error tag: Predefined normalized texts as a basis for correction annotation
2024 (English)In: Continuing Learner Corpus Research: Challenges and Opportunities / [ed] Katherine Ackerley; Erik Castello, Louvain-la-Neuve: Presses universitaires de Louvain, 2024, p. 123-152Chapter in book (Refereed)
Abstract [en]

This article on learner corpus design presents a novel methodological approach, founded on a systematic and clear separation between normalization (‘reconstruction’, ‘standardization’) and correction annotation (‘error annotation’), where the correction annotation is based on a predefined normalized version of each learner text, and where the normalized texts are presented and accessible as a separate corpus, parallel to the corpus of the original learner texts. It is argued that the nature of the normalization process, as distinct from the correction annotation, requires a holistic and interpretation-based approach, which increases the validity of the correction annotation as an analysis of differences between 1) the writer’s choice of expression, and 2) a normadhering way of formulating the message intended by the writer. The case in point is the recently released Swedish learner language corpus (SweLL, Volodina et al. 2019).

Place, publisher, year, edition, pages
Louvain-la-Neuve: Presses universitaires de Louvain, 2024
Series
Corpora and Language in Use Proceedings, ISSN 2034-6417 ; 7
Keywords
normalization, language learner corpus, Swedish as a second language, normalisering, inlärarkorpus, svenska som andraspråk
National Category
Comparative Language Studies and Linguistics
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-242154 (URN)978-2-39061-525-5 (ISBN)
Projects
SweLL Swedish Learner Language Corpus
Funder
Riksbankens Jubileumsfond, IN16-0464:1
Available from: 2025-04-14 Created: 2025-04-14 Last updated: 2025-04-15Bibliographically approved
Nelson, M., Michanek, M., Rydell, M., Sayehli, S., Skogmyr Marian, K. & Sundberg, G. (2023). Språk i praktiken - i en föränderlig värld: Rapport från ASLA-symposiet, Stockholms universitet, 7-8 april 2022. Stockholm: Svenska föreningen för tillämpad språkvetenskap
Open this publication in new window or tab >>Språk i praktiken - i en föränderlig värld: Rapport från ASLA-symposiet, Stockholms universitet, 7-8 april 2022
Show others...
2023 (Swedish)Report (Refereed)
Alternative title[en]
Language in practice - in a changing world
Place, publisher, year, edition, pages
Stockholm: Svenska föreningen för tillämpad språkvetenskap, 2023. p. 465
Series
ASLA:s skriftserie, ISSN 1100-5629 ; 30
Keywords
ASLA, tillämpad språkvetenskap, språkbruk, flerspråkighet, språkutveckling
National Category
General Language Studies and Linguistics
Research subject
Linguistics
Identifiers
urn:nbn:se:su:diva-223306 (URN)10.17045/sthlmuni.24321526.v1 (DOI)978-91-87884-30-6 (ISBN)
Available from: 2023-10-25 Created: 2023-10-25 Last updated: 2023-11-27Bibliographically approved
Sundberg, G. & Prentice, J. (2023). SweLL: En svensk inlärarkorpus. In: Marie Nelson; Mårten Michanek; Maria Rydell; Susan Sayehli; Klara Skogmyr Marian; Gunlög Sundberg (Ed.), Språk i praktiken – i en föränderlig värld. Rapport från ASLA-symposiet. Stockholms universitet, 7–8, april 2022: [Language in practice – in a changing world. Papers from the ASLA symposium. Stockholm University, 7–8 April, 2022]. Paper presented at 29:e ASLA-symposiet "Språk i praktiken – i en föränderlig värld", Stockholm, Sverige, 7-8 april, 2022 (pp. 428-452). ASLA
Open this publication in new window or tab >>SweLL: En svensk inlärarkorpus
2023 (Swedish)In: Språk i praktiken – i en föränderlig värld. Rapport från ASLA-symposiet. Stockholms universitet, 7–8, april 2022: [Language in practice – in a changing world. Papers from the ASLA symposium. Stockholm University, 7–8 April, 2022] / [ed] Marie Nelson; Mårten Michanek; Maria Rydell; Susan Sayehli; Klara Skogmyr Marian; Gunlög Sundberg, ASLA , 2023, p. 428-452Conference paper, Published paper (Refereed)
Abstract [sv]

Språkbanken lanserade 2021 en ny annoterad digital inlärarkorpus med texter skrivna av inlärare av svenska, SweLL (Swedish Learner Language corpus). Syftet med artikeln är att diskutera och ge exempel på hur digitala inlärarkorpusar kan bidra till språkforskningen om inlärares språkanvändning, vilket är en fråga som fått allt mer uppmärksamhet (Le Bruyn & Paquot, 2021; Myles, 2015). Genom att utgå ifrån diskussionen om relationen mellan forskningsfälten LCR (Learner Corpus Research) och SLA (Second Language Acquisition) kommer vi i artikeln att presentera den nya inlärarkorpusen och diskutera några viktiga teoretiska och metodologiska aspekter i samband med dess utveckling (Volodina m.fl., 2019).Vi ger även exempel på några tillämpningar av korpusen genom en explorativ studie av förflyttningsverbkonstruktioner i inlärartexterna. Resultaten på gruppnivå stärker tidigare forskning som visar en överrepresentation av verbkonstruktioner med gå hos tidiga inlärare.

Place, publisher, year, edition, pages
ASLA, 2023
Series
ASLA:s skriftserie, ISSN 1100-5629, E-ISSN 2004-108X ; 30
Keywords
Inlärarkorpus, Svenska som andraspråk, Learner corpus research, Second language acquisition
National Category
General Language Studies and Linguistics
Research subject
Scandinavian Languages
Identifiers
urn:nbn:se:su:diva-223841 (URN)978-91-87884-30-6 (ISBN)
Conference
29:e ASLA-symposiet "Språk i praktiken – i en föränderlig värld", Stockholm, Sverige, 7-8 april, 2022
Projects
SweLL Swedish Learner Language Corpus RJ Infrastruktur
Funder
Riksbankens Jubileumsfond, IN16-0464:1
Available from: 2023-12-08 Created: 2023-12-08 Last updated: 2026-02-06Bibliographically approved
Rydell, M., Sundberg, G., Nelson, M., Michanek, M., Sayehli, S. & Skogmyr Marian, K. (2023). Tillämpad språkvetenskap – i en föränderlig tid. In: Marie Nelson, Mårten Michanek, Maria Rydell, Susan Sayehli, Klara Skogmyr Marian, Gunlög Sundberg (Ed.), Språk i praktiken – i en föränderlig värld: Language in practice – in a changing world. Paper presented at ASLA symposium, Stockholm, Sverige, 7-8 april, 2022 (pp. 7-15). Stockholm: ASLA
Open this publication in new window or tab >>Tillämpad språkvetenskap – i en föränderlig tid
Show others...
2023 (Swedish)In: Språk i praktiken – i en föränderlig värld: Language in practice – in a changing world / [ed] Marie Nelson, Mårten Michanek, Maria Rydell, Susan Sayehli, Klara Skogmyr Marian, Gunlög Sundberg, Stockholm: ASLA , 2023, p. 7-15Conference paper, Published paper (Other academic)
Place, publisher, year, edition, pages
Stockholm: ASLA, 2023
Series
ASLA:s skriftserie, ISSN 1100-5629, E-ISSN 2004-108X ; 30
Keywords
tillämpad språkvetenskap, ASLA
National Category
Specific Languages
Research subject
Scandinavian Languages
Identifiers
urn:nbn:se:su:diva-223556 (URN)978-91-87884-30-6 (ISBN)
Conference
ASLA symposium, Stockholm, Sverige, 7-8 april, 2022
Available from: 2023-11-02 Created: 2023-11-02 Last updated: 2024-09-06Bibliographically approved
Rudebeck, L. & Sundberg, G. (2021). SweLL correction annotation guidelines. Göteborg: Göteborgs universitet
Open this publication in new window or tab >>SweLL correction annotation guidelines
2021 (English)Report (Other academic)
Place, publisher, year, edition, pages
Göteborg: Göteborgs universitet, 2021. p. 78
Series
Forskningsrapporter från institutionen för svenska språket, Göteborgs universitet, ISSN 1401-5919 ; GU-ISS-2021-04
Keywords
learner corpus, Swedish as a second language, writing
National Category
Specific Languages
Research subject
Computational Linguistics; Scandinavian Languages
Identifiers
urn:nbn:se:su:diva-196024 (URN)
Projects
SweLL Infrastructure for Learner Language Corpus
Funder
Riksbankens Jubileumsfond, IN16-0464:1
Available from: 2021-08-30 Created: 2021-08-30 Last updated: 2024-09-03Bibliographically approved
Rudebeck, L., Sundberg, G. & Wirén, M. (2021). SweLL normalization guidelines: The SweLL guideline series nr 3. Göteborg: Inst för svenska, Göteborgs universitet
Open this publication in new window or tab >>SweLL normalization guidelines: The SweLL guideline series nr 3
2021 (English)Report (Other academic)
Place, publisher, year, edition, pages
Göteborg: Inst för svenska, Göteborgs universitet, 2021. p. 9
Series
Forskningsrapporter från institutionen för svenska språket, Göteborgs universitet, ISSN 1401-5919 ; GU-ISS-2021-03
Keywords
Learner corpus, Swedish as a second language, writing
National Category
Specific Languages
Research subject
Computational Linguistics; Scandinavian Languages
Identifiers
urn:nbn:se:su:diva-196023 (URN)
Projects
SweLL Infrastructure for Swedish Learner Language corpus
Funder
Riksbankens Jubileumsfond, IN16-0464:1
Available from: 2021-08-30 Created: 2021-08-30 Last updated: 2022-02-25Bibliographically approved
Volodina, E., Granstedt, L., Matsson, A., Megyesi, B., Pilán, I., Prentice, J., . . . Wirén, M. (2019). The Swell Language Learner Corpus: From Design to Annotation. Northern European Journal of Language Technology (NEJLT), 6, 67-104, Article ID 4.
Open this publication in new window or tab >>The Swell Language Learner Corpus: From Design to Annotation
Show others...
2019 (English)In: Northern European Journal of Language Technology (NEJLT), ISSN 2000-1533, Vol. 6, p. 67-104, article id 4Article in journal (Refereed) Published
Abstract [en]

The article presents a new language learner corpus for Swedish, SweLL, and the methodology from collection and pesudonymisation to protect personal information of learners to annotation adapted to second language learning. The main aim is to deliver a well-annotated corpus of essays written by second language learners of Swedish and make it available for research through a browsable environment. To that end, a new annotation tool and a new project management tool have been implemented, both with the main purpose to ensure reliability and quality of the final corpus. In the article we discuss reasoning behind metadata selection, principles of gold corpus compilation and argue for separation of normalization from correction annotation.

Keywords
Second-language learning
National Category
General Language Studies and Linguistics
Research subject
Computational Linguistics; Language Education
Identifiers
urn:nbn:se:su:diva-185228 (URN)10.3384/nejlt.2000-1533.19667 (DOI)
Funder
Riksbankens Jubileumsfond, N16-0464:1
Note

Special Issue of Selected Contributions from the Seventh Swedish Language Technology Conference (SLTC 2018)

Available from: 2020-09-18 Created: 2020-09-18 Last updated: 2023-09-08Bibliographically approved
Volodina, E., Granstedt, L., Megyesi, B., Prentice, J., Rosén, D., Schenström, C.-J., . . . Wirén, M. (2018). Annotation of learner corpora: first SweLL insights. In: Proceedings of 7th Workshop on NLP for Computer Assisted Language Learning at SLTC 2018: . Paper presented at The Seventh Swedish Language Technology Conference (SLTC), Stockholm, Sweden, November 7-9, 2018.
Open this publication in new window or tab >>Annotation of learner corpora: first SweLL insights
Show others...
2018 (English)In: Proceedings of 7th Workshop on NLP for Computer Assisted Language Learning at SLTC 2018, 2018Conference paper, Published paper (Refereed)
National Category
General Language Studies and Linguistics
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-195286 (URN)
Conference
The Seventh Swedish Language Technology Conference (SLTC), Stockholm, Sweden, November 7-9, 2018
Projects
Swedish Language Learner Corpora SweLL
Funder
Riksbankens Jubileumsfond, IN16-0464:1
Available from: 2021-08-11 Created: 2021-08-11 Last updated: 2023-09-08Bibliographically approved
Megyesi, B., Granstedt, L., Johansson, S., Prentice, J., Rosén, D., Schenström, C.-J., . . . Volodina, E. (2018). Learner Corpus Anonymization in the Age of GDPR: Insights from the Creation of a Learner Corpus of Swedish. In: Proceedings of the 7th Workshop on NLP for Computer Assisted Language Learning at SLTC 2018 (NLP4CALL 2018): . Paper presented at 7th Workshop on NLP for Computer Assisted Language Learning at SLTC 2018 (NLP4CALL 2018), Stockholm, Sweden, 7th November, 2018 (pp. 47-56). Linköping: Linköping University Electronic Press, Article ID 006.
Open this publication in new window or tab >>Learner Corpus Anonymization in the Age of GDPR: Insights from the Creation of a Learner Corpus of Swedish
Show others...
2018 (English)In: Proceedings of the 7th Workshop on NLP for Computer Assisted Language Learning at SLTC 2018 (NLP4CALL 2018), Linköping: Linköping University Electronic Press, 2018, p. 47-56, article id 006Conference paper, Published paper (Refereed)
Abstract [en]

This paper reports on the status of learner corpus anonymization for the ongoing research infrastructure project SweLL. The main project aim is to deliver and make available for research a well-annotated corpus of essays written by second language (L2) learners of Swedish. As the practice shows, annotation of learner texts is a sensitive process demanding a lot of compromises between ethical and legal demands on the one hand, and research and technical demands, on the other. Below, is a concise description of the current status of pseudonymization of language learner data to ensure anonymity of the learners, with numerous examples of the above-mentioned compromises.

Place, publisher, year, edition, pages
Linköping: Linköping University Electronic Press, 2018
Series
Linköping Electronic Conference Proceedings, ISSN 1650-3686, E-ISSN 1650-3740 ; 152
National Category
General Language Studies and Linguistics Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-162706 (URN)978-91-7685-173-9 (ISBN)
Conference
7th Workshop on NLP for Computer Assisted Language Learning at SLTC 2018 (NLP4CALL 2018), Stockholm, Sweden, 7th November, 2018
Funder
Riksbankens Jubileumsfond, IN16-0464:1
Available from: 2018-12-07 Created: 2018-12-07 Last updated: 2025-02-01Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0003-0518-9036

Search in DiVA

Show all publications