Change search
Link to record
Permanent link

Direct link
Publications (10 of 58) Show all publications
Volodina, E., Masciolini, A., Megyesi, B., Prentice, J., Rudebeck, L., Sundberg, G. & Wirén, M. (2025). SweLL with pride: How to put a learner corpus to good use. In: Gerlof Bouma, Dana Dannélls Dimitrios Kokkinakis, Elena Volodina (Ed.), Huminfra handbook: Empowering digital and experimental humanities (pp. 251-306). Tartu: University of Tartu Library
Open this publication in new window or tab >>SweLL with pride: How to put a learner corpus to good use
Show others...
2025 (English)In: Huminfra handbook: Empowering digital and experimental humanities / [ed] Gerlof Bouma, Dana Dannélls Dimitrios Kokkinakis, Elena Volodina, Tartu: University of Tartu Library , 2025, p. 251-306Chapter in book (Refereed)
Abstract [en]

Second language (L2) learner corpora are collections of language samples that demonstrate learners’ abilities to perform some learning tasks, e.g. an ability to write essays, answer to reading comprehension questions, or talk on a given topic. Such corpora are necessary for both empirical-based research within Second Language Acquisition (SLA), and for development of methods for automatic processing of such data. L2 corpora are notoriously difficult to collect, and their value depends to a greater degree on the representativeness and balance of the sampled data, type of associated metadata and reliability of manual annotations. In this chapter we thoroughly describe the SweLL-gold corpus of L2 Swedish, its annotation, statistics and metadata, and showcase main types of its use, such as (1) in research on SLA through detailed instructions on how to perform corpus searches given SweLL-specific annotation, combined with guidelines for SVALA usage, a tool for correction annotation; and (2) in NLP research on problems such as grammatical error correction through guidelines on how to use the different available file formats that the SweLL-gold corpus is released in. Both cases are further supported by case studies and, where available, relevant scripts ready for reuse by researchers.

Place, publisher, year, edition, pages
Tartu: University of Tartu Library, 2025
Series
NEALT Proceedings Series, ISSN 1736-8197, E-ISSN 1736-6305 ; 59
Keywords
Learner language corpus, Swedish as a second language, Inlärarkorpus, svenska som andraspråk
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-249680 (URN)10.58009/aere-perennius0178 (DOI)978-99-0853-612-5 (ISBN)978-91-531-7077-8 (ISBN)
Funder
Swedish Research Council, 2021-00176
Available from: 2025-11-17 Created: 2025-11-17 Last updated: 2025-12-01Bibliographically approved
Kurfali, M., Östling, R., Sjons, J. & Wirén, M. (2020). A Multi-Word Expression Dataset for Swedish. In: Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020): . Paper presented at 12th Conference on Language Resources and Evaluation (LREC 2020), Marseille, France, May 11–16, 2020 (pp. 4402-4409). Marseille: European Language Resources Association (ELRA)
Open this publication in new window or tab >>A Multi-Word Expression Dataset for Swedish
2020 (English)In: Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), Marseille: European Language Resources Association (ELRA) , 2020, p. 4402-4409Conference paper, Published paper (Refereed)
Abstract [en]

We present a new set of 96 Swedish multi-word expressions annotated with degree of (non-)compositionality. In contrast to most previous compositionality datasets we also consider syntactically complex constructions and publish a formal specification of each expression. This allows evaluation of computational models beyond word bigrams, which have so far been the norm. Finally, we use the annotations to evaluate a system for automatic compositionality estimation based on distributional semantics. Our analysis of the disagreements between human annotators and the distributional model reveal interesting questions related to the perception of compositionality, and should be informative to future work in the area.

Place, publisher, year, edition, pages
Marseille: European Language Resources Association (ELRA), 2020
Keywords
multi-word expressions, compositionality, distributional semantic
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-192115 (URN)
Conference
12th Conference on Language Resources and Evaluation (LREC 2020), Marseille, France, May 11–16, 2020
Available from: 2021-04-12 Created: 2021-04-12 Last updated: 2025-02-07Bibliographically approved
Kurfali, M. & Wirén, M. (2020). Zero-shot cross-lingual identification of direct speech using distant supervision. In: The 4th Joint SIGHUM Workshopon Computational Linguistics for Cultural Heritage,Social Sciences, Humanities and Literature: Co-located with the 28th International Conferenceon Computational Linguistics COLING’2020. Paper presented at Proceedings of LaTeCH-CLfL 2020, Barcelona, Spain (Online), December 12, 2020 (pp. 105-111).
Open this publication in new window or tab >>Zero-shot cross-lingual identification of direct speech using distant supervision
2020 (English)In: The 4th Joint SIGHUM Workshopon Computational Linguistics for Cultural Heritage,Social Sciences, Humanities and Literature: Co-located with the 28th International Conferenceon Computational Linguistics COLING’2020, 2020, p. 105-111Conference paper, Published paper (Refereed)
National Category
Natural Language Processing
Identifiers
urn:nbn:se:su:diva-189701 (URN)978-1-952148-34-7 (ISBN)
Conference
Proceedings of LaTeCH-CLfL 2020, Barcelona, Spain (Online), December 12, 2020
Available from: 2021-01-31 Created: 2021-01-31 Last updated: 2025-02-07Bibliographically approved
Wirén, M., Ek, A. & Kasaty, A. (2019). Annotation Guideline No. 7: Guidelines for annotation of narrative structure. Journal of Cultural Analytics, 4(3)
Open this publication in new window or tab >>Annotation Guideline No. 7: Guidelines for annotation of narrative structure
2019 (English)In: Journal of Cultural Analytics, ISSN 2371-4549, Vol. 4, no 3Article in journal (Refereed) Published
Abstract [en]

Analysis of narrative structure can be said to answer the question “Who tells what, and how?”. The first part of the question thus concerns aspects such as who is narrating, whether it is a character in the story or not, and if it is a first-person or third-person narrator. The second part is related to the story and its basic elements: characters and events, and how the sequence of events forms a plot. The third part concerns how the narrative text is constructed: ordering of the events, the perspective from which the story is seen, how much information the narrator has access to, etc.

Keywords
literary computing, narratology, annotation
National Category
General Language Studies and Linguistics
Research subject
Literature; Computational Linguistics
Identifiers
urn:nbn:se:su:diva-197006 (URN)10.22148/001c.11772 (DOI)
Funder
Swedish Research Council, 2017-00626
Available from: 2021-09-21 Created: 2021-09-21 Last updated: 2021-11-26Bibliographically approved
Ek, A. & Wirén, M. (2019). Distinguishing Narration and Speech in Prose Fiction Dialogues. In: Costanza Navarretta, Manex Agirrezabal, Bente Maegaard (Ed.), Proceedings of the Digital Humanities in the Nordic Countries 4th Conference: . Paper presented at Digital Humanities in the Nordic Countries 4th Conference (DHN), Copenhagen, Denmark, March 5-8, 2019 (pp. 124-132). CEUR-WS.org
Open this publication in new window or tab >>Distinguishing Narration and Speech in Prose Fiction Dialogues
2019 (English)In: Proceedings of the Digital Humanities in the Nordic Countries 4th Conference / [ed] Costanza Navarretta, Manex Agirrezabal, Bente Maegaard, CEUR-WS.org , 2019, p. 124-132Conference paper, Published paper (Refereed)
Abstract [en]

This paper presents a supervised method for a novel task, namely, detecting elements of narration in passages of dialogue in prose fiction. The method achieves an F1-score of 80.8%, exceeding the best baseline by almost 33 percentage points. The purpose of the method is to enable a more fine-grained analysis of fictional dialogue than has previously been possible, and to provide a component for the further analysis of narrative structure in general.

Place, publisher, year, edition, pages
CEUR-WS.org, 2019
Series
CEUR Workshop Proceedings, E-ISSN 1613-0073 ; 2364
Keywords
Prose fiction, Literary dialogue, Characters’ discourse, Narrative structure
National Category
General Language Studies and Linguistics
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-170361 (URN)
Conference
Digital Humanities in the Nordic Countries 4th Conference (DHN), Copenhagen, Denmark, March 5-8, 2019
Funder
Swedish Research Council, 821-2013-2003
Available from: 2019-06-27 Created: 2019-06-27 Last updated: 2022-02-26Bibliographically approved
Dalianis, H., Östling, R., Weegar, R. & Wirén, M. (Eds.). (2019). Special Issue of Selected Contributions from the Seventh Swedish Language Technology Conference (SLTC 2018). Paper presented at Seventh Swedish Language Technology Conference (SLTC 2018), Stockholm, Sweden, 8–9 November, 2018. Linköping University Electronic Press
Open this publication in new window or tab >>Special Issue of Selected Contributions from the Seventh Swedish Language Technology Conference (SLTC 2018)
2019 (English)Conference proceedings (editor) (Other academic)
Abstract [en]

This Special Issue contains three papers that are extended versions of abstracts presented at the Seventh Swedish Language Technology Conference (SLTC 2018), held at Stockholm University 8–9 November 2018.1 SLTC 2018 received 34 submissions, of which 31 were accepted for presentation. The number of registered participants was 113, including both attendees at SLTC 2018 and two co-located workshops that took place on 7 November. 32 participants were internationally affiliated, of which 14 were from outside the Nordic countries. Overall participation was thus on a par with previous editions of SLTC, but international participation was higher.

Place, publisher, year, edition, pages
Linköping University Electronic Press, 2019
National Category
Natural Language Processing
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-177731 (URN)10.3384/nejlt.2000-1533.196 (DOI)
Conference
Seventh Swedish Language Technology Conference (SLTC 2018), Stockholm, Sweden, 8–9 November, 2018
Note

Special Issue: Northern European Journal of Language Technology, 2019, Vol. 6, Article 1, pp 1–3.

Available from: 2020-01-07 Created: 2020-01-07 Last updated: 2025-02-07Bibliographically approved
Wirén, M., Matsson, A., Rosén, D. & Volodina, E. (2019). SVALA: Annotation of Second-Language Learner Text Based on Mostly Automatic Alignment of Parallel Corpora. In: Inguna Skadina, Maria Eskevich (Ed.), Selected papers from the CLARIN Annual Conference 2018, Pisa, 8-10 October 2018: . Paper presented at CLARIN Annual Conference, Pisa, Italy, 8-10 October, 2018 (pp. 222-234). Linköping: Linköping University Electronic Press, Article ID 023.
Open this publication in new window or tab >>SVALA: Annotation of Second-Language Learner Text Based on Mostly Automatic Alignment of Parallel Corpora
2019 (English)In: Selected papers from the CLARIN Annual Conference 2018, Pisa, 8-10 October 2018 / [ed] Inguna Skadina, Maria Eskevich, Linköping: Linköping University Electronic Press, 2019, p. 222-234, article id 023Conference paper, Published paper (Refereed)
Abstract [en]

Annotation of second-language learner text is a cumbersome manual task which in turn requires interpretation to postulate the intended meaning of the learner’s language. This paper describes SVALA, a tool which separates the logical steps in this process while providing rich visual support for each of them. The first step is to pseudonymize the learner text to fulfil the legal and ethical requirements for a distributable learner corpus. The second step is to correct the text, which is carried out in the simplest possible way by text editing. During the editing, SVALA automatically maintains a parallel corpus with alignments between words in the learner source text and corrected text, while the annotator may repair inconsistent word alignments. Finally, the actual labelling of the corrections (the postulated errors) is performed. We describe the objectives, design and workflow of SVALA, and our plans for further development.

Place, publisher, year, edition, pages
Linköping: Linköping University Electronic Press, 2019
Series
Linköping Electronic Conference Proceedings, ISSN 1650-3686, E-ISSN 1650-3740 ; 159
Keywords
Normalization, Error annotation, Learner corpora, Parallel corpora, Word alignment
National Category
General Language Studies and Linguistics
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-170363 (URN)978-91-7685-034-3 (ISBN)
Conference
CLARIN Annual Conference, Pisa, Italy, 8-10 October, 2018
Funder
Riksbankens Jubileumsfond, IN16- 0464:1
Available from: 2019-06-27 Created: 2019-06-27 Last updated: 2022-02-26Bibliographically approved
Volodina, E., Granstedt, L., Matsson, A., Megyesi, B., Pilán, I., Prentice, J., . . . Wirén, M. (2019). The Swell Language Learner Corpus: From Design to Annotation. Northern European Journal of Language Technology (NEJLT), 6, 67-104, Article ID 4.
Open this publication in new window or tab >>The Swell Language Learner Corpus: From Design to Annotation
Show others...
2019 (English)In: Northern European Journal of Language Technology (NEJLT), ISSN 2000-1533, Vol. 6, p. 67-104, article id 4Article in journal (Refereed) Published
Abstract [en]

The article presents a new language learner corpus for Swedish, SweLL, and the methodology from collection and pesudonymisation to protect personal information of learners to annotation adapted to second language learning. The main aim is to deliver a well-annotated corpus of essays written by second language learners of Swedish and make it available for research through a browsable environment. To that end, a new annotation tool and a new project management tool have been implemented, both with the main purpose to ensure reliability and quality of the final corpus. In the article we discuss reasoning behind metadata selection, principles of gold corpus compilation and argue for separation of normalization from correction annotation.

Keywords
Second-language learning
National Category
General Language Studies and Linguistics
Research subject
Computational Linguistics; Language Education
Identifiers
urn:nbn:se:su:diva-185228 (URN)10.3384/nejlt.2000-1533.19667 (DOI)
Funder
Riksbankens Jubileumsfond, N16-0464:1
Note

Special Issue of Selected Contributions from the Seventh Swedish Language Technology Conference (SLTC 2018)

Available from: 2020-09-18 Created: 2020-09-18 Last updated: 2023-09-08Bibliographically approved
Volodina, E., Granstedt, L., Megyesi, B., Prentice, J., Rosén, D., Schenström, C.-J., . . . Wirén, M. (2018). Annotation of learner corpora: first SweLL insights. In: Proceedings of 7th Workshop on NLP for Computer Assisted Language Learning at SLTC 2018: . Paper presented at The Seventh Swedish Language Technology Conference (SLTC), Stockholm, Sweden, November 7-9, 2018.
Open this publication in new window or tab >>Annotation of learner corpora: first SweLL insights
Show others...
2018 (English)In: Proceedings of 7th Workshop on NLP for Computer Assisted Language Learning at SLTC 2018, 2018Conference paper, Published paper (Refereed)
National Category
General Language Studies and Linguistics
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-195286 (URN)
Conference
The Seventh Swedish Language Technology Conference (SLTC), Stockholm, Sweden, November 7-9, 2018
Projects
Swedish Language Learner Corpora SweLL
Funder
Riksbankens Jubileumsfond, IN16-0464:1
Available from: 2021-08-11 Created: 2021-08-11 Last updated: 2023-09-08Bibliographically approved
Rosén, D., Wirén, M. & Volodina, E. (2018). Error Coding of Second-Language Learner Texts Based on Mostly Automatic Alignment of Parallel Corpora. In: Inguna Skadina, Maria Eskevich (Ed.), CLARIN Annual Conference 2018: Proceedings. Paper presented at CLARIN Annual Conference 2018, Pisa, Italy, 8–10 October, 2018 (pp. 181-184).
Open this publication in new window or tab >>Error Coding of Second-Language Learner Texts Based on Mostly Automatic Alignment of Parallel Corpora
2018 (English)In: CLARIN Annual Conference 2018: Proceedings / [ed] Inguna Skadina, Maria Eskevich, 2018, p. 181-184Conference paper, Published paper (Refereed)
Abstract [en]

Error coding of second-language learner text, that is, detecting, correcting and annotating errors, is a cumbersome task which in turn requires interpretation of the text to decide what the errors are. This paper describes a system with which the annotator corrects the learner text by editing it prior to the actual error annotation. During the editing, the system automatically generates a parallel corpus of the learner and corrected texts. Based on this, the work of the annotator consists of three independent tasks that are otherwise often conflated in error coding: correcting the learner text, repairing inconsistent alignments, and performing the actual error annotation.

Keywords
Second language learning, error coding, annotation, parallel corpora
National Category
General Language Studies and Linguistics Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-162705 (URN)
Conference
CLARIN Annual Conference 2018, Pisa, Italy, 8–10 October, 2018
Projects
SweLL project
Funder
Riksbankens Jubileumsfond, N16-0464:1
Available from: 2018-12-07 Created: 2018-12-07 Last updated: 2025-02-01Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0003-4040-3544

Search in DiVA

Show all publications