Change search
Link to record
Permanent link

Direct link
Publications (10 of 52) Show all publications
Sand Aronsson, F., Öijerstedt, L., Jelic, V., Östling, R., Östberg, P. & Graff, C. (2026). Grammatical complexity and productivity in written text in presymptomatic frontotemporal dementia: A repeated measures study. International Journal of Speech-Language Pathology
Open this publication in new window or tab >>Grammatical complexity and productivity in written text in presymptomatic frontotemporal dementia: A repeated measures study
Show others...
2026 (English)In: International Journal of Speech-Language Pathology, ISSN 1754-9507, E-ISSN 1754-9515Article in journal (Refereed) Epub ahead of print
Abstract [en]

Purpose: To assess if the writing performance of presymptomatic frontotemporal dementia mutation carriers differ from that of non-carrier controls in syntactic complexity and productivity.

Method: Using an automated analysis pipeline, average dependency distance and word count were calculated from written picture descriptions and sentences of 42 cohort participants from The Genetic Frontotemporal Dementia Initiative across three timepoints. Linear mixed models compared several fixed effects including mutation status with text measures over time.

Result: Genetic status had a positive effect on the number of words in a written picture description, with the presymptomatic mutation carriers producing more words than the non-carrier controls. A significant interaction between genetic status and age was also observed. No effect was found for the models using average dependency distance as the dependent variable.

Conclusion: We found limited evidence for differences in the written language between mutation carriers and non-carrier controls. As previous studies suggest, cognitive changes in genetic frontotemporal dementia manifest nearer to symptom onset and our participants were on average 11.8 years from expected onset, the anticipated changes might be undetectable at this stage.

Keywords
automated text analysis, average dependency distance, frontotemporal dementia, repeated measures, syntax, writing
National Category
Neurosciences Comparative Language Studies and Linguistics
Identifiers
urn:nbn:se:su:diva-252491 (URN)10.1080/17549507.2025.2602580 (DOI)001658780500001 ()2-s2.0-105028024094 (Scopus ID)
Available from: 2026-02-12 Created: 2026-02-12 Last updated: 2026-02-12
Ploeger, E., Bjerva, J., Tiedemann, J. & Östling, R. (2025). A Cross-Lingual Perspective on Neural Machine Translation Difficulty. In: Barry Haddow; Tom Kocmi; Philipp Koehn; Christof Monz (Ed.), The Tenth Conference on Machine Translation: Proceedings of the Conference. Paper presented at Tenth Conference on Machine Translation (WMT 2025), Suzhou, China, November 8-9, 2025 (pp. 340-354). Stroudsburg: Association for Computational Linguistics
Open this publication in new window or tab >>A Cross-Lingual Perspective on Neural Machine Translation Difficulty
2025 (English)In: The Tenth Conference on Machine Translation: Proceedings of the Conference / [ed] Barry Haddow; Tom Kocmi; Philipp Koehn; Christof Monz, Stroudsburg: Association for Computational Linguistics, 2025, p. 340-354Conference paper, Published paper (Refereed)
Abstract [en]

Intuitively, machine translation (MT) between closely related languages, such as Swedish and Danish, is easier than MT between more distant pairs, such as Finnish and Danish. Yet, the notions of ‘closely related’ languages and ‘easier’ translation have so far remained underspecified. Moreover, in the context of neural MT, this assumption was almost exclusively evaluated in scenarios where English was either the source or target language, leaving a broader cross-lingual view unexplored. In this work, we present a controlled study of language similarity and neural MT difficulty for 56 European translation directions. We test a range of language similarity metrics, some of which are reasonable predictors of MT difficulty. On a text-level, we reassess previously introduced indicators of MT difficulty, and find that they are not well-suited to our domain, or neural MT more generally. Ultimately, we hope that this work inspires further cross-lingual investigations of neural MT difficulty

Place, publisher, year, edition, pages
Stroudsburg: Association for Computational Linguistics, 2025
National Category
Natural Language Processing
Identifiers
urn:nbn:se:su:diva-252888 (URN)10.18653/v1/2025.wmt-1.21 (DOI)2-s2.0-105028902948 (Scopus ID)979-8-89176-341-8 (ISBN)
Conference
Tenth Conference on Machine Translation (WMT 2025), Suzhou, China, November 8-9, 2025
Available from: 2026-02-25 Created: 2026-02-25 Last updated: 2026-02-25Bibliographically approved
Kurfali, M. & Östling, R. (2025). Conflicting Needles in a Haystack: How LLMs Behave When Faced with Contradictory Information. In: Christos Christodoulopoulos; Tanmoy Chakraborty; Carolyn Rose; Violet Peng (Ed.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: . Paper presented at the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China (pp. 34361-34376). Association for Computational Linguistics
Open this publication in new window or tab >>Conflicting Needles in a Haystack: How LLMs Behave When Faced with Contradictory Information
2025 (English)In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing / [ed] Christos Christodoulopoulos; Tanmoy Chakraborty; Carolyn Rose; Violet Peng, Association for Computational Linguistics , 2025, p. 34361-34376Conference paper, Published paper (Refereed)
Abstract [en]

Large Language Models (LLMs) have demonstrated an impressive ability to retrieve and summarize complex information, but their reliability in conflicting contexts remains poorly understood. We introduce an adversarial extension of the Needle-in-a-Haystack framework in which three mutually exclusive “needles” are embedded within long documents. By systematically manipulating factors such as position, repetition, layout, and domain relevance, we evaluate how LLMs handle contradictions. We find that models almost always fail to signal uncertainty and instead confidently select a single answer, exhibiting strong and consistent biases toward repetition, recency, and particular surface forms. We further analyze whether these patterns persist across model families and sizes, and we evaluate both probability-based and generation-based retrieval. Our framework highlights critical limitations in the robustness of current LLMs-including commercial systems-to contradiction. These limitations reveal potential shortcomings in RAG systems' ability to handle noisy or manipulated inputs and exposes risks for deployment in high-stakes applications.

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2025
National Category
Natural Language Processing
Identifiers
urn:nbn:se:su:diva-257290 (URN)10.18653/v1/2025.emnlp-main.1742 (DOI)2-s2.0-105040272200 (Scopus ID)
Conference
the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China
Available from: 2026-06-25 Created: 2026-06-25 Last updated: 2026-06-25Bibliographically approved
Tudor, C. M., Megyesi, B. & Östling, R. (2025). Prompting the Past: Exploring Zero-Shot Learning for Named Entity Recognition in Historical Texts Using Prompt-Answering LLMs. In: Anna Kazantseva, Stan Szpakowicz, Stefania Degaetano-Ortlieb, Yuri Bizzoni, Janis Pagel (Ed.), Proceedings of the 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL2025): . Paper presented at The 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL2025). NAACL (pp. 216-226). Association for Computational Linguistics
Open this publication in new window or tab >>Prompting the Past: Exploring Zero-Shot Learning for Named Entity Recognition in Historical Texts Using Prompt-Answering LLMs
2025 (English)In: Proceedings of the 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL2025) / [ed] Anna Kazantseva, Stan Szpakowicz, Stefania Degaetano-Ortlieb, Yuri Bizzoni, Janis Pagel, Association for Computational Linguistics , 2025, p. 216-226Conference paper, Published paper (Refereed)
Abstract [en]

This paper investigates the application of prompt-answering Large Language Models (LLMs) for the task of Named Entity Recognition (NER) in historical texts. Historical NER presents unique challenges due to language change through time, spelling variation, limited availability of digitized data (and, in particular, labeled data), and errors introduced by Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) processes. Leveraging the zero-shot capabilities of prompt-answering LLMs, we address these challenges by prompting the model to extract entities such as persons, locations, organizations, and dates from historical documents. We then conduct an extensive error analysis of the model output in order to identify and address potential weaknesses in the entity recognition process. The results show that, while such models display ability for extracting named entities, their overall performance is lackluster. Our analysis reveals that model performance is significantly affected by hallucinations in the model output, as well as by challenges imposed by the evaluation of NER output.

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2025
National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-243266 (URN)979-8-89176-241-1 (ISBN)
Conference
The 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL2025). NAACL
Funder
Swedish Research Council, 2018-06074
Available from: 2025-05-20 Created: 2025-05-20 Last updated: 2025-05-23Bibliographically approved
Masciolini, A., Caines, A., de Clercq, O., Kruijsbergen, J., Kurfalı, M., Sánchez, R. M., . . . Östling, R. (2025). The MultiGEC-2025 Shared Task on Multilingual Grammatical Error Correction at NLP4CALL. In: Ricardo Muñoz Sánchez, David Alfter, Elena Volodina, Jelena Kallas (Ed.), Proceedings of the 14th Workshop on Natural Language Processing for Computer Assisted Language Learning (NLP4CALL 2025): . Paper presented at 14th Workshop on Natural Language Processing for Computer Assisted Language Learning (NLP4CALL 2025), Tallin, Estonia, March 5, 2025 (pp. 1-33). Tartu: University of Tartu
Open this publication in new window or tab >>The MultiGEC-2025 Shared Task on Multilingual Grammatical Error Correction at NLP4CALL
Show others...
2025 (English)In: Proceedings of the 14th Workshop on Natural Language Processing for Computer Assisted Language Learning (NLP4CALL 2025) / [ed] Ricardo Muñoz Sánchez, David Alfter, Elena Volodina, Jelena Kallas, Tartu: University of Tartu, 2025, p. 1-33Conference paper, Published paper (Refereed)
Abstract [en]

This paper reports on MultiGEC-2025, the first shared task in text-level Multilingual Grammatical Error Correction. The shared task features twelve European languages (Czech, English, Estonian, German, Greek, Icelandic, Italian, Latvian, Russian, Slovene, Swedish and Ukrainian) and is organized into two tracks, one for systems producing minimally corrected texts, thus preserving as much as possible of the original language use, and one dedicated to systems that prioritize fluency and idiomaticity. We introduce the task setup, data, evaluation metrics and baseline; present results obtained by the submitted systems and discuss key takeaways and ideas for future work.

Place, publisher, year, edition, pages
Tartu: University of Tartu, 2025
Keywords
grammatical error correction, GEC, natural language processing
National Category
Natural Language Processing Comparative Language Studies and Linguistics
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-243205 (URN)978-9908-53-112-0 (ISBN)
Conference
14th Workshop on Natural Language Processing for Computer Assisted Language Learning (NLP4CALL 2025), Tallin, Estonia, March 5, 2025
Funder
Swedish Research Council, 2019-04129
Available from: 2025-05-15 Created: 2025-05-15 Last updated: 2025-12-04Bibliographically approved
Masciolini, A., Caines, A., De Clercq, O., Kruijsbergen, J., Kurfalı, M., Sánchez, R. M., . . . Zesch, T. (2025). Towards better language representation in Natural Language Processing: A multilingual dataset for text-level Grammatical Error Correction. International Journal of Learner Corpus Research, 11(2), 309-335
Open this publication in new window or tab >>Towards better language representation in Natural Language Processing: A multilingual dataset for text-level Grammatical Error Correction
Show others...
2025 (English)In: International Journal of Learner Corpus Research, ISSN 2215-1478, E-ISSN 2215-1486, Vol. 11, no 2, p. 309-335Article in journal (Refereed) Published
Abstract [en]

This paper introduces MultiGEC, a dataset for multilingual Grammatical Error Correction (GEC) in twelve European languages: Czech, English, Estonian, German, Greek, Icelandic, Italian, Latvian, Russian, Slovene, Swedish and Ukrainian. MultiGEC distinguishes itself from previous GEC datasets in that it covers several underrepresented languages, which we argue should be included in resources used to train models for Natural Language Processing tasks which, as GEC itself, have implications for Learner Corpus Research and Second Language Acquisition. Aside from multilingualism, the novelty of the MultiGEC dataset is that it consists of full texts — typically learner essays — rather than individual sentences, making it possible to train systems that take a broader context into account. The dataset was built for MultiGEC-2025, the first shared task in multilingual text-level GEC, but it remains accessible after its competitive phase, serving as a resource to train new error correction systems and perform cross-lingual GEC studies.

Keywords
grammatical error correction, learner corpora, Matthew effect, MultiGEC shared task, multilingual corpora
National Category
Natural Language Processing
Identifiers
urn:nbn:se:su:diva-243066 (URN)10.1075/ijlcr.24033.mas (DOI)001457603500001 ()2-s2.0-105003035015 (Scopus ID)
Available from: 2025-05-09 Created: 2025-05-09 Last updated: 2025-09-22Bibliographically approved
Levshina, N., Koptjevskaja-Tamm, M. & Östling, R. (2024). Revered and reviled: a sentiment analysis of female and male referents in three languages. Frontiers in Communication, 9, Article ID 1266407.
Open this publication in new window or tab >>Revered and reviled: a sentiment analysis of female and male referents in three languages
2024 (English)In: Frontiers in Communication, E-ISSN 2297-900X, Vol. 9, article id 1266407Article in journal (Refereed) Published
Abstract [en]

Our study contributes to the less explored domain of lexical typology, focusing on semantic prosody and connotation. Semantic derogation, or pejoration of nouns referring to women, whereby such words acquire connotations and further denotations of social pejoration, immorality and/or loose sexuality, has been a very prominent question in studies on gender and language (change). It has been argued that pejoration emerges due to the general derogatory attitudes toward female referents. However, the evidence for systematic differences in connotations of female- vs. male-related words is fragmentary and often fairly impressionistic; moreover, many researchers argue that expressed sentiments toward women (as well as men) often are ambivalent. One should also expect gender differences in connotations to have decreased in the recent years, thanks to the advances of feminism and social progress. We test these ideas in a study of positive and negative connotations of feminine and masculine term pairs such as woman - man, girl - boy, wife - husband, etc. Sentences containing these words were sampled from diachronic corpora of English, Chinese and Russian, and sentiment scores for every word were obtained using two systems for Aspect-Based Sentiment Analysis: PyABSA, and OpenAI's large language model GPT-3.5. The Generalized Linear Mixed Models of our data provide no indications of significantly more negative sentiment toward female referents in comparison with their male counterparts. However, some of the models suggest that female referents are more infrequently associated with neutral sentiment than male ones. Neither do our data support the hypothesis of the diachronic convergence between the genders. In sum, results suggest that pejoration is unlikely to be explained simply by negative attitudes to female referents in general.

Keywords
semantic derogation, pejoration, sentiment analysis, diachronic corpora, semantic change, semantic prosody, gender stereotypes, prejudice
National Category
General Language Studies and Linguistics
Identifiers
urn:nbn:se:su:diva-228590 (URN)10.3389/fcomm.2024.1266407 (DOI)001199813900001 ()2-s2.0-85189960209 (Scopus ID)
Available from: 2024-04-23 Created: 2024-04-23 Last updated: 2024-04-23Bibliographically approved
Östling, R. & Kurfali, M. (2023). Language Embeddings Sometimes Contain Typological Generalizations. Computational linguistics - Association for Computational Linguistics (Print), 49(4), 1003-1051
Open this publication in new window or tab >>Language Embeddings Sometimes Contain Typological Generalizations
2023 (English)In: Computational linguistics - Association for Computational Linguistics (Print), ISSN 0891-2017, E-ISSN 1530-9312, Vol. 49, no 4, p. 1003-1051Article in journal (Refereed) Published
Abstract [en]

To what extent can neural network models learn generalizations about language structure, and how do we find out what they have learned? We explore these questions by training neural models for a range of natural language processing tasks on a massively multilingual dataset of Bible translations in 1,295 languages. The learned language representations are then compared to existing typological databases as well as to a novel set of quantitative syntactic and morphological features obtained through annotation projection. We conclude that some generalizations are surprisingly close to traditional features from linguistic typology, but that most of our models, as well as those of previous work, do not appear to have made linguistically meaningful generalizations. Careful attention to details in the evaluation turns out to be essential to avoid false positives. Furthermore, to encourage continued work in this field, we release several resources covering most or all of the languages in our data: (1) multiple sets of language representations, (2) multilingual word embeddings, (3) projected and predicted syntactic and morphological features, (4) software to provide linguistically sound evaluations of language representations.

Keywords
computational typology, language models, multilingual neural models, multilingual NLP, linguistic typology
National Category
General Language Studies and Linguistics Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-226219 (URN)10.1162/coli_a_00491 (DOI)001152974700001 ()2-s2.0-85175520799 (Scopus ID)
Funder
Swedish Research Council, 2019-04129
Available from: 2024-02-02 Created: 2024-02-02 Last updated: 2025-02-01Bibliographically approved
Kurfali, M. & Östling, R. (2021). Let’s be explicit about that: Distant supervision for implicit discourse relation classification via connective prediction. In: : . Paper presented at The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Bangkok, Thailand, August 1-6, 2021.
Open this publication in new window or tab >>Let’s be explicit about that: Distant supervision for implicit discourse relation classification via connective prediction
2021 (English)Conference paper, Oral presentation with published abstract (Refereed)
Abstract [en]

In implicit discourse relation classification, we want to predict the relation between adjacent sentences in the absence of any overt discourse connectives. This is challenging even for humans, leading to shortage of annotated data, a fact that makes the task even more difficult for supervised machine learning approaches. In the current study, we perform implicit discourse relation classification without relying on any labeled implicit relation. We sidestep the lack of data through explicitation of implicit relations to reduce the task to two sub-problems: language modeling and explicit discourse relation classification, a much easier problem. Our experimental results show that this method can even marginally outperform the state-of-the-art, in spite of being much simpler than alternative models of comparable performance. Moreover, we show that the achieved performance is robust across domains as suggested by the zero-shot experiments on a completely different domain. This indicates that recent advances in language modeling have made language models sufficiently good at capturing inter-sentence relations without the help of explicit discourse markers.

National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-201395 (URN)
Conference
The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Bangkok, Thailand, August 1-6, 2021
Available from: 2022-01-25 Created: 2022-01-25 Last updated: 2025-02-07Bibliographically approved
Kurfali, M. & Östling, R. (2021). Probing Multilingual Language Models for Discourse. In: : . Paper presented at The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Bangkok, Thailand, August 1-6, 2021.
Open this publication in new window or tab >>Probing Multilingual Language Models for Discourse
2021 (English)Conference paper, Oral presentation with published abstract (Refereed)
Abstract [en]

Pre-trained multilingual language models have become an important building block in multilingual natural language processing. In the present paper, we investigate a range of such models to find out how well they transfer discourse-level knowledge across languages. This is done with a systematic evaluation on a broader set of discourse-level tasks than has been previously been assembled. We find that the XLM-RoBERTa family of models consistently show the best performance, by simultaneously being good monolingual models and degrading relatively little in a zero-shot setting. Our results also indicate that model distillation may hurt the ability of cross-lingual transfer of sentence representations, while language dissimilarity at most has a modest effect. We hope that our test suite, covering 5 tasks with a total of 22 languages in 10 distinct families, will serve as a useful evaluation platform for multilingual performance at and beyond the sentence level. 

National Category
Natural Language Processing
Research subject
Computational Linguistics
Identifiers
urn:nbn:se:su:diva-201394 (URN)
Conference
The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Bangkok, Thailand, August 1-6, 2021
Available from: 2022-01-25 Created: 2022-01-25 Last updated: 2025-02-07Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-6027-4156

Search in DiVA

Show all publications