Change search
Link to record
Permanent link

Direct link
Pavlopoulos, IoannisORCID iD iconorcid.org/0000-0001-9188-7425
Alternative names
Publications (10 of 22) Show all publications
Bakagianni, J., Pouli, K., Gavriilidou, M. & Pavlopoulos, J. (2025). A systematic survey of natural language processing for the Greek language. Patterns, 6(11), Article ID 101313.
Open this publication in new window or tab >>A systematic survey of natural language processing for the Greek language
2025 (English)In: Patterns, E-ISSN 2666-3899, Vol. 6, no 11, article id 101313Article in journal (Refereed) Published
Abstract [en]

Comprehensive monolingual natural language processing (NLP) surveys are essential for assessing language-specific challenges, resource availability, and research gaps. However, existing surveys often lack standardized methodologies, leading to selection bias and fragmented coverage of NLP tasks and resources. This study introduces a generalizable framework for systematic monolingual NLP surveys. Our approach integrates a structured search protocol to minimize bias, an NLP task taxonomy for classification, and language resource taxonomies to identify potential benchmarks and highlight opportunities for improving resource availability. We apply this framework to Greek NLP (2012–2023), providing an in-depth analysis of its current state, task-specific progress, and resource gaps. The survey results are publicly available and are regularly updated to provide an evergreen resource. This systematic survey of Greek NLP serves as a case study, demonstrating the effectiveness of our framework and its potential for broader application to other not-so-well-resourced languages as regards NLP.

Keywords
Greek NLP, language resources, monolingual NLP survey, search protocol, task taxonomy
National Category
Natural Language Processing
Identifiers
urn:nbn:se:su:diva-246282 (URN)10.1016/j.patter.2025.101313 (DOI)001621298100011 ()2-s2.0-105011258794 (Scopus ID)
Available from: 2025-09-03 Created: 2025-09-03 Last updated: 2026-04-15Bibliographically approved
Randl, K. R., Pavlopoulos, I., Henriksson, A. & Lindgren, T. (2025). Evaluating the Reliability of Self-Explanations in Large Language Models. In: Dino Pedreschi; Anna Monreale; Riccardo Guidotti; Roberto Pellungrini; Francesca Naretto (Ed.), Discovery Science: 27th International Conference, DS 2024, Pisa, Italy, October 14–16, 2024, Proceedings, Part I. Paper presented at Discovery Science, 27th International Conference, DS 2024, 14-16 October 2024, Pisa, Italy. (pp. 36-51). Springer Publishing Company
Open this publication in new window or tab >>Evaluating the Reliability of Self-Explanations in Large Language Models
2025 (English)In: Discovery Science: 27th International Conference, DS 2024, Pisa, Italy, October 14–16, 2024, Proceedings, Part I / [ed] Dino Pedreschi; Anna Monreale; Riccardo Guidotti; Roberto Pellungrini; Francesca Naretto, Springer Publishing Company , 2025, p. 36-51Conference paper, Published paper (Refereed)
Abstract [en]

This paper investigates the reliability of explanations generated by large language models~(LLMs) when prompted to explain their previous output. We evaluate two kinds of such self-explanations -- extractive and counterfactual -- using three state-of-the-art LLMs (2B to 8B parameters) on two different classification tasks (objective and subjective).

Our findings reveal, that, while these self-explanations can correlate with human judgement, they do not fully and accurately follow the model's decision process, indicating a gap between perceived and actual model reasoning.

We show that this gap can be bridged because prompting LLMs for counterfactual explanations can produce faithful, informative, and easy-to-verify results. These counterfactuals offer a promising alternative to traditional explainability methods (e.g. SHAP, LIME), provided that prompts are tailored to specific tasks and checked for validity.

Place, publisher, year, edition, pages
Springer Publishing Company, 2025
Series
Lecture Notes in Computer Science (LNCS), ISSN 0302-9743, E-ISSN 1611-3349 ; 15243
Keywords
Large Language Models, Self-Explanations, Counterfactuals
National Category
Computer Sciences
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-239126 (URN)10.1007/978-3-031-78977-9_3 (DOI)2-s2.0-85218499264 (Scopus ID)978-3-031-78976-2 (ISBN)978-3-031-78977-9 (ISBN)
Conference
Discovery Science, 27th International Conference, DS 2024, 14-16 October 2024, Pisa, Italy.
Available from: 2025-02-06 Created: 2025-02-06 Last updated: 2025-04-09Bibliographically approved
Randl, K. R., Pavlopoulos, I., Henriksson, A., Lindgren, T. & Bakagianni, J. (2025). SemEval-2025 Task 9: The Food Hazard Detection Challenge. In: Sara Rosenthal; Aiala Rosá; Debanjan Ghosh; Marcos Zampieri (Ed.), Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025): . Paper presented at The 19th International Workshop on Semantic Evaluation, July 2025, Vienna, Austria. (pp. 2523-2534). Association for Computational Linguistics
Open this publication in new window or tab >>SemEval-2025 Task 9: The Food Hazard Detection Challenge
Show others...
2025 (English)In: Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025) / [ed] Sara Rosenthal; Aiala Rosá; Debanjan Ghosh; Marcos Zampieri, Association for Computational Linguistics , 2025, p. 2523-2534Conference paper, Published paper (Refereed)
Abstract [en]

In this challenge, we explored text-based food hazard prediction with long tail distributed classes. The task was divided into two subtasks: (1) predicting whether a web text implies one of ten food-hazard categories and identifying the associated food category, and (2) providing a more fine-grained classification by assigning a specific label to both the hazard and the product. Our findings highlight that large language model-generated synthetic data can be highly effective for oversampling long-tail distributions. Furthermore, we find that fine-tuned encoder-only, encoder-decoder, and decoder-only systems achieve comparable maximum performance across both subtasks. During this challenge, we gradually released (under CC BY-NC-SA 4.0) a novel set of 6,644 manually labeled food-incident reports.

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2025
National Category
Computer Sciences
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-247395 (URN)979-8-89176-273-2 (ISBN)
Conference
The 19th International Workshop on Semantic Evaluation, July 2025, Vienna, Austria.
Available from: 2025-09-24 Created: 2025-09-24 Last updated: 2025-09-24Bibliographically approved
Pavlopoulos, I., Romell, A., Curman, J., Steinert, O., Lindgren, T., Borg, M. & Randl, K. (2024). Automotive fault nowcasting with machine learning and natural language processing. Machine Learning, 113(2), 843-861
Open this publication in new window or tab >>Automotive fault nowcasting with machine learning and natural language processing
Show others...
2024 (English)In: Machine Learning, ISSN 0885-6125, E-ISSN 1573-0565, Vol. 113, no 2, p. 843-861Article in journal (Refereed) Published
Abstract [en]

Automated fault diagnosis can facilitate diagnostics assistance, speedier troubleshooting, and better-organised logistics. Currently, most AI-based prognostics and health management in the automotive industry ignore textual descriptions of the experienced problems or symptoms. With this study, however, we propose an ML-assisted workflow for automotive fault nowcasting that improves on current industry standards. We show that a multilingual pre-trained Transformer model can effectively classify the textual symptom claims from a large company with vehicle fleets, despite the task’s challenging nature due to the 38 languages and 1357 classes involved. Overall, we report an accuracy of more than 80% for high-frequency classes and above 60% for classes with reasonable minimum support, bringing novel evidence that automotive troubleshooting management can benefit from multilingual symptom text classification.

Keywords
Automotive fault nowcasting, Natural language processing, Multilingual text classification
National Category
Computer Sciences
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-222622 (URN)10.1007/s10994-023-06398-7 (DOI)001075969400001 ()2-s2.0-85173121401 (Scopus ID)
Available from: 2023-10-13 Created: 2023-10-13 Last updated: 2024-02-22Bibliographically approved
Randl, K. R., Pavlopoulos, I., Henriksson, A. & Lindgren, T. (2024). CICLe: Conformal In-Context Learning for Largescale Multi-Class Food Risk Classification. In: Lun-Wei Ku; Andre Martins; Vivek Srikumar (Ed.), Findings of the Association for Computational Linguistics: ACL 2024. Paper presented at The 62nd Annual Meeting of the Association for Computational Linguistics, August 11-16 2024, Bangkok, Thailand. (pp. 7695-7715). Association for Computational Linguistics
Open this publication in new window or tab >>CICLe: Conformal In-Context Learning for Largescale Multi-Class Food Risk Classification
2024 (English)In: Findings of the Association for Computational Linguistics: ACL 2024 / [ed] Lun-Wei Ku; Andre Martins; Vivek Srikumar, Association for Computational Linguistics , 2024, p. 7695-7715Conference paper, Published paper (Refereed)
Abstract [en]

Contaminated or adulterated food poses a substantial risk to human health. Given sets of labeled web texts for training, Machine Learning and Natural Language Processing can be applied to automatically detect such risks. We publish a dataset of 7,546 short texts describing public food recall announcements. Each text is manually labeled, on two granularity levels (coarse and fine), for food products and hazards that the recall corresponds to. We describe the dataset and benchmark naive, traditional, and Transformer models. Based on our analysis, Logistic Regression based on a tf-idf representation outperforms RoBERTa and XLM-R on classes with low support. Finally, we discuss different prompting strategies and present an LLM-in-the-loop framework, based on Conformal Prediction, which boosts the performance of the base classifier while reducing energy consumption compared to normal prompting.

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2024
Keywords
In-Context-Learning, Prompting, Text Classification, Food-Risk, Conformal Prediction
National Category
Computer Sciences
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-237876 (URN)10.18653/v1/2024.findings-acl.459 (DOI)979-8-89176-099-8 (ISBN)
Conference
The 62nd Annual Meeting of the Association for Computational Linguistics, August 11-16 2024, Bangkok, Thailand.
Available from: 2025-01-14 Created: 2025-01-14 Last updated: 2025-01-15Bibliographically approved
Pavlopoulos, J., Konstantinidou, M., Perdiki, E., Marthot-Santaniello, I., Essler, H., Vardakas, G. & Likas, A. (2024). Explainable dating of greek papyri images. Machine Learning, 113(9), 6765-6786
Open this publication in new window or tab >>Explainable dating of greek papyri images
Show others...
2024 (English)In: Machine Learning, ISSN 0885-6125, E-ISSN 1573-0565, Vol. 113, no 9, p. 6765-6786Article in journal (Refereed) Published
Abstract [en]

Greek literary papyri, which are unique witnesses of antique literature, do not usually bear a date. They are thus currently dated based on palaeographical methods, with broad approximations which often span more than a century. We created a dataset of 242 images of papyri written in “bookhand” scripts whose date can be securely assigned, and we used it to train algorithms for the task of dating, showing its challenging nature. To address data scarcity, we extended our dataset by segmenting each image into its respective text lines. By using the line-based version of our dataset, we trained a Convolutional Neural Network, equipped with a fragmentation-based augmentation strategy, and we achieved a mean absolute error of 54 years. The results improve further when the task is cast as a multi-class classification problem, predicting the century. Using our network, we computed precise date estimations for papyri whose date is disputed or vaguely defined, employing explainability to understand dating-driving features.

Keywords
Attribution, Dating, Explainability, Papyri images
National Category
Natural Language Processing
Identifiers
urn:nbn:se:su:diva-237923 (URN)10.1007/s10994-024-06589-w (DOI)001269865600002 ()2-s2.0-85198093774 (Scopus ID)
Available from: 2025-01-14 Created: 2025-01-14 Last updated: 2025-01-14Bibliographically approved
Boulieris, P., Pavlopoulos, J., Xenos, A. & Vassalos, V. (2024). Fraud detection with natural language processing. Machine Learning, 113, 5087-5108
Open this publication in new window or tab >>Fraud detection with natural language processing
2024 (English)In: Machine Learning, ISSN 0885-6125, E-ISSN 1573-0565, Vol. 113, p. 5087-5108Article in journal (Refereed) Published
Abstract [en]

Automated fraud detection can assist organisations to safeguard user accounts, a task that is very challenging due to the great sparsity of known fraud transactions. Many approaches in the literature focus on credit card fraud and ignore the growing field of online banking. However, there is a lack of publicly available data for both. The lack of publicly available data hinders the progress of the field and limits the investigation of potential solutions. With this work, we: (a) introduce FraudNLP, the first anonymised, publicly available dataset for online fraud detection, (b) benchmark machine and deep learning methods with multiple evaluation measures, (c) argue that online actions do follow rules similar to natural language and hence can be approached successfully by natural language processing methods.

Keywords
Fraud detection, Natural language processing, E-banking, Feature engineering, Varying class imbalance
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:su:diva-221316 (URN)10.1007/s10994-023-06354-5 (DOI)001032046200001 ()2-s2.0-85165213807 (Scopus ID)
Available from: 2023-09-19 Created: 2023-09-19 Last updated: 2024-09-16Bibliographically approved
Pavlopoulos, I. & Konstantinidou, M. (2023). Computational authorship analysis of the homeric poems. International Journal of Digital Humanities, 5(1), 45-64
Open this publication in new window or tab >>Computational authorship analysis of the homeric poems
2023 (English)In: International Journal of Digital Humanities, ISSN 2524-7832, Vol. 5, no 1, p. 45-64Article in journal (Refereed) Published
Abstract [en]

Natural language modeling is used to predict or generate the next word or character of modern languages. Furthermore, statistical character-based language models have been found useful in authorship attribution analyses by studying the linguistic proximity of excerpts unknown to the model. In prior work, we modeled Homeric language and provided empirical findings regarding the authorship nature of the 48 Iliad and Odyssey books. Following this line of work, and considering the current philological views and trends, we break down the two poems further into smaller portions. By employing language modeling we identify outlying passages, indicating reduced linguistic affinity with the main body of the two works and, by extension, potentially different authorship. Our results show that some of the passages isolated as outliers by the language models were also identified as such by human researchers. We further test our methodology and models on texts of similar language and genre created by other authors, namely Hesiod’s “Theogony” and “Work and Days”.

Keywords
Natural language processing, Computational authorship analysis, Language modeling, Homeric language
National Category
Information Systems
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-213550 (URN)10.1007/s42803-022-00046-7 (DOI)
Available from: 2023-01-09 Created: 2023-01-09 Last updated: 2023-06-12Bibliographically approved
Pavlopoulos, I., Konstantinidou, M., Marthot-Santaniello, I., Essler, H. & Paparigopoulou, A. (2023). Dating Greek Papyri with Text Regression. In: Anna Rogers; Jordan Boyd-Graber; Naoaki Okazaki (Ed.), The 61st Conference of the the Association for Computational Linguistics: Proceedings of the Conference Volume 1: Long Papers. Paper presented at 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023), Toronto, Canada, July 9-14, 2023 (pp. 10001-10013). Stroudsburg: Association for Computational Linguistics
Open this publication in new window or tab >>Dating Greek Papyri with Text Regression
Show others...
2023 (English)In: The 61st Conference of the the Association for Computational Linguistics: Proceedings of the Conference Volume 1: Long Papers / [ed] Anna Rogers; Jordan Boyd-Graber; Naoaki Okazaki, Stroudsburg: Association for Computational Linguistics, 2023, p. 10001-10013Conference paper, Published paper (Refereed)
Abstract [en]

Dating Greek papyri accurately is crucial not only to edit their texts but also to understand numerous other aspects of ancient writing, document and book production and circulation, as well as various other aspects of administration, everyday life and intellectual history of antiquity. Although a substantial number of Greek papyri documents bear a date or other conclusive data as to their chronological placement, an even larger number can only be dated tentatively or in approximation, due to the lack of decisive evidence. By creating a dataset of 389 transcriptions of documentary Greek papyri, we train 389 regression models and we predict a date for the papyri with an average MAE of 54 years and an MSE of 1.17, outperforming image classifiers and other baselines. Last, we release date estimations for 159 manuscripts, for which only the upper limit is known.

Place, publisher, year, edition, pages
Stroudsburg: Association for Computational Linguistics, 2023
National Category
Computer Sciences
Identifiers
urn:nbn:se:su:diva-235288 (URN)10.18653/v1/2023.acl-long.556 (DOI)2-s2.0-85174419510 (Scopus ID)978-1-959429-72-2 (ISBN)
Conference
61st Annual Meeting of the Association for Computational Linguistics (ACL 2023), Toronto, Canada, July 9-14, 2023
Available from: 2024-11-07 Created: 2024-11-07 Last updated: 2024-11-07Bibliographically approved
Pavlopoulos, I. & Likas, A. (2023). Distance from Unimodality for the Assessment of Opinion Polarization. Cognitive Computation, 15(2), 731-738
Open this publication in new window or tab >>Distance from Unimodality for the Assessment of Opinion Polarization
2023 (English)In: Cognitive Computation, ISSN 1866-9956, E-ISSN 1866-9964, Vol. 15, no 2, p. 731-738Article in journal (Refereed) Published
Abstract [en]

Commonsense knowledge is often approximated by the fraction of annotators who classified an item as belonging to the positive class. Instances for which this fraction is equal to or above 50% are considered positive, including however ones that receive polarized opinions. This is a problematic encoding convention that disregards the potentially polarized nature of opinions and which is often employed to estimate subjectivity, sentiment polarity, and toxic language. We present the distance from unimodality (DFU), a novel measure that estimates the extent of polarization on a distribution of opinions and which correlates well with human judgment. We applied DFU to two use cases. The first case concerns tweets created over 9 months during the pandemic. The second case concerns textual posts crowd-annotated for toxicity. We specified the days for which the sentiment-annotated tweets were determined as polarized based on the DFU measure and we found that polarization occurred on different days for two different states in the USA. Regarding toxicity, we found that polarized opinions are more likely by annotators originating from different countries. Moreover, we show that DFU can be exploited as an objective function to train models to predict whether a post will provoke polarized opinions in the future.

National Category
Information Systems
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-213551 (URN)10.1007/s12559-022-10088-2 (DOI)000905531900002 ()2-s2.0-85145030762 (Scopus ID)
Available from: 2023-01-09 Created: 2023-01-09 Last updated: 2023-10-09Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0001-9188-7425

Search in DiVA

Show all publications