Change search
Link to record
Permanent link

Direct link
Alternative names
Publications (10 of 43) Show all publications
Kreutzer, Y., Rahu, I., Norinder, U. & Kruve, A. (2026). Molecular networking, conformal predictions and revised fingerprint-based models for discovering endocrine disruptors in mixtures. Analytical and Bioanalytical Chemistry, 418, 1445-1457
Open this publication in new window or tab >>Molecular networking, conformal predictions and revised fingerprint-based models for discovering endocrine disruptors in mixtures
2026 (English)In: Analytical and Bioanalytical Chemistry, ISSN 1618-2642, E-ISSN 1618-2650, Vol. 418, p. 1445-1457Article in journal (Refereed) Published
Abstract [en]

Prioritizing high-risk features is a key step to reduce workload in non-targeted screening (NTS) when identifying environmental contaminants. Machine learning models from the MS2Tox toolbox have shown promise for feature prioritization, but rely heavily on the accuracy of molecular formulas and fingerprint features provided by SIRIUS + CSI:FingerID. In this study, we introduce and evaluate two new approaches—molecular networking (MN) and conformal predictions—to discover unidentified compounds potentially posing endocrine-disrupting activity based on tandem mass spectral similarity. Furthermore, we revised the previously published MS2Tox models, leveraging molecular fingerprints for seven Tox21 Data Challenge endpoints. The fingerprint-based MS2Tox models achieved the lowest false positive rate, 0.35, at 90% recall on the test set, while MN and CP yielded 0.82 and 0.68, respectively. In a case study of transformation products and persistent chemicals in wastewater, these three approaches prioritized 29 features in influent and effluent samples as potentially associated with AhR agonism among 189 LC/HRMS features corresponding to transformation products and persistent chemicals. All candidate structures for prioritized features showed scaffolds related to AhR binding affinity. Three features were identified on level 1, showcasing potential in using combined feature prioritization strategies.

Keywords
Hazard, High-resolution mass spectrometry, Machine learning, Toxicity, Untargeted screening
National Category
Analytical Chemistry
Identifiers
urn:nbn:se:su:diva-252403 (URN)10.1007/s00216-025-06303-2 (DOI)001666416500001 ()41566047 (PubMedID)2-s2.0-105028184634 (Scopus ID)
Available from: 2026-02-11 Created: 2026-02-11 Last updated: 2026-03-26Bibliographically approved
Margarita, C., Pierozan, P., Subramaniyan, S., Shatskiy, A., Pakarinen, D., Fritz, A., . . . Lundberg, H. (2026). Safe-and-sustainable-by-design approach to polyesters from non-oestrogenic bisphenols. Nature Sustainability, 9, 86-95
Open this publication in new window or tab >>Safe-and-sustainable-by-design approach to polyesters from non-oestrogenic bisphenols
Show others...
2026 (English)In: Nature Sustainability, E-ISSN 2398-9629, Vol. 9, p. 86-95Article in journal (Refereed) Published
Abstract [en]

Most contemporary chemical processes rely on non-renewable resources and reagents associated with negative impact on environment and human health. As a result, the safe-and-sustainable-by-design (SSbD) framework is launched to guide the innovation towards safe and sustainable materials and chemical products. Bisphenol A (BPA) is a widely used chemical in the production of plastics but known to activate oestrogen receptors and linked by numerous studies to adverse effects on both human health and the environment. Here we demonstrate how SSbD can lead a multidisciplinary study for the identification of non-oestrogenic BPA analogues suitable for incorporation into high-performance polymeric materials. Toxicological evaluation of a library of 172 bisphenols using an in silico model identified 20 promising candidates that are synthesized from renewable lignin-sourced feedstocks via benign dehydrative catalytic routes. Subsequent in vitro assessment of their oestrogen receptor activity identifies bisguaiacol F as optimal BPA analogue, which is incorporated into a polyester with attractive thermal stability and flexibility. This work demonstrates an effective workflow for the discovery of renewable and non-oestrogenic bisphenols by taking advantage of the synergy of synthetic chemistry, toxicology and computational modelling.

National Category
Other Chemistry Topics Pharmacology and Toxicology
Identifiers
urn:nbn:se:su:diva-251214 (URN)10.1038/s41893-025-01672-z (DOI)001630545000001 ()2-s2.0-105024011066 (Scopus ID)
Available from: 2026-01-16 Created: 2026-01-16 Last updated: 2026-03-25Bibliographically approved
Norinder, U., Zheng, Z. & Cotgreave, I. (2025). Prediction of the classification, labelling and packaging regulation H-statements with confidence using conformal prediction with N-grams and molecular fingerprints. Current Research in Toxicology, 8, Article ID 100242.
Open this publication in new window or tab >>Prediction of the classification, labelling and packaging regulation H-statements with confidence using conformal prediction with N-grams and molecular fingerprints
2025 (English)In: Current Research in Toxicology, E-ISSN 2666-027X, Vol. 8, article id 100242Article in journal (Refereed) Published
Abstract [en]

Effective chemical hazard labelling systems are essential for safeguarding human health and the environment as a result of widespread chemical use, and machine-learning models can be used to predict hazard labels efficiently and reduce the use of animal tests. This investigation shows the utility of N-grams and other fingerprint featurization procedures for predicting classification, labelling and packaging (CLP). Regulation H-statements, particularly in an ensemble (consensus) setting. Consensus modelling by class or Conformal Prediction median p-values seems to be particularly advantageous in order to obtain both high conformal prediction validity and efficiency as well as good balanced accuracy, sensitivity and specificity. Utilization of the N-grams allows handling of all symbols in SMILES strings including those related to metals and salts that may be important for the compounds to exhibit their experimental determined toxicities. The models developed in this study are efficient tools to access hazard classification H-statements of chemicals, which can be useful for chemical hazard assessment, read-across as well as risk management.

Keywords
CLP Regulation, Conformal prediction, Consensus modeling, H-statements, Molecular fingerprints, N-grams, Random forest
National Category
Artificial Intelligence
Identifiers
urn:nbn:se:su:diva-244164 (URN)10.1016/j.crtox.2025.100242 (DOI)001503552100002 ()2-s2.0-105006900609 (Scopus ID)
Available from: 2025-06-16 Created: 2025-06-16 Last updated: 2025-06-16Bibliographically approved
Geylan, G., De Maria, L., Engkvist, O., David, F. & Norinder, U. (2024). A methodology to correctly assess the applicability domain of cell membrane permeability predictors for cyclic peptides. Digital Discovery, 3(9), 1761-1775
Open this publication in new window or tab >>A methodology to correctly assess the applicability domain of cell membrane permeability predictors for cyclic peptides
Show others...
2024 (English)In: Digital Discovery, E-ISSN 2635-098X, Vol. 3, no 9, p. 1761-1775Article in journal (Refereed) Published
Abstract [en]

Being able to predict the cell permeability of cyclic peptides is essential for unlocking their potential as a drug modality for intracellular targets. With a wide range of studies of cell permeability but a limited number of data points, the reliability of the machine learning (ML) models to predict previously unexplored chemical spaces becomes a challenge. In this work, we systemically investigate the predictive capability of ML models from the perspective of their extrapolation to never-before-seen applicability domains, with a particular focus on the permeability task. Four predictive algorithms, namely Support-Vector Machine, Random Forest, LightGBM and XGBoost, jointly with a conformal prediction framework were employed to characterize and evaluate the applicability through uncertainty quantification. Efficiency and validity of the models' predictions with multiple calibration strategies were assessed with respect to several external datasets from different parts of the chemical space through a set of experiments. The experiments showed that the predictors generalizing well to the applicability domain defined by the training data, can fail to achieve similar model performance on other parts of the chemical spaces. Our study proposes an approach to overcome such limitations by the means of improving the efficiency of models without sacrificing the validity. The trade-off between the reliability and informativeness was balanced when the models were calibrated with a subset of the data from the new targeted domain. This study outlines an approach to enable the extrapolation of predictive power and restore the models' reliability via a recalibration strategy without the need for retraining the underlying model.

National Category
Bioinformatics (Computational Biology)
Identifiers
urn:nbn:se:su:diva-238174 (URN)10.1039/d4dd00056k (DOI)001279737000001 ()2-s2.0-85200371105 (Scopus ID)
Available from: 2025-01-22 Created: 2025-01-22 Last updated: 2025-01-22Bibliographically approved
Arvidsson McShane, S., Norinder, U., Alvarsson, J., Ahlberg, E., Carlsson, L. & Spjuth, O. (2024). CPSign: conformal prediction for cheminformatics modeling. Journal of Cheminformatics, 16, Article ID 75.
Open this publication in new window or tab >>CPSign: conformal prediction for cheminformatics modeling
Show others...
2024 (English)In: Journal of Cheminformatics, E-ISSN 1758-2946, Vol. 16, article id 75Article in journal (Refereed) Published
Abstract [en]

Conformal prediction has seen many applications in pharmaceutical science, being able to calibrate outputs of machine learning models and producing valid prediction intervals. We here present the open source software CPSign that is a complete implementation of conformal prediction for cheminformatics modeling. CPSign implements inductive and transductive conformal prediction for classification and regression, and probabilistic prediction with the Venn-ABERS methodology. The main chemical representation is signatures but other types of descriptors are also supported. The main modeling methodology is support vector machines (SVMs), but additional modeling methods are supported via an extension mechanism, e.g. DeepLearning4J models. We also describe features for visualizing results from conformal models including calibration and efficiency plots, as well as features to publish predictive models as REST services. We compare CPSign against other common cheminformatics modeling approaches including random forest, and a directed message-passing neural network. The results show that CPSign produces robust predictive performance with comparative predictive efficiency, with superior runtime and lower hardware requirements compared to neural network based models. CPSign has been used in several studies and is in production-use in multiple organizations. The ability to work directly with chemical input files, perform descriptor calculation and modeling with SVM in the conformal prediction framework, with a single software package having a low footprint and fast execution time makes CPSign a convenient and yet flexible package for training, deploying, and predicting on chemical data. CPSign can be downloaded from GitHub at https://github.com/arosbio/cpsign.

National Category
Computer Sciences Other Chemistry Topics
Identifiers
urn:nbn:se:su:diva-237024 (URN)10.1186/s13321-024-00870-9 (DOI)001258657400001 ()2-s2.0-85197657994 (Scopus ID)
Available from: 2024-12-12 Created: 2024-12-12 Last updated: 2024-12-12Bibliographically approved
Söderberg, E., von Borries, K., Norinder, U., Petchey, M., Ranjani, G., Chavan, S., . . . Syrén, P.-O. (2024). Toward safer and more sustainable by design biocatalytic amide-bond coupling. Green Chemistry, 26(22), 11147-11163
Open this publication in new window or tab >>Toward safer and more sustainable by design biocatalytic amide-bond coupling
Show others...
2024 (English)In: Green Chemistry, ISSN 1463-9262, E-ISSN 1463-9270, Vol. 26, no 22, p. 11147-11163Article in journal (Refereed) Published
Abstract [en]

Amide bond synthesis is ranked as the second most important challenge in key green chemistry research areas identified by the ACS Green Chemistry Institute. While developing more sustainable amide bond forming reactions has been in focus, significantly less attention has been given to human toxicity and environmental aspects of the underlying amine and acid substrates and their corresponding coupled products, a potentially important contribution to the overall sustainability of the amide-bond-forming reactions. Here, we explore biocatalytic amide bond formation from a safer-and-more-sustainable-by-design perspective in which commercially available amines and acids as well as their corresponding amide products were evaluated in silico based on potential human toxicity and environmental fate and exposure. This in silico filtering resulted in a panel of 188 amine and 54 acid building blocks that could be classified as safe, referred to herein as “safechems”. To enable couplings of safechems, we generated a panel of robust and promiscuous ancestral ATP-dependent amide bond synthetases (ABS) using McbA from Marinactinospora thermotolerans SCSIO 00652 as a template. Ancestral ABS enzymes exhibited complementary specificities in the coupling of a representative safechem subset of 17 amines and 16 acids while showing an increased thermostability of up to 20 °C compared to the extant biocatalyst. Finally, the pool of safechems and their corresponding amides were evaluated by USEtox (the UNEP-SETAC toxicity model), analysing not only the intrinsic properties of the compounds but evaluating their complete impact pathway including fate, exposure and effects. The amides were in general predicted as more toxic compared to the starting acids and amines through non-additive effects, emphasising that focusing on the toxicity of the building blocks alone is not sufficient to strive towards low human and ecotoxicity impact. Pursuing a safer and more sustainable by design perspective in the implementation of safechems did not prevent us from generating an array of novel products with potentially potent applications as exemplified here by enzymatic synthesis of substructures that are part of drug candidates for e.g. cancer treatment.

National Category
Organic Chemistry Environmental Sciences
Identifiers
urn:nbn:se:su:diva-238902 (URN)10.1039/d4gc03665d (DOI)001329141300001 ()2-s2.0-85206544471 (Scopus ID)
Available from: 2025-02-03 Created: 2025-02-03 Last updated: 2025-02-03Bibliographically approved
Ylipää, E., Chavan, S., Bånkestad, M., Broberg, J., Glinghammar, B., Norinder, U. & Cotgreave, I. (2023). hERG-toxicity prediction using traditional machine learning and advanced deep learning techniques. Current Research in Toxicology, 5, Article ID 100121.
Open this publication in new window or tab >>hERG-toxicity prediction using traditional machine learning and advanced deep learning techniques
Show others...
2023 (English)In: Current Research in Toxicology, E-ISSN 2666-027X, Vol. 5, article id 100121Article in journal (Refereed) Published
Abstract [en]

The rise of artificial intelligence (AI) based algorithms has gained a lot of interest in the pharmaceutical development field. Our study demonstrates utilization of traditional machine learning techniques such as random forest (RF), support-vector machine (SVM), extreme gradient boosting (XGBoost), deep neural network (DNN) as well as advanced deep learning techniques like gated recurrent unit-based DNN (GRU-DNN) and graph neural network (GNN), towards predicting human ether-á-go-go related gene (hERG) derived toxicity. Using the largest hERG dataset derived to date, we have utilized 203,853 and 87,366 compounds for training and testing the models, respectively. The results show that GNN, SVM, XGBoost, DNN, RF, and GRU-DNN all performed well, with validation set AUC ROC scores equals 0.96, 0.95, 0.95, 0.94, 0.94 and 0.94, respectively. The GNN was found to be the top performing model based on predictive power and generalizability. The GNN technique is free of any feature engineering steps while having a minimal human intervention. The GNN approach may serve as a basis for comprehensive automation in predictive toxicology. We believe that the models presented here may serve as a promising tool, both for academic institutes as well as pharmaceutical industries, in predicting hERG-liability in new molecular structures.

Keywords
Deep Learning, Graph -neural Network, hERG Channel, Random Forest, Recurrent -neural Network, Support -vector Machines
National Category
Computer Sciences
Identifiers
urn:nbn:se:su:diva-223947 (URN)10.1016/j.crtox.2023.100121 (DOI)001074877800001 ()37701072 (PubMedID)2-s2.0-85171525959 (Scopus ID)
Available from: 2023-11-27 Created: 2023-11-27 Last updated: 2023-11-29Bibliographically approved
Andersson, M., Norinder, U., Chavan, S. & Cotgreave, I. (2023). In Silico Prediction of Eye Irritation Using Hansen Solubility Parameters and Predicted pKa Values. ATLA (Alternatives to Laboratory Animals), 51(3), 204-209
Open this publication in new window or tab >>In Silico Prediction of Eye Irritation Using Hansen Solubility Parameters and Predicted pKa Values
2023 (English)In: ATLA (Alternatives to Laboratory Animals), ISSN 0261-1929, Vol. 51, no 3, p. 204-209Article in journal (Refereed) Published
Abstract [en]

An in silico method has been developed that permits the binary differentiation between pure liquids causing serious eye damage or eye irritation, and pure liquids with no need for such classification, according to the UN GHS system. The method is based on the finding that the Hansen Solubility Parameters (HSP) of a liquid are collectively important predictors for eye irritation. Thus, by applying a two-tier approach in which in silico-predicted pKa values (firstly) and a trained model based solely on in silico-predicted HSP data (secondly) were used, we have developed, and validated, a fully in silico approach for predicting the outcome of a Draize test (in terms of UN GHS Cat. 1/Cat. 2A/Cat. 2B or UN GHS No Cat.) with high validation set performance (sensitivity = 0.846, specificity = 0.818, balanced accuracy = 0.832) using SMILES only. The method is applicable to pure non-ionic liquids with molecular weight below 500 g/mol, fewer than six hydrogen bond donors (e.g. nitrogen–hydrogen or oxygen–hydrogen bonds) and fewer than eleven hydrogen bond acceptors (e.g. nitrogen or oxygen atoms). Due to its fully in silico characteristics, this method can be applied to pure liquids that are still at the desktop design stage and not yet in production.

Keywords
computational toxicology, eye irritation, genetic algorithm optimisation, Hansen Solubility Parameters, in silico prediction
National Category
Pharmacology and Toxicology
Identifiers
urn:nbn:se:su:diva-229956 (URN)10.1177/02611929231175676 (DOI)001012108200005 ()37184299 (PubMedID)2-s2.0-85159096294 (Scopus ID)
Available from: 2024-06-03 Created: 2024-06-03 Last updated: 2024-06-03Bibliographically approved
Sapounidou, M., Norinder, U. & Andersson, P. L. (2023). Predicting Endocrine Disruption Using Conformal Prediction – A Prioritization Strategy to Identify Hazardous Chemicals with Confidence. Chemical Research in Toxicology, 36(1), 53-65
Open this publication in new window or tab >>Predicting Endocrine Disruption Using Conformal Prediction – A Prioritization Strategy to Identify Hazardous Chemicals with Confidence
2023 (English)In: Chemical Research in Toxicology, ISSN 0893-228X, E-ISSN 1520-5010, Vol. 36, no 1, p. 53-65Article in journal (Refereed) Published
Abstract [en]

Receptor-mediated molecular initiating events (MIEs) and their relevance in endocrine activity (EA) have been highlighted in literature. More than 15 receptors have been associated with neurodevelopmental adversity and metabolic disruption. MIEs describe chemical interactions with defined biological outcomes, a relationship that could be described with quantitative structure–activity relationship (QSAR) models. QSAR uncertainty can be assessed using the conformal prediction (CP) framework, which provides similarity (i.e., nonconformity) scores relative to the defined classes per prediction. CP calibration can indirectly mitigate data imbalance during model development, and the nonconformity scores serve as intrinsic measures of chemical applicability domain assessment during screening. The focus of this work was to propose an in silico predictive strategy for EA. First, 23 QSAR models for MIEs associated with EA were developed using high-throughput data for 14 receptors. To handle the data imbalance, five protocols were compared, and CP provided the most balanced class definition. Second, the developed QSAR models were applied to a large data set (∼55,000 chemicals), comprising chemicals representative of potential risk for human exposure. Using CP, it was possible to assess the uncertainty of the screening results and identify model strengths and out of domain chemicals. Last, two clustering methods, t-distributed stochastic neighbor embedding and Tanimoto similarity, were used to identify compounds with potential EA using known endocrine disruptors as reference. The cluster overlap between methods produced 23 chemicals with suspected or demonstrated EA potential. The presented models could be utilized for first-tier screening and identification of compounds with potential biological activity across the studied MIEs. 

National Category
Pharmacology and Toxicology
Identifiers
urn:nbn:se:su:diva-213845 (URN)10.1021/acs.chemrestox.2c00267 (DOI)000903383200001 ()36534483 (PubMedID)2-s2.0-85144410434 (Scopus ID)
Available from: 2023-01-18 Created: 2023-01-18 Last updated: 2023-01-18Bibliographically approved
Norinder, U. & Lowry, S. (2023). Predicting Larch Casebearer damage with confidence using Yolo network models and conformal prediction. Remote Sensing Letters, 14(10), 1023-1035
Open this publication in new window or tab >>Predicting Larch Casebearer damage with confidence using Yolo network models and conformal prediction
2023 (English)In: Remote Sensing Letters, ISSN 2150-704X, E-ISSN 2150-7058, Vol. 14, no 10, p. 1023-1035Article in journal (Refereed) Published
Abstract [en]

This investigation shows that successful forecasting models for monitoring forest health status with respect to Larch Casebearer damages can be derived using a combination of a confidence predictor framework (Conformal Prediction) in combination with a deep learning architecture (Yolo v5). A confidence predictor framework can predict the current types of diseases used to develop the model and also provide indication of new, unseen, types or degrees of disease. The user of the models is also, at the same time, provided with reliable predictions and a well-established applicability domain for the model where such reliable predictions can and cannot be expected. Furthermore, the framework gracefully handles class imbalances without explicit over- or under-sampling or category weighting which may be of crucial importance in cases of highly imbalanced datasets. The present approach also provides indication of when insufficient information has been provided as input to the model at the level of accuracy (reliability) need by the user to make subsequent decisions based on the model predictions. 

Keywords
Yolo network, Larch Casebearer moth, conformal prediction, forest health, tree damage
National Category
Earth Observation
Identifiers
urn:nbn:se:su:diva-222259 (URN)10.1080/2150704X.2023.2258460 (DOI)001071044000001 ()2-s2.0-85171885925 (Scopus ID)
Available from: 2023-10-12 Created: 2023-10-12 Last updated: 2025-02-10Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0003-3107-331x

Search in DiVA

Show all publications