Change search
Link to record
Permanent link

Direct link
Sonnhammer, Erik L. L.ORCID iD iconorcid.org/0000-0002-9015-5588
Alternative names
Publications (10 of 104) Show all publications
Garbulowski, M., Mosca, R., Gallardo-Dodd, C. J., Kutter, C. & Sonnhammer, E. L. L. (2026). Comprehensive analysis of the RBP regulome reveals functional modules and drug candidates in liver cancer. Scientific Reports, 16, Article ID 19626.
Open this publication in new window or tab >>Comprehensive analysis of the RBP regulome reveals functional modules and drug candidates in liver cancer
Show others...
2026 (English)In: Scientific Reports, E-ISSN 2045-2322, Vol. 16, article id 19626Article in journal (Refereed) Published
Abstract [en]

RNA binding proteins (RBPs) are essential components of the transcriptomic regulome. Understanding the RBP regulome in cancer cells is crucial for uncovering carcinogenesis mechanisms and identifying novel therapeutic targets. In this study, we aimed to reveal the regulome of liver cancer upon specific perturbations. To achieve this, we applied a consensus Gene Regulatory Network (GRN) approach using knockdown data focusing on the liver cancer cell line HepG2. By integrating multiple GRNs inferred from diverse computational methods, we constructed a comprehensive regulatory network. To validate our findings, we evaluated the consensus GRN by focusing on characterizing key regulatory interactions in liver cancer. We used eCLIP-seq and RAP-seq data to verify RBP interactions and binding sites. In addition, we performed an enrichment analysis of network modules and in silico drug repurposing based on the inferred GRN. Taken together, our findings highlight the critical role of RBP-mediated regulation in liver cancer, which can be used to improve treatment strategies and develop further research.

National Category
Medical Genetics and Genomics Medical Bioinformatics and Systems Biology
Identifiers
urn:nbn:se:su:diva-258129 (URN)10.1038/s41598-026-58864-6 (DOI)001805737200004 ()42362666 (PubMedID)2-s2.0-105043107558 (Scopus ID)
Available from: 2026-08-20 Created: 2026-08-20 Last updated: 2026-08-20Bibliographically approved
Hillerton, T., Björk, A., Lundqvist, N., Zhivkoplias, E. K., Garbulowski, M. & Sonnhammer, E. L. L. (2026). GeneSNAKE: a Python package for simulation of gene regulatory networks and perturbation-induced expression data. Bioinformatics Advances, 6(1), Article ID vbag039.
Open this publication in new window or tab >>GeneSNAKE: a Python package for simulation of gene regulatory networks and perturbation-induced expression data
Show others...
2026 (English)In: Bioinformatics Advances, E-ISSN 2635-0041, Vol. 6, no 1, article id vbag039Article in journal (Refereed) Published
Abstract [en]

Motivation: Understanding how genes interact with and regulate each other is a key challenge in systems biology. One of the primary methods to study this is through gene regulatory networks (GRNs). The field of GRN inference faces many challenges, which necessitate effective tools for evaluating inference methods. Data that corresponds to a known GRN, from various conditions and experimental setups is necessary for this purpose, which is only possible to attain via simulation. However, most existing tools for GRN-based simulation are limited either in network or data properties, with few or no options to modify these properties.

Results: We present GeneSNAKE, a Python package designed to allow users to generate biologically realistic GRNs and expression data for benchmarking purposes. GeneSNAKE improves on previous work by providing a unique combination of modules, allowing users to control a wide range of GRN and data properties. It provides full control of the noise level, several noise models, full control of the perturbation design, and a wide range of pre-defined perturbation schemes. For benchmarking, GeneSNAKE offers several functions both for comparing network similarity, and properties in data and GRNs. These functions can further be used to study properties of biological data to produce simulated data with more realistic properties.

Availability and implementation: GeneSNAKE is an open-source, comprehensive simulation and benchmarking package with powerful capabilities that are not combined in any other single package. Thanks to the Python implementation, it can be extended and modified by users. The tool is available at: https://bitbucket.org/sonnhammergrni/genesnake/

Keywords
Gene regulatory networks, simulation, benchmarking, method development
National Category
Bioinformatics and Computational Biology
Research subject
Biochemistry towards Bioinformatics
Identifiers
urn:nbn:se:su:diva-221154 (URN)10.1093/bioadv/vbag039 (DOI)001754320500001 ()2-s2.0-105037843987 (Scopus ID)
Available from: 2023-09-14 Created: 2023-09-14 Last updated: 2026-06-11Bibliographically approved
Qin, H., Garbulowski, M., Sonnhammer, E. L. L. & Chatterjee, S. (2025). BiGSM: Bayesian inference of gene regulatory network via sparse modelling. Bioinformatics, 41(6), Article ID btaf318.
Open this publication in new window or tab >>BiGSM: Bayesian inference of gene regulatory network via sparse modelling
2025 (English)In: Bioinformatics, ISSN 1367-4803, E-ISSN 1367-4811, Vol. 41, no 6, article id btaf318Article in journal (Refereed) Published
Abstract [en]

Motivation Inference of gene regulatory network (GRN) is challenging due to the inherent sparsity of the GRN matrix and noisy expression data, often leading to a high possibility of false positive or negative predictions. To address this, it is essential to leverage the sparsity of the GRN matrix and develop a robust method capable of handling varying levels of noise in the data. Moreover, most existing GRN inference methods produce only fixed point estimates, which lack the flexibility and informativeness for comprehensive network analysis. In contrast, a Bayesian approach that yields closed-form posterior distributions allows probabilistic link selection, offering insights into the statistical confidence of each possible link. Consequently, it is important to engineer a Bayesian GRN inference method and rigorously execute a benchmark evaluation compared to state-of-the-art methods. Results We propose a method - Bayesian inference of GRN via Sparse Modelling (BiGSM). BiGSM effectively exploits the sparsity of the GRN matrix and infers the posterior distributions of GRN links from noisy expression data by using the maximum likelihood based learning. We thoroughly benchmarked BiGSM using biological and simulated datasets including GeneNetWeaver, GeneSPIDER, and GRNbenchmark. The benchmark test evaluates its accuracy and robustness across varying noise levels and data models. Using point-estimate based performance measures, BiGSM provides an overall best performance in comparison with several state-of-the-art methods including GENIE3, LASSO, LSCON, and Zscore. Additionally, BiGSM is the only method in the set of competing methods that provides posteriors for the GRN weights, helping to decipher confidence across predictions.

National Category
Bioinformatics and Computational Biology
Identifiers
urn:nbn:se:su:diva-245961 (URN)10.1093/bioinformatics/btaf318 (DOI)001505003500001 ()40484997 (PubMedID)2-s2.0-105008280617 (Scopus ID)
Available from: 2025-08-28 Created: 2025-08-28 Last updated: 2025-10-03Bibliographically approved
Majidian, S., Hadziahmetovic, A., Langschied, F., Pascarelli, S., Prieto-Baños, S., Rojas-Vargas, J., . . . Julca, I. (2025). Quest for Orthologs in the era of Data Deluge and AI: Challenges and Innovations in Orthology Prediction and Data Integration. Journal of Molecular Evolution, 93(6), 702-719
Open this publication in new window or tab >>Quest for Orthologs in the era of Data Deluge and AI: Challenges and Innovations in Orthology Prediction and Data Integration
Show others...
2025 (English)In: Journal of Molecular Evolution, ISSN 0022-2844, E-ISSN 1432-1432, Vol. 93, no 6, p. 702-719Article, review/survey (Refereed) Published
Abstract [en]

The rapid advancement of DNA sequencing technologies and computational algorithms has led to an unprecedented surge in genomic data, driven by several large-scale sequencing projects worldwide. Orthology plays a crucial role in understanding evolutionary patterns of genes and their functions. At the last Quest for Orthologs meeting (Montréal, Canada—2024), we discussed recent advances in orthology inference, with a focus on the impact of artificial intelligence (AI), protein structures, RNA splicing isoforms, and protein domain evolution together with other evolutionary considerations. A long-standing challenge in the field is the functional annotation of paralogs, for which we present novel approaches. The meeting also emphasised strategies for integrating diverse genetic features into the concept of orthology, encouraging frameworks that account for elements like alternative splicing, domain organisation, and regulatory sequences. We discuss various applications of orthology and paralogy to environmental research, agriculture, and comparative genomics. Additionally, we report recent progress in orthology inference methodologies and resources. This work represents a collaborative synthesis of insights and innovations presented at the 8th Quest for Orthologs meeting, highlighting current progress while outlining future directions for orthology research.

Keywords
Artificial intelligence, Gene function, Orthology, Paralogy, Protein domains
National Category
Genetics and Genomics Bioinformatics and Computational Biology
Identifiers
urn:nbn:se:su:diva-251236 (URN)10.1007/s00239-025-10272-6 (DOI)001592742500001 ()41085653 (PubMedID)2-s2.0-105023447396 (Scopus ID)
Available from: 2026-01-15 Created: 2026-01-15 Last updated: 2026-01-30Bibliographically approved
Buzzao, D., Steininger, L., Guala, D. & Sonnhammer, E. L. L. (2025). The FunCoup Cytoscape App: Multi-species network analysis and visualization. Bioinformatics, 41(1), Article ID btae739.
Open this publication in new window or tab >>The FunCoup Cytoscape App: Multi-species network analysis and visualization
2025 (English)In: Bioinformatics, ISSN 1367-4803, E-ISSN 1367-4811, Vol. 41, no 1, article id btae739Article in journal (Refereed) Published
Abstract [en]

Motivation: Functional association networks, such as FunCoup, are crucial for analyzing complex gene interactions. To facilitate the analysis and visualization of such genome-wide networks, there is a need for seamless integration with powerful network analysis tools like Cytoscape. Results: The FunCoup Cytoscape App integrates the FunCoup web service API with Cytoscape, allowing users to visualize and analyze gene interaction networks for 640 species. Users can input gene identifiers and customize search parameters, using various network expansion algorithms like group or independent gene search, MaxLink, and TOPAS. The app maintains consistent visualizations with the FunCoup website, providing detailed node and link information, including tissue and pathway gene annotations. The integration with Cytoscape plugins, such as ClusterMaker2, enhances the analytical capabilities of FunCoup, as exemplified by the identification of the Myasthenia gravis disease module along with potential new therapeutic targets.

National Category
Bioinformatics and Computational Biology
Identifiers
urn:nbn:se:su:diva-240400 (URN)10.1093/bioinformatics/btae739 (DOI)001388812400001 ()39700425 (PubMedID)2-s2.0-85214320388 (Scopus ID)
Available from: 2025-03-10 Created: 2025-03-10 Last updated: 2025-03-10Bibliographically approved
Paysan-Lafosse, T., Andreeva, A., Blum, M., Chuguransky, S. R., Grego, T., Pinto, B. L., . . . Bateman, A. (2025). The Pfam protein families database: Embracing AI/ML. Nucleic Acids Research, 53(D1), D523-D534
Open this publication in new window or tab >>The Pfam protein families database: Embracing AI/ML
Show others...
2025 (English)In: Nucleic Acids Research, ISSN 0305-1048, E-ISSN 1362-4962, Vol. 53, no D1, p. D523-D534Article in journal (Refereed) Published
Abstract [en]

The Pfam protein families database is a comprehensive collection of protein domains and families used for genome annotation and protein structure and function analysis (https://www.ebi.ac.uk/interpro/). This update describes major developments in Pfam since 2020, including decommissioning the Pfam website and integration with InterPro, harmonization with the ECOD structural classification, and expanded curation of metagenomic, microprotein and repeat-containing families. We highlight how AlphaFold structure predictions are being leveraged to refine domain boundaries and identify new domains. New families discovered through large-scale sequence similarity analysis of AlphaFold models are described. We also detail the development of Pfam-N, which uses deep learning to expand family coverage, achieving an 8.8% increase in UniProtKB coverage compared to standard Pfam. We discuss plans for more frequent Pfam releases integrated with InterPro and the potential for artificial intelligence to further assist curation. Despite recent advances, many protein families remain to be classified, and Pfam continues working toward comprehensive coverage of the protein universe.

National Category
Molecular Biology
Identifiers
urn:nbn:se:su:diva-240064 (URN)10.1093/nar/gkae997 (DOI)001354632400001 ()39540428 (PubMedID)2-s2.0-85214397377 (Scopus ID)
Available from: 2025-03-03 Created: 2025-03-03 Last updated: 2025-03-03Bibliographically approved
Lundqvist, N., Garbulowski, M., Hillerton, T. & Sonnhammer, E. L. L. (2025). Topology-based metrics for finding the optimal sparsity in gene regulatory network inference. Bioinformatics, 41(5), Article ID btaf120.
Open this publication in new window or tab >>Topology-based metrics for finding the optimal sparsity in gene regulatory network inference
2025 (English)In: Bioinformatics, ISSN 1367-4803, E-ISSN 1367-4811, Vol. 41, no 5, article id btaf120Article in journal (Refereed) Published
Abstract [en]

Motivation: Gene regulatory network (GRN) inference is a complex task aiming to unravel regulatory interactions between genes in a cell. A major shortcoming of most GRN inference methods is that they do not attempt to find the optimal sparsity, i.e. the single best GRN, which is important when applying GRN inference in a real situation. Instead, the sparsity tends to be controlled by an arbitrarily set hyperparameter. Results: In this paper, two new methods for predicting the optimal sparsity of GRNs are formulated and benchmarked on simulated perturbation-based gene expression data using four GRN inference methods: LASSO, Zscore, LSCON, and GENIE3. Both sparsity prediction methods are defined using the hypothesis that the topology of real GRNs is scale-free, and are evaluated based on their ability to predict the sparsity of the true GRN. The results show that the new topology-based approaches reliably predict a sparsity close to the true one. This ability is valuable for real-world applications where a single GRN is inferred from real data. In such situations, it is vital to be able to infer a GRN with the correct sparsity.

National Category
Biochemistry
Identifiers
urn:nbn:se:su:diva-243341 (URN)10.1093/bioinformatics/btaf120 (DOI)001483462800001 ()2-s2.0-105004690157 (Scopus ID)
Available from: 2025-05-22 Created: 2025-05-22 Last updated: 2025-05-22Bibliographically approved
Buzzao, D., Castresana-Aguirre, M., Guala, D. & Sonnhammer, E. L. L. (2024). Benchmarking enrichment analysis methods with the disease pathway network. Briefings in Bioinformatics, 25(2), Article ID bbae069.
Open this publication in new window or tab >>Benchmarking enrichment analysis methods with the disease pathway network
2024 (English)In: Briefings in Bioinformatics, ISSN 1467-5463, E-ISSN 1477-4054, Vol. 25, no 2, article id bbae069Article in journal (Refereed) Published
Abstract [en]

Enrichment analysis (EA) is a common approach to gain functional insights from genome-scale experiments. As a consequence, a large number of EA methods have been developed, yet it is unclear from previous studies which method is the best for a given dataset. The main issues with previous benchmarks include the complexity of correctly assigning true pathways to a test dataset, and lack of generality of the evaluation metrics, for which the rank of a single target pathway is commonly used. We here provide a generalized EA benchmark and apply it to the most widely used EA methods, representing all four categories of current approaches. The benchmark employs a new set of 82 curated gene expression datasets from DNA microarray and RNA-Seq experiments for 26 diseases, of which only 13 are cancers. In order to address the shortcomings of the single target pathway approach and to enhance the sensitivity evaluation, we present the Disease Pathway Network, in which related Kyoto Encyclopedia of Genes and Genomes pathways are linked. We introduce a novel approach to evaluate pathway EA by combining sensitivity and specificity to provide a balanced evaluation of EA methods. This approach identifies Network Enrichment Analysis methods as the overall top performers compared with overlap-based methods. By using randomized gene expression datasets, we explore the null hypothesis bias of each method, revealing that most of them produce skewed P-values.

Keywords
disease pathway network, functional enrichment, gene expression data, gene set enrichment analysis, pathway enrichment analysis, systems biology
National Category
Bioinformatics and Computational Biology
Identifiers
urn:nbn:se:su:diva-235218 (URN)10.1093/bib/bbae069 (DOI)001281650100007 ()2-s2.0-85186679428 (Scopus ID)
Funder
Swedish Research Council, 2022-06725Swedish Research Council, 2018-05973Swedish Research Council, 2019-04095Stockholm University
Available from: 2024-11-01 Created: 2024-11-01 Last updated: 2025-02-07Bibliographically approved
Garbulowski, M., Hillerton, T., Morgan, D., Seçilmiş, D., Sonnhammer, L., Tjärnberg, A., . . . Sonnhammer, E. L. L. (2024). GeneSPIDER2: large scale GRN simulation and benchmarking with perturbed single-cell data. NAR Genomics and Bioinformatics, 6(3), Article ID lqae121.
Open this publication in new window or tab >>GeneSPIDER2: large scale GRN simulation and benchmarking with perturbed single-cell data
Show others...
2024 (English)In: NAR Genomics and Bioinformatics, E-ISSN 2631-9268, Vol. 6, no 3, article id lqae121Article in journal (Refereed) Published
Abstract [en]

Single-cell data is increasingly used for gene regulatory network (GRN) inference, and benchmarks for this have been developed based on simulated data. However, existing single-cell simulators cannot model the effects of gene perturbations. A further challenge lies in generating large-scale GRNs that often struggle with computational and stability issues. We present GeneSPIDER2, an update of the GeneSPIDER MATLAB toolbox for GRN benchmarking, inference, and analysis. Several software modules have improved capabilities and performance, and new functionalities have been added. A major improvement is the ability to generate large GRNs with biologically realistic topological properties in terms of scale-free degree distribution and modularity. Another major addition is a simulation of single-cell data, which is becoming increasingly popular as input for GRN inference. Specifically, we introduced the unique feature to generate single-cell data based on genetic perturbations. Finally, the simulated single-cell data was compared to real single-cell Perturb-seq data from two cell lines, showing that the synthetic and real data exhibit similar properties.

National Category
Biochemistry Molecular Biology
Identifiers
urn:nbn:se:su:diva-237834 (URN)10.1093/nargab/lqae121 (DOI)001314667300003 ()2-s2.0-85204555831 (Scopus ID)
Available from: 2025-01-16 Created: 2025-01-16 Last updated: 2025-10-03Bibliographically approved
Altenhoff, A., Nevers, Y., Tran, V., Jyothi, D., Martin, M., Cosentino, S., . . . Sonnhammer, E. L. L. (2024). New developments for the Quest for Orthologs benchmark service. NAR Genomics and Bioinformatics, 6(4), Article ID lqae167.
Open this publication in new window or tab >>New developments for the Quest for Orthologs benchmark service
Show others...
2024 (English)In: NAR Genomics and Bioinformatics, E-ISSN 2631-9268, Vol. 6, no 4, article id lqae167Article in journal (Refereed) Published
Abstract [en]

The Quest for Orthologs (QfO) orthology benchmark service (https://orthology.benchmarkservice.org) hosts a wide range of standardized benchmarks for orthology inference evaluation. It is supported and maintained by the QfO consortium, and is used to gather ortholog predictions and to examine strengths and weaknesses of newly developed and existing orthology inference methods. The web server allows different inference methods to be compared in a standardized way using the same proteome data. The benchmark results are useful for developing new methods and can help researchers to guide their choice of orthology method for applications in comparative genomics and phylogenetic analysis. We here present a new release of the Orthology Benchmark Service with a new benchmark based on feature architecture similarity as well as updated reference proteomes. We further provide a meta-analysis of the public predictions from 18 different orthology assignment methods to reveal how they relate in terms of ortholog predictions and benchmark performance. These results can guide users of orthologs to the best suited method for their purpose.

National Category
Biochemistry
Identifiers
urn:nbn:se:su:diva-240704 (URN)10.1093/nargab/lqae167 (DOI)001374275400001 ()2-s2.0-85211996425 (Scopus ID)
Available from: 2025-03-14 Created: 2025-03-14 Last updated: 2025-03-14Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-9015-5588

Search in DiVA

Show all publications