Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
De novo clustering of large long-read transcriptome datasets with isONclust3
Stockholm University, Faculty of Science, Department of Mathematics. Stockholm University, Science for Life Laboratory (SciLifeLab). University of Helsinki, Finland.ORCID iD: 0009-0005-9397-0341
Stockholm University, Faculty of Science, Department of Mathematics. Stockholm University, Science for Life Laboratory (SciLifeLab).ORCID iD: 0000-0001-7378-2320
2025 (English)In: Bioinformatics, ISSN 1367-4803, E-ISSN 1367-4811, Vol. 41, no 5, article id btaf207Article in journal (Refereed) Published
Abstract [en]

Motivation

Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies.

Results

We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10–100 times faster on large datasets. Also, using a 256 Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine.

Place, publisher, year, edition, pages
2025. Vol. 41, no 5, article id btaf207
National Category
Bioinformatics (Computational Biology)
Identifiers
URN: urn:nbn:se:su:diva-243269DOI: 10.1093/bioinformatics/btaf207ISI: 001483472300001PubMedID: 40265453Scopus ID: 2-s2.0-105004673060OAI: oai:DiVA.org:su-243269DiVA, id: diva2:1959525
Funder
Swedish Research Council, 2021–04000Available from: 2025-05-20 Created: 2025-05-20 Last updated: 2025-06-02Bibliographically approved
In thesis
1. Computational methods for long-read sequencing data analysis
Open this publication in new window or tab >>Computational methods for long-read sequencing data analysis
2025 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

This thesis presents algorithms developed for long-read sequencing techniques, which, since their introduction in the 2010’s become increasingly important approaches in modern bioscientific research. The first two papers cover the development of algorithms for our de novo transcriptome prediction pipeline, the isON pipeline, while the third paper describes an algorithm used for biotechnological analysis of ligated fragments. Paper I introduces isONform, an algorithm capable of predicting different gene products, called isoforms, from a set of long reads sequenced from complementary DNA without the need to rely on a reference or annotation. IsONform is a tool that is part of a larger long-read transcriptome pipeline, isON pipeline, that consists of clustering and error correction steps prior to the isoform prediction. The isONform algorithm is based on the construction of a directed acyclic graph with minimizer-pairs as nodes and connecting neighboring minimizer-pairs on the reads with edges. The algorithm then employs an iterative bubble-popping scheme to merge nodes to ultimately follow all distinct paths through the graph generating the final isoform predictions. The algorithm has been shown to outperform existing state-of-the-art algorithms, while showing comparable results to approaches requiring information of a reference genome and an annotation. Paper II introduces isONclust3, an algorithm used for clustering transcriptomic reads by gene family. The algorithm constitutes the first step employed in pipelines for reference-free prediction of isoforms. The algorithm is based on the minimizer indexing scheme with its novelties being a dynamic clustering approach, assessing and storing minimizers by confidence, and an iterative post-cluster merging step. The algorithm has been shown to scale better, in terms of runtime and memory usage, on large datasets than existing methods while yielding comparable or even better results with respect to clustering quality assessments. We demonstrate that isONclust3 is the only algorithm that can process the clustering of PacBio’s new Revio datasets with tens of millions of reads using typical cluster computing resources (256Gb RAM). These algorithms help to improve the accuracy and efficiency of transcriptomic analysis based on long-read techniques, which is crucial for understanding complex biological systems and diseases. Paper III presents an algorithmic solution, cONcat, to the detection of concatenated fragments in long-read sequencing reads with typical error profiles. The algorithm is based on a greedy heuristic that employs the edit distance measure to find best-fitting fragments and divides the sequence around those points to search for fragment hits on the remaining areas of the read. The algorithm has been shown to be resilient to errors in the data and to be scalable on large numbers of reads.

Place, publisher, year, edition, pages
Stockholm: Department of Mathematics, Stockholm University, 2025. p. 46
National Category
Bioinformatics (Computational Biology)
Research subject
Computational Mathematics
Identifiers
urn:nbn:se:su:diva-243271 (URN)978-91-8107-294-5 (ISBN)978-91-8107-295-2 (ISBN)
Public defence
2025-08-27, Lärosal 10, vån 2, hus 2, Albano, Albanovägen 18, Stockholm, 13:00 (English)
Opponent
Supervisors
Available from: 2025-06-03 Created: 2025-05-20 Last updated: 2025-05-23Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textPubMedScopus

Authority records

Petri, Alexander J.Sahlin, Kristoffer

Search in DiVA

By author/editor
Petri, Alexander J.Sahlin, Kristoffer
By organisation
Department of MathematicsScience for Life Laboratory (SciLifeLab)
In the same journal
Bioinformatics
Bioinformatics (Computational Biology)

Search outside of DiVA

GoogleGoogle Scholar

doi
pubmed
urn-nbn

Altmetric score

doi
pubmed
urn-nbn
Total: 138 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf