Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
isONform: reference-free transcriptome reconstruction from Oxford Nanopore data
Stockholm University, Faculty of Science, Department of Mathematics. Stockholm University, Science for Life Laboratory (SciLifeLab).
Stockholm University, Faculty of Science, Department of Mathematics. Stockholm University, Science for Life Laboratory (SciLifeLab).ORCID iD: 0000-0001-7378-2320
Number of Authors: 22023 (English)In: Bioinformatics, ISSN 1367-4803, E-ISSN 1367-4811, Vol. 39, p. i222-i231Article in journal (Refereed) Published
Abstract [en]

Motivation With advances in long-read transcriptome sequencing, we can now fully sequence transcripts, which greatly improves our ability to study transcription processes. A popular long-read transcriptome sequencing technique is Oxford Nanopore Technologies (ONT), which through its cost-effective sequencing and high throughput, has the potential to characterize the transcriptome in a cell. However, due to transcript variability and sequencing errors, long cDNA reads need substantial bioinformatic processing to produce a set of isoform predictions from the reads. Several genome and annotation-based methods exist to produce transcript predictions. However, such methods require high-quality genomes and annotations and are limited by the accuracy of long-read splice aligners. In addition, gene families with high heterogeneity may not be well represented by a reference genome and would benefit from reference-free analysis. Reference-free methods to predict transcripts from ONT, such as RATTLE, exist, but their sensitivity is not comparable to reference-based approaches.Results We present isONform, a high-sensitivity algorithm to construct isoforms from ONT cDNA sequencing data. The algorithm is based on iterative bubble popping on gene graphs built from fuzzy seeds from the reads. Using simulated, synthetic, and biological ONT cDNA data, we show that isONform has substantially higher sensitivity than RATTLE albeit with some loss in precision. On biological data, we show that isONform's predictions have substantially higher consistency with the annotation-based method StringTie2 compared with RATTLE. We believe isONform can be used both for isoform construction for organisms without well-annotated genomes and as an orthogonal method to verify predictions of reference-based methods.Availability and implementation

Place, publisher, year, edition, pages
2023. Vol. 39, p. i222-i231
National Category
Biological Sciences Environmental Biotechnology Computer and Information Sciences Mathematics
Identifiers
URN: urn:nbn:se:su:diva-220840DOI: 10.1093/bioinformatics/btad264ISI: 001027457000029PubMedID: 37387174Scopus ID: 2-s2.0-85163651809OAI: oai:DiVA.org:su-220840DiVA, id: diva2:1797281
Available from: 2023-09-14 Created: 2023-09-14 Last updated: 2025-05-20Bibliographically approved
In thesis
1. Computational methods for long-read sequencing data analysis
Open this publication in new window or tab >>Computational methods for long-read sequencing data analysis
2025 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

This thesis presents algorithms developed for long-read sequencing techniques, which, since their introduction in the 2010’s become increasingly important approaches in modern bioscientific research. The first two papers cover the development of algorithms for our de novo transcriptome prediction pipeline, the isON pipeline, while the third paper describes an algorithm used for biotechnological analysis of ligated fragments. Paper I introduces isONform, an algorithm capable of predicting different gene products, called isoforms, from a set of long reads sequenced from complementary DNA without the need to rely on a reference or annotation. IsONform is a tool that is part of a larger long-read transcriptome pipeline, isON pipeline, that consists of clustering and error correction steps prior to the isoform prediction. The isONform algorithm is based on the construction of a directed acyclic graph with minimizer-pairs as nodes and connecting neighboring minimizer-pairs on the reads with edges. The algorithm then employs an iterative bubble-popping scheme to merge nodes to ultimately follow all distinct paths through the graph generating the final isoform predictions. The algorithm has been shown to outperform existing state-of-the-art algorithms, while showing comparable results to approaches requiring information of a reference genome and an annotation. Paper II introduces isONclust3, an algorithm used for clustering transcriptomic reads by gene family. The algorithm constitutes the first step employed in pipelines for reference-free prediction of isoforms. The algorithm is based on the minimizer indexing scheme with its novelties being a dynamic clustering approach, assessing and storing minimizers by confidence, and an iterative post-cluster merging step. The algorithm has been shown to scale better, in terms of runtime and memory usage, on large datasets than existing methods while yielding comparable or even better results with respect to clustering quality assessments. We demonstrate that isONclust3 is the only algorithm that can process the clustering of PacBio’s new Revio datasets with tens of millions of reads using typical cluster computing resources (256Gb RAM). These algorithms help to improve the accuracy and efficiency of transcriptomic analysis based on long-read techniques, which is crucial for understanding complex biological systems and diseases. Paper III presents an algorithmic solution, cONcat, to the detection of concatenated fragments in long-read sequencing reads with typical error profiles. The algorithm is based on a greedy heuristic that employs the edit distance measure to find best-fitting fragments and divides the sequence around those points to search for fragment hits on the remaining areas of the read. The algorithm has been shown to be resilient to errors in the data and to be scalable on large numbers of reads.

Place, publisher, year, edition, pages
Stockholm: Department of Mathematics, Stockholm University, 2025. p. 46
National Category
Bioinformatics (Computational Biology)
Research subject
Computational Mathematics
Identifiers
urn:nbn:se:su:diva-243271 (URN)978-91-8107-294-5 (ISBN)978-91-8107-295-2 (ISBN)
Public defence
2025-08-27, Lärosal 10, vån 2, hus 2, Albano, Albanovägen 18, Stockholm, 13:00 (English)
Opponent
Supervisors
Available from: 2025-06-03 Created: 2025-05-20 Last updated: 2025-05-23Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textPubMedScopus

Authority records

Petri, Alexander J.Sahlin, Kristoffer

Search in DiVA

By author/editor
Petri, Alexander J.Sahlin, Kristoffer
By organisation
Department of MathematicsScience for Life Laboratory (SciLifeLab)
In the same journal
Bioinformatics
Biological SciencesEnvironmental BiotechnologyComputer and Information SciencesMathematics

Search outside of DiVA

GoogleGoogle Scholar

doi
pubmed
urn-nbn

Altmetric score

doi
pubmed
urn-nbn
Total: 115 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf