Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
cONcat: Computational reconstruction of concatenated fragments from long Oxford Nanopore reads
Stockholm University, Faculty of Science, Department of Mathematics. Stockholm University, Science for Life Laboratory (SciLifeLab).ORCID iD: 0009-0005-9397-0341
Show others and affiliations
2025 (English)In: Article in journal (Other academic) Submitted
Abstract [en]

Synthetic combinatorial DNA libraries are widely used to produce protein variants, optimize binders, and for high throughput studies of protein - DNA interactions. The libraries can be made by researchers or vendors and high-throughput sequencing is used for both quality control and to study the outcome of selection experiments. Oxford nanopore sequencing (ONT) is well suited to this as it allows for long read lengths and can be done rapidly with low-cost instrumentation. However, it suffers from a lower overall read accuracy and an uneven error profile. No current bioinformatics tools are well suited to the challenge of deducing the composition and order of constituent members of combinatorial libraries from ONT reads.

We introduce cONcat, an algorithm to identify the makeup of concatenated DNA fragments in a set of ONT sequencing reads from a pool of known fragments. cONcat uses the edit distance-based recursive covering algorithm for finding the best possible matchings between the fragments and the reads. In our experiments on simulated and experimental data, cONcat could accurately detect the correct fragment coverings given the short fragment sizes (< 20bp) and the sequencing errors present in ONT reads. However, we find that the high error rates in the start of ONT reads make it challenging to get confident coverage there, inferring a need for experimental strategies to avoid key sequence information in the start of reads.

Place, publisher, year, edition, pages
2025.
National Category
Bioinformatics (Computational Biology)
Identifiers
URN: urn:nbn:se:su:diva-243270DOI: 10.1101/2025.03.05.641699OAI: oai:DiVA.org:su-243270DiVA, id: diva2:1959527
Available from: 2025-05-20 Created: 2025-05-20 Last updated: 2025-06-05
In thesis
1. Computational methods for long-read sequencing data analysis
Open this publication in new window or tab >>Computational methods for long-read sequencing data analysis
2025 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

This thesis presents algorithms developed for long-read sequencing techniques, which, since their introduction in the 2010’s become increasingly important approaches in modern bioscientific research. The first two papers cover the development of algorithms for our de novo transcriptome prediction pipeline, the isON pipeline, while the third paper describes an algorithm used for biotechnological analysis of ligated fragments. Paper I introduces isONform, an algorithm capable of predicting different gene products, called isoforms, from a set of long reads sequenced from complementary DNA without the need to rely on a reference or annotation. IsONform is a tool that is part of a larger long-read transcriptome pipeline, isON pipeline, that consists of clustering and error correction steps prior to the isoform prediction. The isONform algorithm is based on the construction of a directed acyclic graph with minimizer-pairs as nodes and connecting neighboring minimizer-pairs on the reads with edges. The algorithm then employs an iterative bubble-popping scheme to merge nodes to ultimately follow all distinct paths through the graph generating the final isoform predictions. The algorithm has been shown to outperform existing state-of-the-art algorithms, while showing comparable results to approaches requiring information of a reference genome and an annotation. Paper II introduces isONclust3, an algorithm used for clustering transcriptomic reads by gene family. The algorithm constitutes the first step employed in pipelines for reference-free prediction of isoforms. The algorithm is based on the minimizer indexing scheme with its novelties being a dynamic clustering approach, assessing and storing minimizers by confidence, and an iterative post-cluster merging step. The algorithm has been shown to scale better, in terms of runtime and memory usage, on large datasets than existing methods while yielding comparable or even better results with respect to clustering quality assessments. We demonstrate that isONclust3 is the only algorithm that can process the clustering of PacBio’s new Revio datasets with tens of millions of reads using typical cluster computing resources (256Gb RAM). These algorithms help to improve the accuracy and efficiency of transcriptomic analysis based on long-read techniques, which is crucial for understanding complex biological systems and diseases. Paper III presents an algorithmic solution, cONcat, to the detection of concatenated fragments in long-read sequencing reads with typical error profiles. The algorithm is based on a greedy heuristic that employs the edit distance measure to find best-fitting fragments and divides the sequence around those points to search for fragment hits on the remaining areas of the read. The algorithm has been shown to be resilient to errors in the data and to be scalable on large numbers of reads.

Place, publisher, year, edition, pages
Stockholm: Department of Mathematics, Stockholm University, 2025. p. 46
National Category
Bioinformatics (Computational Biology)
Research subject
Computational Mathematics
Identifiers
urn:nbn:se:su:diva-243271 (URN)978-91-8107-294-5 (ISBN)978-91-8107-295-2 (ISBN)
Public defence
2025-08-27, Lärosal 10, vån 2, hus 2, Albano, Albanovägen 18, Stockholm, 13:00 (English)
Opponent
Supervisors
Available from: 2025-06-03 Created: 2025-05-20 Last updated: 2025-05-23Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full text

Authority records

Petri, Alexander J.Sahlin, Kristoffer

Search in DiVA

By author/editor
Petri, Alexander J.Thi-Huyen Nguyen, MaiSahlin, Kristoffer
By organisation
Department of MathematicsScience for Life Laboratory (SciLifeLab)
Bioinformatics (Computational Biology)

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 102 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf