Change search
ReferencesLink to record
Permanent link

Direct link
Assembly scaffolding with PE-contaminated mate-pair libraries
Stockholm University, Science for Life Laboratory (SciLifeLab). Stockholm University, Faculty of Science, Numerical Analysis and Computer Science (NADA).
Number of Authors: 3
2016 (English)In: Bioinformatics, ISSN 1367-4803, E-ISSN 1367-4811, Vol. 32, no 13, 1925-1932 p.Article in journal (Refereed) Published
Abstract [en]

Motivation: Scaffolding is often an essential step in a genome assembly process, in which contigs are ordered and oriented using read pairs from a combination of paired-end libraries and longer-range mate-pair libraries. Although a simple idea, scaffolding is unfortunately hard to get right in practice. One source of problems is so-called PE-contamination in mate-pair libraries, in which a non-negligible fraction of the read pairs get the wrong orientation and a much smaller insert size than what is expected. This contamination has been discussed before, in relation to integrated scaffolders, but solutions rely on the orientation being observable, e.g. by finding the junction adapter sequence in the reads. This is not always possible, making orientation and insert size of a read pair stochastic. To our knowledge, there is neither previous work on modeling PE-contamination, nor a study on the effect PE-contamination has on scaffolding quality. Results: We have addressed PE-contamination in an update to our scaffolder BESST. We formulate the problem as an integer linear program which is solved using an efficient heuristic. The new method shows significant improvement over both integrated and stand-alone scaffolders in our experiments. The impact of modeling PE-contamination is quantified by comparing with the previous BESST model. We also show how other scaffolders are vulnerable to PE-contaminated libraries, resulting in an increased number of misassemblies, more conservative scaffolding and inflated assembly sizes.

Place, publisher, year, edition, pages
2016. Vol. 32, no 13, 1925-1932 p.
National Category
Biological Sciences Environmental Biotechnology Computer and Information Science Mathematics
Identifiers
URN: urn:nbn:se:su:diva-132540DOI: 10.1093/bioinformatics/btw064ISI: 000379761500002PubMedID: 27153683OAI: oai:DiVA.org:su-132540DiVA: diva2:955280
Available from: 2016-08-25 Created: 2016-08-15 Last updated: 2016-08-25Bibliographically approved

Open Access in DiVA

No full text

Other links

Publisher's full textPubMed

Search in DiVA

By author/editor
Arvestad, Lars
By organisation
Science for Life Laboratory (SciLifeLab)Numerical Analysis and Computer Science (NADA)
In the same journal
Bioinformatics
Biological SciencesEnvironmental BiotechnologyComputer and Information ScienceMathematics

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

Altmetric score

ReferencesLink to record
Permanent link

Direct link