Home LiteratureArticle Details
PMID: 28298431 Published · ppublish English Journal Article Research Support, N.I.H., Intramural Research Support, U.S. Gov't, Non-P.H.S.

Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation.

Genome research ·Vol. 27 ·No. 5 ·2017-00-00 ·Pages 722-736

Koren S, Walenz BP, Berlin K, Miller JR, Bergman NH, Phillippy AM

Abstract

Long-read single-molecule sequencing has revolutionized de novo genome assembly and enabled the automated reconstruction of reference-quality genomes. However, given the relatively high error rates of such technologies, efficient and accurate assembly of large repeats and closely related haplotypes remains challenging. We address these issues with Canu, a successor of Celera Assembler that is specifically designed for noisy single-molecule sequences. Canu introduces support for nanopore sequencing, halves depth-of-coverage requirements, and improves assembly continuity while simultaneously reducing runtime by an order of magnitude on large genomes versus Celera Assembler 8.2. These advances result from new overlapping and assembly algorithms, including an adaptive overlapping strategy based on tf-idf weighted MinHash and a sparse assembly graph construction that avoids collapsing diverged repeats and haplotypes. We demonstrate that Canu can reliably assemble complete microbial genomes and near-complete eukaryotic chromosomes using either Pacific Biosciences (PacBio) or Oxford Nanopore technologies and achieves a contig NG50 of >21 Mbp on both human and Drosophila melanogaster PacBio data sets. For assembly structures that cannot be linearly represented, Canu provides graph-based assembly outputs in graphical fragment assembly (GFA) format for analysis or integration with complementary phasing and scaffolding techniques. The combination of such highly resolved assembly graphs with long-range scaffolding information promises the complete and automated assembly of complex genomes.

MeSH Terms
Animals Contig Mapping/methods,standards Drosophila melanogaster/genetics Genome, Bacterial Genomics/methods,standards Humans Repetitive Sequences, Nucleic Acid Sequence Analysis, DNA/methods,standards Software
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Koren Sergey
Genome Informatics Section, Computational and Statistical Genomics Branch, National Human Genome Research Institute, National Institutes of Health, Bethesda, Maryland 20892, USA.
Walenz Brian P
Genome Informatics Section, Computational and Statistical Genomics Branch, National Human Genome Research Institute, National Institutes of Health, Bethesda, Maryland 20892, USA.
Berlin Konstantin
Invincea Incorporated, Fairfax, Virginia 22030, USA.
Miller Jason R
J. Craig Venter Institute, Rockville, Maryland 20850, USA.
Bergman Nicholas H
National Biodefense Analysis and Countermeasures Center, Frederick, Maryland 21702, USA.
Phillippy Adam M ORCID
Genome Informatics Section, Computational and Statistical Genomics Branch, National Human Genome Research Institute, National Institutes of Health, Bethesda, Maryland 20892, USA.
References (60)
60 references, click to expand
  1. LoRDEC: accurate and efficient long read error correction.
    Bioinformatics. 2014 Dec 15;30(24):3506-14 PMID: 25165095
  2. Single-molecule sequencing and chromatin conformation capture enable de novo reference assembly of the domestic goat genome.
    Nat Genet. 2017 Apr;49(4):643-650 PMID: 28263316
  3. Fast and accurate de novo genome assembly from long uncorrected reads.
    Genome Res. 2017 May;27(5):737-746 PMID: 28100585
  4. Increased plasmid copy number is essential for Yersinia T3SS function and virulence.
    Science. 2016 Jul 29;353(6298):492-5 PMID: 27365311
  5. Aggressive assembly of pyrosequencing reads with mates.
    Bioinformatics. 2008 Dec 15;24(24):2818-24 PMID: 18952627
  6. proovread: large-scale high-accuracy PacBio correction through iterative short read consensus.
    Bioinformatics. 2014 Nov 1;30(21):3004-11 PMID: 25015988
  7. PBSIM: PacBio reads simulator--toward accurate genome assembly.
    Bioinformatics. 2013 Jan 1;29(1):119-21 PMID: 23129296
  8. One chromosome, one contig: complete microbial genomes from long-read sequencing and assembly.
    Curr Opin Microbiol. 2015 Feb;23:110-20 PMID: 25461581
  9. Exploring variation-aware contig graphs for (comparative) metagenomics using MaryGold.
    Bioinformatics. 2013 Nov 15;29(22):2826-34 PMID: 24058058
  10. Assembling large genomes with single-molecule sequencing and locality-sensitive hashing.
    Nat Biotechnol. 2015 Jun;33(6):623-30 PMID: 26006009
  11. The fragment assembly string graph.
    Bioinformatics. 2005 Sep 1;21 Suppl 2:ii79-85 PMID: 16204131
  12. Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data.
    Nat Methods. 2013 Jun;10(6):563-9 PMID: 23644548
  13. An improved genome assembly uncovers prolific tandem repeats in Atlantic cod.
    BMC Genomics. 2017 Jan 18;18(1):95 PMID: 28100185
  14. Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
    Genome Res. 2008 May;18(5):821-9 PMID: 18349386
  15. Oxford Nanopore sequencing, hybrid error correction, and de novo assembly of a eukaryotic genome.
    Genome Res. 2015 Nov;25(11):1750-6 PMID: 26447147
  16. Finishing the euchromatic sequence of the human genome.
    Nature. 2004 Oct 21;431(7011):931-45 PMID: 15496913
  17. Efficiently detecting polymorphisms during the fragment assembly process.
    Bioinformatics. 2002;18 Suppl 1:S294-302 PMID: 12169559
  18. Hybrid error correction and de novo assembly of single-molecule sequencing reads.
    Nat Biotechnol. 2012 Jul 01;30(7):693-700 PMID: 22750884
  19. Evaluation of hybrid and non-hybrid methods for de novo assembly of nanopore reads.
    Bioinformatics. 2016 Sep 1;32(17):2582-9 PMID: 27162186
  20. Phased diploid genome assembly with single-molecule real-time sequencing.
    Nat Methods. 2016 Dec;13(12 ):1050-1054 PMID: 27749838
  21. Genome assembly forensics: finding the elusive mis-assembly.
    Genome Biol. 2008;9(3):R55 PMID: 18341692
  22. Base-calling of automated sequencer traces using phred. II. Error probabilities.
    Genome Res. 1998 Mar;8(3):186-94 PMID: 9521922
  23. Contiguous and accurate de novo assembly of metazoan genomes with modest long read coverage.
    Nucleic Acids Res. 2016 Nov 2;44(19):e147 PMID: 27458204
  24. GAGE: A critical evaluation of genome assemblies and assembly algorithms.
    Genome Res. 2012 Mar;22(3):557-67 PMID: 22147368
  25. The genome sequence of the malaria mosquito Anopheles gambiae.
    Science. 2002 Oct 4;298(5591):129-49 PMID: 12364791
  26. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement.
    PLoS One. 2014 Nov 19;9(11):e112963 PMID: 25409509
  27. Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences.
    Bioinformatics. 2016 Jul 15;32(14):2103-10 PMID: 27153593
  28. SSAHA: a fast search method for large DNA databases.
    Genome Res. 2001 Oct;11(10):1725-9 PMID: 11591649
  29. Comparison of bacterial genome assembly software for MinION data and their applicability to medical microbiology.
    Microb Genom. 2016 Sep 8;2(9):e000085 PMID: 28348876
  30. Optimal assembly for high throughput shotgun sequencing.
    BMC Bioinformatics. 2013;14 Suppl 5:S18 PMID: 23902516
  31. DNA sequencing with nanopores.
    Nat Biotechnol. 2012 Apr 10;30(4):326-8 PMID: 22491281
  32. Assessing the quality of the DNA sequence from the Human Genome Project.
    Genome Res. 1999 Jan;9(1):1-4 PMID: 9927479
  33. Real-time DNA sequencing from single polymerase molecules.
    Science. 2009 Jan 2;323(5910):133-8 PMID: 19023044
  34. A whole-genome assembly of Drosophila.
    Science. 2000 Mar 24;287(5461):2196-204 PMID: 10731133
  35. Bandage: interactive visualization of de novo genome assemblies.
    Bioinformatics. 2015 Oct 15;31(20):3350-2 PMID: 26099265
  36. Bambus 2: scaffolding metagenomes.
    Bioinformatics. 2011 Nov 1;27(21):2964-71 PMID: 21926123
  37. Parametric complexity of sequence assembly: theory and applications to next generation sequencing.
    J Comput Biol. 2009 Jul;16(7):897-908 PMID: 19580519
  38. Whole-genome haplotype reconstruction using proximity-ligation and shotgun sequencing.
    Nat Biotechnol. 2013 Dec;31(12):1111-8 PMID: 24185094
  39. Versatile and open software for comparing large genomes.
    Genome Biol. 2004;5(2):R12 PMID: 14759262
  40. Whole-genome shotgun assembly and comparison of human genome assemblies.
    Proc Natl Acad Sci U S A. 2004 Feb 17;101(7):1916-21 PMID: 14769938
  41. Mash: fast genome and metagenome distance estimation using MinHash.
    Genome Biol. 2016 Jun 20;17(1):132 PMID: 27323842
  42. The Release 6 reference sequence of the Drosophila melanogaster genome.
    Genome Res. 2015 Mar;25(3):445-58 PMID: 25589440
  43. DBG2OLC: Efficient Assembly of Large Genomes Using Long Erroneous Reads of the Third Generation Sequencing Technologies.
    Sci Rep. 2016 Aug 30;6:31900 PMID: 27573208
  44. Reducing assembly complexity of microbial genomes with single-molecule sequencing.
    Genome Biol. 2013;14(9):R101 PMID: 24034426
  45. Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions.
    Nat Biotechnol. 2013 Dec;31(12):1119-25 PMID: 24185095
  46. Long-read sequence assembly of the gorilla genome.
    Science. 2016 Apr 1;352(6281):aae0344 PMID: 27034376
  47. A flexible and efficient template format for circular consensus sequencing and SNP detection.
    Nucleic Acids Res. 2010 Aug;38(15):e159 PMID: 20571086
  48. Long-read, whole-genome shotgun sequence data for five model organisms.
    Sci Data. 2014 Nov 25;1:140045 PMID: 25977796
  49. High-throughput genome scaffolding from in vivo DNA interaction frequency.
    Nat Biotechnol. 2013 Dec;31(12):1143-7 PMID: 24270850
  50. Quality assessment of the human genome sequence.
    Nature. 2004 May 27;429(6990):365-8 PMID: 15164052
  51. de novo assembly and population genomic survey of natural yeast isolates with the Oxford Nanopore MinION sequencer.
    Gigascience. 2017 Feb 1;6(2):1-13 PMID: 28369459
  52. hybridSPAdes: an algorithm for hybrid assembly of short and long reads.
    Bioinformatics. 2016 Apr 1;32(7):1009-15 PMID: 26589280
  53. The bonobo genome compared with the chimpanzee and human genomes.
    Nature. 2012 Jun 28;486(7404):527-31 PMID: 22722832
  54. Characterizing and measuring bias in sequence data.
    Genome Biol. 2013 May 29;14(5):R51 PMID: 23718773
  55. Rapid genome mapping in nanochannel arrays for highly complete and accurate de novo sequence assembly of the complex Aegilops tauschii genome.
    PLoS One. 2013;8(2):e55864 PMID: 23405223
  56. Improved data analysis for the MinION nanopore sequencer.
    Nat Methods. 2015 Apr;12(4):351-6 PMID: 25686389
  57. Long-read sequencing and de novo assembly of a Chinese genome.
    Nat Commun. 2016 Jun 30;7:12065 PMID: 27356984
  58. A complete bacterial genome assembled de novo using only nanopore sequencing data.
    Nat Methods. 2015 Aug;12(8):733-5 PMID: 26076426
  59. FlyBase: establishing a Gene Group resource for Drosophila melanogaster.
    Nucleic Acids Res. 2016 Jan 4;44(D1):D786-92 PMID: 26467478
  60. Haplotyping germline and cancer genomes with high-throughput linked-read sequencing.
    Nat Biotechnol. 2016 Mar;34(3):303-11 PMID: 26829319
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1549-5469
Published
2017-00-00
Epub
2017-00-15
Pages
722-736
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC5411767
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com