Home LiteratureArticle Details
PMID: 28396521 Published · ppublish English Journal Article Research Support, N.I.H., Intramural Research Support, Non-U.S. Gov't Research Support, N.I.H., Extramural

Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly.

Genome research ·Vol. 27 ·No. 5 ·2017-00-00 ·Pages 849-864

Schneider VA, Graves-Lindsay T, Howe K, Bouk N, Chen HC, Kitts PA, Murphy TD, Pruitt KD, Thibaud-Nissen F, Albracht D, Fulton RS, Kremitzki M, Magrini V, Markovic C, McGrath S, Steinberg KM, Auger K, Chow W, Collins J, Harden G, Hubbard T, Pelan S, Simpson JT, Threadgold G, Torrance J, Wood JM, Clarke L, Koren S, Boitano M, Peluso P, Li H, Chin CS, Phillippy AM, Durbin R, Wilson RK, Flicek P, Eichler EE, Church DM

Abstract

The human reference genome assembly plays a central role in nearly all aspects of today's basic and clinical research. GRCh38 is the first coordinate-changing assembly update since 2009; it reflects the resolution of roughly 1000 issues and encompasses modifications ranging from thousands of single base changes to megabase-scale path reorganizations, gap closures, and localization of previously orphaned sequences. We developed a new approach to sequence generation for targeted base updates and used data from new genome mapping technologies and single haplotype resources to identify and resolve larger assembly issues. For the first time, the reference assembly contains sequence-based representations for the centromeres. We also expanded the number of alternate loci to create a reference that provides a more robust representation of human population variation. We demonstrate that the updates render the reference an improved annotation substrate, alter read alignments in unchanged regions, and impact variant interpretation at clinically relevant loci. We additionally evaluated a collection of new de novo long-read haploid assemblies and conclude that although the new assemblies compare favorably to the reference with respect to continuity, error rate, and gene completeness, the reference still provides the best representation for complex genomic regions and coding sequences. We assert that the collected updates in GRCh38 make the newer assembly a more robust substrate for comprehensive analyses that will promote our understanding of human biology and advance our efforts to improve health.

MeSH Terms
Contig Mapping/methods,standards Genome, Human Genomics/methods,standards Haploidy Haplotypes Humans Polymorphism, Genetic Reference Standards Sequence Analysis, DNA/methods,standards Software
Authors & Affiliations
38 authors, click to expand affiliations / ORCID
Schneider Valerie A
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland 20894, USA.
Graves-Lindsay Tina
McDonnell Genome Institute at Washington University, St. Louis, Missouri 63018, USA.
Howe Kerstin
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Bouk Nathan
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland 20894, USA.
Chen Hsiu-Chuan
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland 20894, USA.
Kitts Paul A
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland 20894, USA.
Murphy Terence D
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland 20894, USA.
Pruitt Kim D
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland 20894, USA.
Thibaud-Nissen Françoise
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland 20894, USA.
Albracht Derek
McDonnell Genome Institute at Washington University, St. Louis, Missouri 63018, USA.
Fulton Robert S
McDonnell Genome Institute at Washington University, St. Louis, Missouri 63018, USA.
Kremitzki Milinn
McDonnell Genome Institute at Washington University, St. Louis, Missouri 63018, USA.
Magrini Vincent
McDonnell Genome Institute at Washington University, St. Louis, Missouri 63018, USA.
Markovic Chris
McDonnell Genome Institute at Washington University, St. Louis, Missouri 63018, USA.
McGrath Sean
McDonnell Genome Institute at Washington University, St. Louis, Missouri 63018, USA.
Steinberg Karyn Meltz
McDonnell Genome Institute at Washington University, St. Louis, Missouri 63018, USA.
Auger Kate
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Chow William
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Collins Joanna
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Harden Glenn
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Hubbard Timothy
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Pelan Sarah
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Simpson Jared T
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Threadgold Glen
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Torrance James
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Wood Jonathan M
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Clarke Laura
European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom.
Koren Sergey
National Human Genome Research Institute, National Institutes of Health, Bethesda, Maryland 20892, USA.
Boitano Matthew
Pacific Biosciences, Menlo Park, California 94025, USA.
Peluso Paul
Pacific Biosciences, Menlo Park, California 94025, USA.
Li Heng
Broad Institute, Cambridge, Massachusetts 02142, USA.
Chin Chen-Shan
Pacific Biosciences, Menlo Park, California 94025, USA.
Phillippy Adam M
National Human Genome Research Institute, National Institutes of Health, Bethesda, Maryland 20892, USA.
Durbin Richard
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, United Kingdom.
Wilson Richard K
McDonnell Genome Institute at Washington University, St. Louis, Missouri 63018, USA.
Flicek Paul
European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom.
Eichler Evan E
Department of Genome Sciences, University of Washington School of Medicine, Seattle, Washington 98195, USA. | Howard Hughes Medical Institute, University of Washington, Seattle, Washington 98195, USA.
Church Deanna M
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland 20894, USA.
References (71)
71 references, click to expand
  1. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  2. Segmental duplications: organization and impact within the current human genome project assembly.
    Genome Res. 2001 Jun;11(6):1005-17 PMID: 11381028
  3. Paternal origins of complete hydatidiform moles proven by whole genome single-nucleotide polymorphism haplotyping.
    Genomics. 2002 Jan;79(1):58-62 PMID: 11827458
  4. The DNA sequence and comparative analysis of human chromosome 5.
    Nature. 2004 Sep 16;431(7006):268-74 PMID: 15372022
  5. Finishing the euchromatic sequence of the human genome.
    Nature. 2004 Oct 21;431(7011):931-45 PMID: 15496913
  6. Segmental duplications and copy-number variation in the human genome.
    Am J Hum Genet. 2005 Jul;77(1):78-88 PMID: 15918152
  7. A haplotype map of the human genome.
    Nature. 2005 Oct 27;437(7063):1299-320 PMID: 16255080
  8. The diploid genome sequence of an individual human.
    PLoS Biol. 2007 Sep 4;5(10):e254 PMID: 17803354
  9. Variation analysis and gene annotation of eight MHC haplotypes: the MHC Haplotype Project.
    Immunogenetics. 2008 Jan;60(1):1-18 PMID: 18193213
  10. Mapping and sequencing of structural variation from eight human genomes.
    Nature. 2008 May 1;453(7191):56-64 PMID: 18451855
  11. Adaptive evolution of UGT2B17 copy-number variation.
    Am J Hum Genet. 2008 Sep;83(3):337-46 PMID: 18760392
  12. The diploid genome sequence of an Asian individual.
    Nature. 2008 Nov 6;456(7218):60-5 PMID: 18987735
  13. Evolutionary toggling of the MAPT 17q21.31 inversion region.
    Nat Genet. 2008 Sep;40(9):1076-83 PMID: 19165922
  14. Building the sequence map of the human pan-genome.
    Nat Biotechnol. 2010 Jan;28(1):57-63 PMID: 19997067
  15. A draft sequence of the Neandertal genome.
    Science. 2010 May 7;328(5979):710-722 PMID: 20448178
  16. High-resolution human genome structure by single-molecule analysis.
    Proc Natl Acad Sci U S A. 2010 Jun 15;107(24):10848-53 PMID: 20534489
  17. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data.
    Genome Res. 2010 Sep;20(9):1297-303 PMID: 20644199
  18. A map of human genome variation from population-scale sequencing.
    Nature. 2010 Oct 28;467(7319):1061-73 PMID: 20981092
  19. Diversity of human copy number variation and multicopy genes.
    Science. 2010 Oct 29;330(6004):641-6 PMID: 21030649
  20. Different patterns of evolution in the centromeric and telomeric regions of group A and B haplotypes of the human killer cell Ig-like receptor locus.
    PLoS One. 2010 Dec 29;5(12):e15115 PMID: 21206914
  21. Efficient storage of high throughput DNA sequencing data using reference-based compression.
    Genome Res. 2011 May;21(5):734-40 PMID: 21245279
  22. Genome assembly has a major impact on gene content: a comparison of annotation in two Bos taurus assemblies.
    PLoS One. 2011;6(6):e21400 PMID: 21731731
  23. Modernizing reference genome assemblies.
    PLoS Biol. 2011 Jul;9(7):e1001091 PMID: 21750661
  24. Assemblathon 1: a competitive assessment of de novo short read assembly methods.
    Genome Res. 2011 Dec;21(12):2224-41 PMID: 21926179
  25. Evolution of human-specific neural SRGAP2 genes by incomplete segmental duplication.
    Cell. 2012 May 11;149(4):912-22 PMID: 22559943
  26. Limitations of the human reference genome for personalized genomics.
    PLoS One. 2012;7(7):e40294 PMID: 22811759
  27. GENCODE: the reference human genome annotation for The ENCODE Project.
    Genome Res. 2012 Sep;22(9):1760-74 PMID: 22955987
  28. DNA template strand sequencing of single-cells maps genomic rearrangements at high resolution.
    Nat Methods. 2012 Nov;9(11):1107-12 PMID: 23042453
  29. An integrated map of genetic variation from 1,092 human genomes.
    Nature. 2012 Nov 1;491(7422):56-65 PMID: 23128226
  30. Clone DB: an integrated NCBI resource for clone-associated data.
    Nucleic Acids Res. 2013 Jan;41(Database issue):D1070-8 PMID: 23193260
  31. Reevaluating assembly evaluations with feature response curves: GAGE and assemblathons.
    PLoS One. 2012;7(12):e52210 PMID: 23284938
  32. Using population admixture to help complete maps of the human genome.
    Nat Genet. 2013 Apr;45(4):406-14, 414e1-2 PMID: 23435088
  33. HAL: a hierarchical format for storing and analyzing multiple genome alignments.
    Bioinformatics. 2013 May 15;29(10):1341-2 PMID: 23505295
  34. Complete haplotype sequence of the human immunoglobulin heavy-chain variable, diversity, and joining genes and characterization of allelic and copy-number variation.
    Am J Hum Genet. 2013 Apr 4;92(4):530-46 PMID: 23541343
  35. Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data.
    Nat Methods. 2013 Jun;10(6):563-9 PMID: 23644548
  36. Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species.
    Gigascience. 2013 Jul 22;2(1):10 PMID: 23870653
  37. Independent specialization of the human and mouse X chromosomes for the male germ line.
    Nat Genet. 2013 Sep;45(9):1083-7 PMID: 23872635
  38. Mapping the human reference genome's missing sequence by three-way admixture in Latino genomes.
    Am J Hum Genet. 2013 Sep 5;93(3):411-21 PMID: 23932108
  39. ClinVar: public archive of relationships among sequence variation and human phenotype.
    Nucleic Acids Res. 2014 Jan;42(Database issue):D980-5 PMID: 24234437
  40. CrossMap: a versatile tool for coordinate conversion between genome assemblies.
    Bioinformatics. 2014 Apr 1;30(7):1006-7 PMID: 24351709
  41. Reconstructing complex regions of genomes using long-read sequencing technology.
    Genome Res. 2014 Apr;24(4):688-96 PMID: 24418700
  42. Centromere reference models for human chromosomes X and Y satellite arrays.
    Genome Res. 2014 Apr;24(4):697-707 PMID: 24501022
  43. Integrating human sequence data sets provides a resource of benchmark SNP and indel genotype calls.
    Nat Biotechnol. 2014 Mar;32(3):246-51 PMID: 24531798
  44. Toward better understanding of artifacts in variant calling from high-coverage samples.
    Bioinformatics. 2014 Oct 15;30(20):2843-51 PMID: 24974202
  45. A framework for the interpretation of de novo mutation in human disease.
    Nat Genet. 2014 Sep;46(9):944-50 PMID: 25086666
  46. Palindromic GOLGA8 core duplicons promote chromosome 15q13.3 microdeletion and evolutionary instability.
    Nat Genet. 2014 Dec;46(12):1293-302 PMID: 25326701
  47. Sequencing of the human IG light chain loci from a hydatidiform mole BAC library reveals locus-specific signatures of genetic diversity.
    Genes Immun. 2015 Jan-Feb;16(1):24-34 PMID: 25338678
  48. Single haplotype assembly of the human genome from a hydatidiform mole.
    Genome Res. 2014 Dec;24(12):2066-76 PMID: 25373144
  49. Resolving the complexity of the human genome using single-molecule sequencing.
    Nature. 2015 Jan 29;517(7536):608-11 PMID: 25383537
  50. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement.
    PLoS One. 2014 Nov 19;9(11):e112963 PMID: 25409509
  51. Building a pan-genome reference for a population.
    J Comput Biol. 2015 May;22(5):387-401 PMID: 25565268
  52. Extending reference assembly models.
    Genome Biol. 2015 Jan 24;16:13 PMID: 25651527
  53. A naturally occurring null variant of the NMDA type glutamate receptor NR3B subunit is a risk factor of schizophrenia.
    PLoS One. 2015 Mar 13;10(3):e0116319 PMID: 25768306
  54. Using optical mapping data for the improvement of vertebrate genome assemblies.
    Gigascience. 2015 Mar 18;4:10 PMID: 25789164
  55. Assessing structural variation in a personal genome-towards a human reference diploid genome.
    BMC Genomics. 2015 Apr 11;16:286 PMID: 25886820
  56. Improved genome inference in the MHC using a population reference graph.
    Nat Genet. 2015 Jun;47(6):682-8 PMID: 25915597
  57. Sharing and Specificity of Co-expression Networks across 35 Human Tissues.
    PLoS Comput Biol. 2015 May 13;11(5):e1004220 PMID: 25970446
  58. De novo assembly of a haplotype-resolved human genome.
    Nat Biotechnol. 2015 Jun;33(6):617-22 PMID: 26006006
  59. Assembling large genomes with single-molecule sequencing and locality-sensitive hashing.
    Nat Biotechnol. 2015 Jun;33(6):623-30 PMID: 26006009
  60. Assembly and diploid architecture of an individual human genome via single-molecule technologies.
    Nat Methods. 2015 Aug;12(8):780-6 PMID: 26121404
  61. Utilizing mapping targets of sequences underrepresented in the reference assembly to reduce false positive alignments.
    Nucleic Acids Res. 2015 Nov 16;43(20):e133 PMID: 26163063
  62. FermiKit: assembly-based variant calling for Illumina resequencing data.
    Bioinformatics. 2015 Nov 15;31(22):3694-6 PMID: 26220959
  63. A global reference for human genetic variation.
    Nature. 2015 Oct 1;526(7571):68-74 PMID: 26432245
  64. Genetic variation and the de novo assembly of human genomes.
    Nat Rev Genet. 2015 Nov;16(11):627-40 PMID: 26442640
  65. Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation.
    Nucleic Acids Res. 2016 Jan 4;44(D1):D733-45 PMID: 26553804
  66. Assembly: a resource for assembled genomes at NCBI.
    Nucleic Acids Res. 2016 Jan 4;44(D1):D73-80 PMID: 26578580
  67. Extensive sequencing of seven human genomes to characterize benchmark reference materials.
    Sci Data. 2016 Jun 07;3:160025 PMID: 27271295
  68. Long-read sequencing and de novo assembly of a Chinese genome.
    Nat Commun. 2016 Jun 30;7:12065 PMID: 27356984
  69. De novo assembly and phasing of a Korean human genome.
    Nature. 2016 Oct 13;538(7624):243-247 PMID: 27706134
  70. Phased diploid genome assembly with single-molecule real-time sequencing.
    Nat Methods. 2016 Dec;13(12):1050-1054 PMID: 27749838
  71. A graph extension of the positional Burrows-Wheeler transform and its applications.
    Algorithms Mol Biol. 2017 Jul 11;12:18 PMID: 28702075
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1549-5469
Published
2017-00-00
Epub
2017-00-10
Pages
849-864
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC5411779
Subset
IM
Grants
Wellcome Trust · United Kingdom
NHGRI NIH HHS · R01 HG002385 · United States
NHGRI NIH HHS · U41 HG007635 · United States
NHGRI NIH HHS · U54 HG003079 · United States
Wellcome Trust · WT095908 · United Kingdom
Wellcome Trust · WT098051 · United Kingdom
Wellcome Trust · WT104947/Z/14/Z · United Kingdom
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com