Abstract
The NCBI Reference Sequence (RefSeq) project and the NIH Mammalian Gene Collection (MGC) together define a set of approximately 30,000 nonredundant human mRNA sequences with identified coding regions representing 17,000 distinct loci. These high-quality mRNA sequences allow for the identification of transcribed regions in the human genome sequence, and many researchers accept them as the correct representation of each defined gene sequence. Computational comparison of these mRNA sequences and the recently published essentially finished human genome sequence reveals several thousand undocumented nonsynonymous substitution and frame shift discrepancies between the two resources. Additional analysis is undertaken to verify that the euchromatic human genome is sufficiently complete--containing nearly the whole mRNA collection, thus allowing for a comprehensive analysis to be undertaken. Many of the discrepancies will prove to be genuine polymorphisms in the human population, somatic cell genomic variants, or examples of RNA editing. It is observed that the genome sequence variant has significant additional support from other mRNAs and ESTs, almost four times more often than does the mRNA variant, suggesting that the genome sequence is more accurate. In approximately 15% of these cases, there is substantial support for both variants, suggestive of an undocumented polymorphism. An initial screening against a 24-individual genomic DNA diversity panel verified 60% of a small set of potential single nucleotide polymorphisms from which successful results could be obtained. We also find statistical evidence that a few of these discrepancies are due to RNA editing. Overall, these results suggest that the mRNA collections may contain a substantial number of errors. For current and future mRNA collections, it may be prudent to fully reconcile each genome sequence discrepancy, classifying each as a polymorphism, site of RNA editing or somatic cell variation, or genome sequence error.
MeSH Terms
Computational Biology
Expressed Sequence Tags
Genetic Variation
Genome, Human
Human Genome Project
Humans
Polymorphism, Genetic
RNA Editing
RNA, Messenger/analysis
Sequence Analysis, DNA
Authors & Affiliations
9 authors, click to expand affiliations / ORCID
Furey Terrence S
Center for Biomolecular Science and Engineering, Department of Computer Science, University of California, Santa Cruz, Santa Cruz, California 95064, USA. booch@cse.ucsc.edu
Diekhans Mark
Lu Yontao
Graves Tina A
Oddy Lachlan
Randall-Maher Jennifer
Hillier LaDeana W
Wilson Richard K
Haussler David
References (28)
28 references, click to expand
-
A Drosophila full-length cDNA resource.
Genome Biol. 2002;3(12):RESEARCH0080
PMID: 12537569
-
Analysis of the mouse transcriptome based on functional annotation of 60,770 full-length cDNAs.
Nature. 2002 Dec 5;420(6915):563-73
PMID: 12466851
-
The EMBL Nucleotide Sequence Database.
Nucleic Acids Res. 2004 Jan 1;32(Database issue):D27-30
PMID: 14681351
-
DDBJ in the stream of various biological data.
Nucleic Acids Res. 2004 Jan 1;32(Database issue):D31-4
PMID: 14681352
-
Generation and initial analysis of more than 15,000 full-length human and mouse cDNA sequences.
Proc Natl Acad Sci U S A. 2002 Dec 24;99(26):16899-903
PMID: 12477932
-
NCBI Reference Sequence project: update and current status.
Nucleic Acids Res. 2003 Jan 1;31(1):34-7
PMID: 12519942
-
SOURCE: a unified genomic resource of functional annotations, ontologies, and gene expression data.
Nucleic Acids Res. 2003 Jan 1;31(1):219-23
PMID: 12519986
-
Concatenation cDNA sequencing for transcriptome analysis.
C R Biol. 2003 Oct-Nov;326(10-11):971-7
PMID: 14744103
-
Quality assessment of the human genome sequence.
Nature. 2004 May 27;429(6990):365-8
PMID: 15164052
-
The status, quality, and expansion of the NIH full-length cDNA project: the Mammalian Gene Collection (MGC).
Genome Res. 2004 Oct;14(10B):2121-7
PMID: 15489334
-
Basic local alignment search tool.
J Mol Biol. 1990 Oct 5;215(3):403-10
PMID: 2231712
-
RNA editing in brain controls a determinant of ion flow in glutamate-gated channels.
Cell. 1991 Oct 4;67(1):11-9
PMID: 1717158
-
RNA editing of the glutamate receptor subunits GluR2 and GluR6 in human brain tissue.
J Neurochem. 1994 Nov;63(5):1596-602
PMID: 7523595
-
A comprehensive genetic map of the human genome based on 5,264 microsatellites.
Nature. 1996 Mar 14;380(6570):152-4
PMID: 8600387
-
Large-scale concatenation cDNA sequencing.
Genome Res. 1997 Apr;7(4):353-8
PMID: 9110174
-
Comprehensive human genetic maps: individual and sex-specific variation in recombination.
Am J Hum Genet. 1998 Sep;63(3):861-9
PMID: 9718341
-
The mammalian gene collection.
Science. 1999 Oct 15;286(5439):455-7
PMID: 10521335
-
RefSeq and LocusLink: NCBI gene-centered resources.
Nucleic Acids Res. 2001 Jan 1;29(1):137-40
PMID: 11125071
-
Initial sequencing and analysis of the human genome.
Nature. 2001 Feb 15;409(6822):860-921
PMID: 11237011
-
RNA editing by base deamination: more enzymes, more targets, new mysteries.
Trends Biochem Sci. 2001 Jun;26(6):376-84
PMID: 11406411
-
Spidey: a tool for mRNA-to-genomic alignments.
Genome Res. 2001 Nov;11(11):1952-7
PMID: 11691860
-
GenBank.
Nucleic Acids Res. 2002 Jan 1;30(1):17-20
PMID: 11752243
-
BLAT--the BLAST-like alignment tool.
Genome Res. 2002 Apr;12(4):656-64
PMID: 11932250
-
RNA editing by adenosine deaminases that act on RNA.
Annu Rev Biochem. 2002;71:817-46
PMID: 12045112
-
A high-resolution recombination map of the human genome.
Nat Genet. 2002 Jul;31(3):241-7
PMID: 12053178
-
Recent segmental duplications in the human genome.
Science. 2002 Aug 9;297(5583):1003-7
PMID: 12169732
-
Initial sequencing and comparative analysis of the mouse genome.
Nature. 2002 Dec 5;420(6915):520-62
PMID: 12466850
-
Low editing efficiency of GluR2 mRNA is associated with a low relative abundance of ADAR2 mRNA in white matter of normal human brain.
Eur J Neurosci. 2003 Jul;18(1):23-33
PMID: 12859334