Home LiteratureArticle Details
PMID: 24939910 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

Multiple evidence strands suggest that there may be as few as 19,000 human protein-coding genes.

Human molecular genetics ·Vol. 23 ·No. 22 ·2014-11-15 ·Pages 5866-78

Ezkurdia I, Juan D, Rodriguez JM, Frankish A, Diekhans M, Harrow J, Vazquez J, Valencia A, Tress ML

Abstract

Determining the full complement of protein-coding genes is a key goal of genome annotation. The most powerful approach for confirming protein-coding potential is the detection of cellular protein expression through peptide mass spectrometry (MS) experiments. Here, we mapped peptides detected in seven large-scale proteomics studies to almost 60% of the protein-coding genes in the GENCODE annotation of the human genome. We found a strong relationship between detection in proteomics experiments and both gene family age and cross-species conservation. Most of the genes for which we detected peptides were highly conserved. We found peptides for >96% of genes that evolved before bilateria. At the opposite end of the scale, we identified almost no peptides for genes that have appeared since primates, for genes that did not have any protein-like features or for genes with poor cross-species conservation. These results motivated us to describe a set of 2001 potential non-coding genes based on features such as weak conservation, a lack of protein features, or ambiguous annotations from major databases, all of which correlated with low peptide detection across the seven experiments. We identified peptides for just 3% of these genes. We show that many of these genes behave more like non-coding genes than protein-coding genes and suggest that most are unlikely to code for proteins under normal circumstances. We believe that their inclusion in the human protein-coding gene catalogue should be revised as part of the ongoing human genome annotation effort.

MeSH Terms
Computational Biology Genome, Human Humans Open Reading Frames Peptides/genetics Proteins/genetics,metabolism Proteomics
Chemicals
Peptides Proteins
Authors & Affiliations
9 authors, click to expand affiliations / ORCID
Ezkurdia Iakes
Unidad de Proteómica and.
Juan David
Structural Biology and Bioinformatics Programme and.
Rodriguez Jose Manuel
National Bioinformatics Institute (INB), Spanish National Cancer Research Centre (CNIO), Melchor Fernández Almagro, 3, 28029, Madrid, Spain.
Frankish Adam
Wellcome Trust Sanger Institute, Wellcome Trust Campus, Hinxton, Cambridge CB10 1SA, UK and.
Diekhans Mark
Center for Biomolecular Science and Engineering, School of Engineering, University of California Santa Cruz (UCSC), 1156 High Street, Santa Cruz, CA 95064, USA.
Harrow Jennifer
Wellcome Trust Sanger Institute, Wellcome Trust Campus, Hinxton, Cambridge CB10 1SA, UK and.
Vazquez Jesus
Laboratorio de Proteómica Cardiovascular, Centro Nacional de Investigaciones Cardiovasculares, CNIC, Melchor Fernández Almagro, 3, 28029, Madrid, Spain.
Valencia Alfonso
Structural Biology and Bioinformatics Programme and, National Bioinformatics Institute (INB), Spanish National Cancer Research Centre (CNIO), Melchor Fernández Almagro, 3, 28029, Madrid, Spain, mtress@cnio.es valencia@cnio.es.
Tress Michael L
Structural Biology and Bioinformatics Programme and, mtress@cnio.es valencia@cnio.es.
References (62)
62 references, click to expand
  1. A phylogenomic study of human, dog, and mouse.
    PLoS Comput Biol. 2007 Jan 5;3(1):e2 PMID: 17206860
  2. Towards a knowledge-based Human Protein Atlas.
    Nat Biotechnol. 2010 Dec;28(12):1248-50 PMID: 21139605
  3. A high-resolution map of human evolutionary constraint using 29 mammals.
    Nature. 2011 Oct 12;478(7370):476-82 PMID: 21993624
  4. The quantitative proteomes of human-induced pluripotent stem cells and embryonic stem cells.
    Mol Syst Biol. 2011 Nov 22;7:550 PMID: 22108792
  5. GENCODE: the reference human genome annotation for The ENCODE Project.
    Genome Res. 2012 Sep;22(9):1760-74 PMID: 22955987
  6. Combining quantitative proteomics data processing workflows for greater sensitivity.
    Nat Methods. 2011 Jun;8(6):481-3 PMID: 21552256
  7. Proteomics: a pragmatic perspective.
    Nat Biotechnol. 2010 Jul;28(7):695-709 PMID: 20622844
  8. Database resources of the National Center for Biotechnology Information.
    Nucleic Acids Res. 2014 Jan;42(Database issue):D7-17 PMID: 24259429
  9. Best alpha-helical transmembrane protein topology predictions are achieved using hidden Markov models and evolutionary information.
    Protein Sci. 2004 Jul;13(7):1908-17 PMID: 15215532
  10. Deep proteome and transcriptome mapping of a human cancer cell line.
    Mol Syst Biol. 2011 Nov 08;7:548 PMID: 22068331
  11. Age-dependent gain of alternative splice forms and biased duplication explain the relation between splicing and duplication.
    Genome Res. 2011 Mar;21(3):357-63 PMID: 21173032
  12. Comparative proteomics reveals a significant bias toward alternative protein isoforms with conserved structure and function.
    Mol Biol Evol. 2012 Sep;29(9):2265-83 PMID: 22446687
  13. The state of the human proteome in 2012 as viewed through PeptideAtlas.
    J Proteome Res. 2013 Jan 4;12(1):162-71 PMID: 23215161
  14. Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project.
    Nature. 2007 Jun 14;447(7146):799-816 PMID: 17571346
  15. GENCODE: producing a reference annotation for ENCODE.
    Genome Biol. 2006;7 Suppl 1:S4.1-9 PMID: 16925838
  16. Locating proteins in the cell using TargetP, SignalP and related tools.
    Nat Protoc. 2007;2(4):953-71 PMID: 17446895
  17. Systematic and integrative analysis of large gene lists using DAVID bioinformatics resources.
    Nat Protoc. 2009;4(1):44-57 PMID: 19131956
  18. Update on activities at the Universal Protein Resource (UniProt) in 2013.
    Nucleic Acids Res. 2013 Jan;41(Database issue):D43-7 PMID: 23161681
  19. Kalign--an accurate and fast multiple sequence alignment algorithm.
    BMC Bioinformatics. 2005 Dec 12;6:298 PMID: 16343337
  20. Mass spectrometry-based proteomics.
    Nature. 2003 Mar 13;422(6928):198-207 PMID: 12634793
  21. Human genome. A low number wins the GeneSweep Pool.
    Science. 2003 Jun 6;300(5625):1484 PMID: 12791949
  22. Aligning multiple genomic sequences with the threaded blockset aligner.
    Genome Res. 2004 Apr;14(4):708-15 PMID: 15060014
  23. EnsemblCompara GeneTrees: Complete, duplication-aware phylogenetic trees in vertebrates.
    Genome Res. 2009 Feb;19(2):327-35 PMID: 19029536
  24. Qscore: an algorithm for evaluating SEQUEST database search results.
    J Am Soc Mass Spectrom. 2002 Apr;13(4):378-86 PMID: 11951976
  25. The Ensembl genome database project.
    Nucleic Acids Res. 2002 Jan 1;30(1):38-41 PMID: 11752248
  26. Robust prediction of the MASCOT score for an improved quality assessment in mass spectrometric proteomics.
    J Proteome Res. 2008 Sep;7(9):3708-17 PMID: 18707158
  27. Late-replicating CNVs as a source of new genes.
    Biol Open. 2013 Dec 15;2(12):1402-11 PMID: 24285712
  28. The RCSB Protein Data Bank: redesigned web site and web services.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D392-401 PMID: 21036868
  29. A method for reducing the time required to match protein sequences with tandem mass spectra.
    Rapid Commun Mass Spectrom. 2003;17(20):2310-6 PMID: 14558131
  30. A combined transmembrane topology and signal peptide prediction method.
    J Mol Biol. 2004 May 14;338(5):1027-36 PMID: 15111065
  31. Protein synthesis rate is the predominant regulator of protein expression during differentiation.
    Mol Syst Biol. 2013;9:689 PMID: 24045637
  32. High performance computational analysis of large-scale proteome data sets to assess incremental contribution to coverage of the human genome.
    J Proteome Res. 2013 Jun 7;12(6):2858-68 PMID: 23611042
  33. Andromeda: a peptide search engine integrated into the MaxQuant environment.
    J Proteome Res. 2011 Apr 1;10(4):1794-805 PMID: 21254760
  34. Consequences of the discontinuation of the International Protein Index (IPI) database and its substitution by the UniProtKB "complete proteome" sets.
    Proteomics. 2011 Nov;11(22):4434-8 PMID: 21932440
  35. A phylostratigraphy approach to uncover the genomic history of major adaptations in metazoan lineages.
    Trends Genet. 2007 Nov;23(11):533-9 PMID: 18029048
  36. firestar--advances in the prediction of functionally important residues.
    Nucleic Acids Res. 2011 Jul;39(Web Server issue):W235-41 PMID: 21672959
  37. Metrics for the Human Proteome Project 2013-2014 and strategies for finding missing proteins.
    J Proteome Res. 2014 Jan 3;13(1):15-20 PMID: 24364385
  38. Ensembl 2013.
    Nucleic Acids Res. 2013 Jan;41(Database issue):D48-55 PMID: 23203987
  39. An integrated encyclopedia of DNA elements in the human genome.
    Nature. 2012 Sep 6;489(7414):57-74 PMID: 22955616
  40. PASSEL: the PeptideAtlas SRMexperiment library.
    Proteomics. 2012 Apr;12(8):1170-5 PMID: 22318887
  41. APPRIS: annotation of principal and alternative splice isoforms.
    Nucleic Acids Res. 2013 Jan;41(Database issue):D110-7 PMID: 23161672
  42. Improving gene annotation using peptide mass spectrometry.
    Genome Res. 2007 Feb;17(2):231-9 PMID: 17189379
  43. Ensembl 2011.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D800-6 PMID: 21045057
  44. Detecting amino acid sites under positive selection and purifying selection.
    Genetics. 2005 Mar;169(3):1753-62 PMID: 15654091
  45. The quantitative proteome of a human cell line.
    Mol Syst Biol. 2011 Nov 08;7:549 PMID: 22068332
  46. Shotgun proteomics aids discovery of novel protein-coding genes, alternative splicing, and "resurrected" pseudogenes in the mouse genome.
    Genome Res. 2011 May;21(5):756-67 PMID: 21460061
  47. The sequence of the human genome.
    Science. 2001 Feb 16;291(5507):1304-51 PMID: 11181995
  48. Distinguishing protein-coding and noncoding genes in the human genome.
    Proc Natl Acad Sci U S A. 2007 Dec 4;104(49):19428-33 PMID: 18040051
  49. The human phylome.
    Genome Biol. 2007;8(6):R109 PMID: 17567924
  50. The Pfam protein families database.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D290-301 PMID: 22127870
  51. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  52. An algorithm for progressive multiple alignment of sequences with insertions.
    Proc Natl Acad Sci U S A. 2005 Jul 26;102(30):10557-62 PMID: 16000407
  53. H-InvDB in 2013: an omics study platform for human functional gene and transcript discovery.
    Nucleic Acids Res. 2013 Jan;41(Database issue):D915-9 PMID: 23197657
  54. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  55. Comparative proteomic analysis of eleven common cell lines reveals ubiquitous but varying expression of most proteins.
    Mol Cell Proteomics. 2012 Mar;11(3):M111.014050 PMID: 22278370
  56. Comprehensive genome-wide proteomic analysis of human placental tissue for the Chromosome-Centric Human Proteome Project.
    J Proteome Res. 2013 Jun 7;12(6):2458-66 PMID: 23362793
  57. Open source system for analyzing, validating, and storing protein identification data.
    J Proteome Res. 2004 Nov-Dec;3(6):1234-42 PMID: 15595733
  58. Finishing the euchromatic sequence of the human genome.
    Nature. 2004 Oct 21;431(7011):931-45 PMID: 15496913
  59. EGASP: the human ENCODE Genome Annotation Assessment Project.
    Genome Biol. 2006;7 Suppl 1:S2.1-31 PMID: 16925836
  60. The potentially deleterious functional variant flavin-containing monooxygenase 2*1 is at high frequency throughout sub-Saharan Africa.
    Pharmacogenet Genomics. 2008 Oct;18(10):877-86 PMID: 18794725
  61. Quantifying the mechanisms of domain gain in animal proteins.
    Genome Biol. 2010;11(7):R74 PMID: 20633280
  62. Improving the accuracy of transmembrane protein topology prediction using evolutionary information.
    Bioinformatics. 2007 Mar 1;23(5):538-44 PMID: 17237066
Article Info
Journal
Human molecular genetics
Abbr.
Hum Mol Genet
ISSN
1460-2083
Published
2014-11-15
Epub
2014-00-16
Pages
5866-78
Language
English
Region
England
NLM ID
9208958
PMCID
PMC4204768
Subset
IM
Grants
NHGRI NIH HHS · U41 HG007234 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com