Home LiteratureArticle Details
PMID: 24217909 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, N.I.H., Intramural Research Support, Non-U.S. Gov't

Current status and new features of the Consensus Coding Sequence database.

Nucleic acids research ·Vol. 42 ·No. Database issue ·2014-01-00 ·Pages D865-72

Farrell CM, O'Leary NA, Harte RA, Loveland JE, Wilming LG, Wallin C, Diekhans M, Barrell D, Searle SM, Aken B, Hiatt SM, Frankish A, Suner MM, Rajput B, Steward CA, Brown GR, Bennett R, Murphy M, Wu W, Kay MP, Hart J, Rajan J, Weber J, Snow C, Riddick LD, Hunt T, Webb D, Thomas M, Tamez P, Rangwala SH, McGarvey KM, Pujar S, Shkeda A, Mudge JM, Gonzalez JM, Gilbert JG, Trevanion SJ, Baertsch R, Harrow JL, Hubbard T, Ostell JM, Haussler D, Pruitt KD

Abstract

The Consensus Coding Sequence (CCDS) project (http://www.ncbi.nlm.nih.gov/CCDS/) is a collaborative effort to maintain a dataset of protein-coding regions that are identically annotated on the human and mouse reference genome assemblies by the National Center for Biotechnology Information (NCBI) and Ensembl genome annotation pipelines. Identical annotations that pass quality assurance tests are tracked with a stable identifier (CCDS ID). Members of the collaboration, who are from NCBI, the Wellcome Trust Sanger Institute and the University of California Santa Cruz, provide coordinated and continuous review of the dataset to ensure high-quality CCDS representations. We describe here the current status and recent growth in the CCDS dataset, as well as recent changes to the CCDS web and FTP sites. These changes include more explicit reporting about the NCBI and Ensembl annotation releases being compared, new search and display options, the addition of biologically descriptive information and our approach to representing genes for which support evidence is incomplete. We also present a summary of recent and future curation targets.

MeSH Terms
Animals Databases, Genetic Exons Genomics Humans Internet Mice Molecular Sequence Annotation Proteins/genetics Sequence Analysis
Chemicals
Proteins
Authors & Affiliations
43 authors, click to expand affiliations / ORCID
Farrell Catherine M
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Building 38A, 8600 Rockville Pike, Bethesda, MD 20894, USA, Center for Biomolecular Science and Engineering, University of California Santa Cruz (UCSC), Santa Cruz, CA 95064, USA, Wellcome Trust Sanger Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge, CB10 1SA, UK and Howard Hughes Medical Institute, University of California Santa Cruz, Santa Cruz, CA 95064, USA.
O'Leary Nuala A
Harte Rachel A
Loveland Jane E
Wilming Laurens G
Wallin Craig
Diekhans Mark
Barrell Daniel
Searle Stephen M J
Aken Bronwen
Hiatt Susan M
Frankish Adam
Suner Marie-Marthe
Rajput Bhanu
Steward Charles A
Brown Garth R
Bennett Ruth
Murphy Michael
Wu Wendy
Kay Mike P
Hart Jennifer
Rajan Jeena
Weber Janet
Snow Catherine
Riddick Lillian D
Hunt Toby
Webb David
Thomas Mark
Tamez Pamela
Rangwala Sanjida H
McGarvey Kelly M
Pujar Shashikant
Shkeda Andrei
Mudge Jonathan M
Gonzalez Jose M
Gilbert James G R
Trevanion Stephen J
Baertsch Robert
Harrow Jennifer L
Hubbard Tim
Ostell James M
Haussler David
Pruitt Kim D
References (24)
24 references, click to expand
  1. Genome-wide annotation and quantitation of translation by ribosome profiling.
    Curr Protoc Mol Biol. 2013 Jul;Chapter 4:Unit 4.18 PMID: 23821443
  2. 5' end-centered expression profiling using cap-analysis gene expression and next-generation sequencing.
    Nat Protoc. 2012 Feb 23;7(3):542-61 PMID: 22362160
  3. An analysis of exome sequencing for diagnostic testing of the genes associated with muscle disease and spastic paraplegia.
    Hum Mutat. 2012 Apr;33(4):614-26 PMID: 22311686
  4. The genetic landscape of mutations in Burkitt lymphoma.
    Nat Genet. 2012 Dec;44(12):1321-5 PMID: 23143597
  5. Quantifying single nucleotide variant detection sensitivity in exome sequencing.
    BMC Bioinformatics. 2013 Jun 18;14:195 PMID: 23773188
  6. NCBI Reference Sequences (RefSeq): current status, new features and genome annotation policy.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D130-5 PMID: 22121212
  7. L1 elements, processed pseudogenes and retrogenes in mammalian genomes.
    IUBMB Life. 2006 Dec;58(12):677-85 PMID: 17424906
  8. The International Nucleotide Sequence Database Collaboration.
    Nucleic Acids Res. 2013 Jan;41(Database issue):D21-4 PMID: 23180798
  9. The Genotype-Tissue Expression (GTEx) project.
    Nat Genet. 2013 Jun;45(6):580-5 PMID: 23715323
  10. GenBank.
    Nucleic Acids Res. 2013 Jan;41(Database issue):D36-42 PMID: 23193287
  11. The vertebrate genome annotation (Vega) database.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D753-60 PMID: 18003653
  12. Next-generation proteomics: towards an integrative view of proteome dynamics.
    Nat Rev Genet. 2013 Jan;14(1):35-48 PMID: 23207911
  13. Genome-wide DNA methylation profiling using Infinium® assay.
    Epigenomics. 2009 Oct;1(1):177-200 PMID: 22122642
  14. The consensus coding sequence (CCDS) project: Identifying a common protein-coding gene set for the human and mouse genomes.
    Genome Res. 2009 Jul;19(7):1316-23 PMID: 19498102
  15. A comparative analysis of exome capture.
    Genome Biol. 2011 Sep 29;12(9):R97 PMID: 21958622
  16. Ensembl 2013.
    Nucleic Acids Res. 2013 Jan;41(Database issue):D48-55 PMID: 23203987
  17. An integrated encyclopedia of DNA elements in the human genome.
    Nature. 2012 Sep 6;489(7414):57-74 PMID: 22955616
  18. Tracking and coordinating an international curation effort for the CCDS Project.
    Database (Oxford). 2012 Mar 20;2012:bas008 PMID: 22434842
  19. Modernizing reference genome assemblies.
    PLoS Biol. 2011 Jul;9(7):e1001091 PMID: 21750661
  20. The Universal Protein Resource (UniProt) in 2010.
    Nucleic Acids Res. 2010 Jan;38(Database issue):D142-8 PMID: 19843607
  21. Global mapping of translation initiation sites in mammalian cells at single-nucleotide resolution.
    Proc Natl Acad Sci U S A. 2012 Sep 11;109(37):E2424-32 PMID: 22927429
  22. GENCODE: the reference human genome annotation for The ENCODE Project.
    Genome Res. 2012 Sep;22(9):1760-74 PMID: 22955987
  23. NMD: a multifaceted response to premature translational termination.
    Nat Rev Mol Cell Biol. 2012 Nov;13(11):700-12 PMID: 23072888
  24. Locus Reference Genomic sequences: an improved basis for describing human DNA variants.
    Genome Med. 2010 Apr 15;2(4):24 PMID: 20398331
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2014-01-00
Epub
2013-00-11
Pages
D865-72
Language
English
Region
England
NLM ID
0411011
PMCID
PMC3965069
Subset
IM
Grants
NHGRI NIH HHS · 5U54HG00455-04 · United States
Howard Hughes Medical Institute · United States
Wellcome Trust · WT077198 · United Kingdom
Wellcome Trust · 095908 · United Kingdom
Intramural NIH HHS · United States
Wellcome Trust · WT062023 · United Kingdom
NHGRI NIH HHS · U54 HG004555 · United States
NHGRI NIH HHS · 10U41 HG007234-01 · United States
NHGRI NIH HHS · U41 HG007234 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com