Home LiteratureArticle Details
PMID: 26110515 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

Comparison of GENCODE and RefSeq gene annotation and the impact of reference geneset on variant effect prediction.

BMC genomics ·Vol. 16 Suppl 8 ·2015-00-00 ·Pages S2

Frankish A, Uszczynska B, Ritchie GR, Gonzalez JM, Pervouchine D, Petryszak R, Mudge JM, Fonseca N, Brazma A, Guigo R, Harrow J

Abstract

A vast amount of DNA variation is being identified by increasingly large-scale exome and genome sequencing projects. To be useful, variants require accurate functional annotation and a wide range of tools are available to this end. McCarthy et al recently demonstrated the large differences in prediction of loss-of-function (LoF) variation when RefSeq and Ensembl transcripts are used for annotation, highlighting the importance of the reference transcripts on which variant functional annotation is based. We describe a detailed analysis of the similarities and differences between the gene and transcript annotation in the GENCODE and RefSeq genesets. We demonstrate that the GENCODE Comprehensive set is richer in alternative splicing, novel CDSs, novel exons and has higher genomic coverage than RefSeq, while the GENCODE Basic set is very similar to RefSeq. Using RNAseq data we show that exons and introns unique to one geneset are expressed at a similar level to those common to both. We present evidence that the differences in gene annotation lead to large differences in variant annotation where GENCODE and RefSeq are used as reference transcripts, although this is predominantly confined to non-coding transcripts and UTR sequence, with at most ~30% of LoF variants annotated discordantly. We also describe an investigation of dominant transcript expression, showing that it both supports the utility of the GENCODE Basic set in providing a smaller set of more highly expressed transcripts and provides a useful, biologically-relevant filter for further reducing the complexity of the transcriptome. The reference transcripts selected for variant functional annotation do have a large effect on the outcome. The GENCODE Comprehensive transcripts contain more exons, have greater genomic coverage and capture many more variants than RefSeq in both genome and exome datasets, while the GENCODE Basic set shows a higher degree of concordance with RefSeq and has fewer unique features. We propose that the GENCODE Comprehensive set has great utility for the discovery of new variants with functional potential, while the GENCODE Basic set is more suitable for applications demanding less complex interpretation of functional variants.

MeSH Terms
Alternative Splicing Computational Biology Databases, Genetic Genome, Human Humans Molecular Sequence Annotation Protein Isoforms/genetics,metabolism Software Transcriptome
Chemicals
Protein Isoforms
Authors & Affiliations
11 authors, click to expand affiliations / ORCID
Frankish Adam
Uszczynska Barbara
Ritchie Graham R S
Gonzalez Jose M
Pervouchine Dmitri
Petryszak Robert
Mudge Jonathan M
Fonseca Nuno
Brazma Alvis
Guigo Roderic
Harrow Jennifer
References (35)
35 references, click to expand
  1. APPRIS: annotation of principal and alternative splice isoforms.
    Nucleic Acids Res. 2013 Jan;41(Database issue):D110-7 PMID: 23161672
  2. Disrupted post-transcriptional regulation of the cystic fibrosis transmembrane conductance regulator (CFTR) by a 5'UTR mutation is associated with a CFTR-related disease.
    Hum Mutat. 2011 Oct;32(10):E2266-82 PMID: 21837768
  3. Nonsense codons can reduce the abundance of nuclear mRNA without affecting the abundance of pre-mRNA or the half-life of cytoplasmic mRNA.
    Mol Cell Biol. 1993 Mar;13(3):1892-902 PMID: 8441420
  4. Analysis of 6,515 exomes reveals the recent origin of most human protein-coding variants.
    Nature. 2013 Jan 10;493(7431):216-20 PMID: 23201682
  5. Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project.
    Nature. 2007 Jun 14;447(7146):799-816 PMID: 17571346
  6. Whole exome sequencing of familial hypercholesterolaemia patients negative for LDLR/APOB/PCSK9 mutations.
    J Med Genet. 2014 Aug;51(8):537-44 PMID: 24987033
  7. An integrated map of genetic variation from 1,092 human genomes.
    Nature. 2012 Nov 1;491(7422):56-65 PMID: 23128226
  8. Evolution and functional impact of rare coding variation from deep sequencing of human exomes.
    Science. 2012 Jul 6;337(6090):64-9 PMID: 22604720
  9. Genome-wide search for exonic variants affecting translational efficiency.
    Nat Commun. 2013;4:2260 PMID: 23900168
  10. AceView: a comprehensive cDNA-supported gene and transcripts annotation.
    Genome Biol. 2006;7 Suppl 1:S12.1-14 PMID: 16925834
  11. The human genome browser at UCSC.
    Genome Res. 2002 Jun;12(6):996-1006 PMID: 12045153
  12. GENCODE: producing a reference annotation for ENCODE.
    Genome Biol. 2006;7 Suppl 1:S4.1-9 PMID: 16925838
  13. A 3'UTR polymorphism modulates mRNA stability of the oncogene and drug target Polo-like Kinase 1.
    Mol Cancer. 2014;13:87 PMID: 24767679
  14. An integrated encyclopedia of DNA elements in the human genome.
    Nature. 2012 Sep 6;489(7414):57-74 PMID: 22955616
  15. GENCODE: the reference human genome annotation for The ENCODE Project.
    Genome Res. 2012 Sep;22(9):1760-74 PMID: 22955987
  16. A method and server for predicting damaging missense mutations.
    Nat Methods. 2010 Apr;7(4):248-9 PMID: 20354512
  17. The Vertebrate Genome Annotation browser 10 years on.
    Nucleic Acids Res. 2014 Jan;42(Database issue):D771-9 PMID: 24316575
  18. A probabilistic disease-gene finder for personal genomes.
    Genome Res. 2011 Sep;21(9):1529-42 PMID: 21700766
  19. Widespread intron retention in mammals functionally tunes transcriptomes.
    Genome Res. 2014 Nov;24(11):1774-86 PMID: 25258385
  20. Combining RT-PCR-seq and RNA-seq to catalog all genic elements encoded in the human genome.
    Genome Res. 2012 Sep;22(9):1698-710 PMID: 22955982
  21. Sequence variants within the 3'-UTR of the COL5A1 gene alters mRNA stability: implications for musculoskeletal soft tissue injuries.
    Matrix Biol. 2011 Jun;30(5-6):338-45 PMID: 21609763
  22. Orchestrated intron retention regulates normal granulocyte differentiation.
    Cell. 2013 Aug 1;154(3):583-95 PMID: 23911323
  23. VAT: a computational framework to functionally annotate variants in personal genomes within a cloud-computing environment.
    Bioinformatics. 2012 Sep 1;28(17):2267-9 PMID: 22743228
  24. Alternative isoform regulation in human tissue transcriptomes.
    Nature. 2008 Nov 27;456(7221):470-6 PMID: 18978772
  25. PhyloCSF: a comparative genomics method to distinguish protein coding and non-coding regions.
    Bioinformatics. 2011 Jul 1;27(13):i275-82 PMID: 21685081
  26. RefSeq: an update on mammalian reference sequences.
    Nucleic Acids Res. 2014 Jan;42(Database issue):D756-63 PMID: 24259432
  27. PseudoPipe: an automated pseudogene identification pipeline.
    Bioinformatics. 2006 Jun 15;22(12):1437-9 PMID: 16574694
  28. Evolution's cauldron: duplication, deletion, and rearrangement in the mouse and human genomes.
    Proc Natl Acad Sci U S A. 2003 Sep 30;100(20):11484-9 PMID: 14500911
  29. Predicting the effects of coding non-synonymous variants on protein function using the SIFT algorithm.
    Nat Protoc. 2009;4(7):1073-81 PMID: 19561590
  30. Assembly information services in the European Nucleotide Archive.
    Nucleic Acids Res. 2014 Jan;42(Database issue):D38-43 PMID: 24214989
  31. Ensembl 2015.
    Nucleic Acids Res. 2015 Jan;43(Database issue):D662-9 PMID: 25352552
  32. Choice of transcripts and software has a large effect on variant annotation.
    Genome Med. 2014 Mar 31;6(3):26 PMID: 24944579
  33. Deriving the consequences of genomic variants with the Ensembl API and SNP Effect Predictor.
    Bioinformatics. 2010 Aug 15;26(16):2069-70 PMID: 20562413
  34. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data.
    Nucleic Acids Res. 2010 Sep;38(16):e164 PMID: 20601685
  35. Activities at the Universal Protein Resource (UniProt).
    Nucleic Acids Res. 2014 Jan;42(Database issue):D191-8 PMID: 24253303
Article Info
Journal
BMC genomics
Abbr.
BMC Genomics
ISSN
1471-2164
Published
2015-00-00
Epub
2015-00-18
Pages
S2
Language
English
Region
England
NLM ID
100965258
PMCID
PMC4502323
Subset
IM
Grants
NHGRI NIH HHS · U41 HG007234 · United States
NHGRI NIH HHS · U41 HG007000 · United States
NHGRI NIH HHS · U54 HG007004 · United States
Wellcome Trust · WT098051 · United Kingdom
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com