Abstract
High-throughput sequencing platforms are generating massive amounts of genetic variation data for diverse genomes, but it remains a challenge to pinpoint a small subset of functionally important variants. To fill these unmet needs, we developed the ANNOVAR tool to annotate single nucleotide variants (SNVs) and insertions/deletions, such as examining their functional consequence on genes, inferring cytogenetic bands, reporting functional importance scores, finding variants in conserved regions, or identifying variants reported in the 1000 Genomes Project and dbSNP. ANNOVAR can utilize annotation databases from the UCSC Genome Browser or any annotation data set conforming to Generic Feature Format version 3 (GFF3). We also illustrate a 'variants reduction' protocol on 4.7 million SNVs and indels from a human genome, including two causal mutations for Miller syndrome, a rare recessive disease. Through a stepwise procedure, we excluded variants that are unlikely to be causal, and identified 20 candidate genes including the causal gene. Using a desktop computer, ANNOVAR requires ∼4 min to perform gene-based annotation and ∼15 min to perform variants reduction on 4.7 million variants, making it practical to handle hundreds of human genomes in a day. ANNOVAR is freely available at http://www.openbioinformatics.org/annovar/.
MeSH Terms
Genes
Genetic Predisposition to Disease
Genetic Variation
Genome, Human
Genomics
High-Throughput Screening Assays
Humans
Sequence Analysis, DNA
Software
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Wang Kai
Center for Applied Genomics, Children's Hospital of Philadelphia, PA 19104, USA. kai@openbioinformatics.org
Li Mingyao
Hakonarson Hakon
References (21)
21 references, click to expand
-
NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
Nucleic Acids Res. 2007 Jan;35(Database issue):D61-5
PMID: 17130148
-
Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes.
Genome Res. 2005 Aug;15(8):1034-50
PMID: 16024819
-
Nomenclature for the description of human sequence variations.
Hum Genet. 2001 Jul;109(1):121-4
PMID: 11479744
-
WGAViewer: software for genomic annotation of whole genome association studies.
Genome Res. 2008 Apr;18(4):640-3
PMID: 18256235
-
Human non-synonymous SNPs: server and survey.
Nucleic Acids Res. 2002 Sep 1;30(17):3894-900
PMID: 12202775
-
Genome variation discovery with high-throughput sequencing data.
Brief Bioinform. 2010 Jan;11(1):3-14
PMID: 20053733
-
Exome sequencing identifies the cause of a mendelian disorder.
Nat Genet. 2010 Jan;42(1):30-5
PMID: 19915526
-
Detection of nonneutral substitution rates on mammalian phylogenies.
Genome Res. 2010 Jan;20(1):110-21
PMID: 19858363
-
Targeted capture and massively parallel sequencing of 12 human exomes.
Nature. 2009 Sep 10;461(7261):272-6
PMID: 19684571
-
High carrier frequency of the 35delG deafness mutation in European populations. Genetic Analysis Consortium of GJB2 35delG.
Eur J Hum Genet. 2000 Jan;8(1):19-23
PMID: 10713883
-
SIFT: Predicting amino acid changes that affect protein function.
Nucleic Acids Res. 2003 Jul 1;31(13):3812-4
PMID: 12824425
-
The UCSC Known Genes.
Bioinformatics. 2006 May 1;22(9):1036-46
PMID: 16500937
-
F-SNP: computationally predicted functional SNPs for disease association studies.
Nucleic Acids Res. 2008 Jan;36(Database issue):D820-4
PMID: 17986460
-
SCAN: SNP and copy number annotation.
Bioinformatics. 2010 Jan 15;26(2):259-62
PMID: 19933162
-
The UCSC Genome Browser database: update 2010.
Nucleic Acids Res. 2010 Jan;38(Database issue):D613-9
PMID: 19906737
-
The Ensembl automatic gene annotation system.
Genome Res. 2004 May;14(5):942-50
PMID: 15123590
-
Accurate whole human genome sequencing using reversible terminator chemistry.
Nature. 2008 Nov 6;456(7218):53-9
PMID: 18987734
-
Next generation tools for the annotation of human SNPs.
Brief Bioinform. 2009 Jan;10(1):35-52
PMID: 19181721
-
Segmental duplications: organization and impact within the current human genome project assembly.
Genome Res. 2001 Jun;11(6):1005-17
PMID: 11381028
-
Snap: an integrated SNP annotation platform.
Nucleic Acids Res. 2007 Jan;35(Database issue):D707-10
PMID: 17135198
-
How to map billions of short reads onto genomes.
Nat Biotechnol. 2009 May;27(5):455-7
PMID: 19430453