Home LiteratureArticle Details
PMID: 20459804 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

ChIPpeakAnno: a Bioconductor package to annotate ChIP-seq and ChIP-chip data.

BMC bioinformatics ·Vol. 11 ·2010-05-11 ·Pages 237

Zhu LJ, Gazin C, Lawson ND, Pagès H, Lin SM, Lapointe DS, Green MR

Abstract

Chromatin immunoprecipitation (ChIP) followed by high-throughput sequencing (ChIP-seq) or ChIP followed by genome tiling array analysis (ChIP-chip) have become standard technologies for genome-wide identification of DNA-binding protein target sites. A number of algorithms have been developed in parallel that allow identification of binding sites from ChIP-seq or ChIP-chip datasets and subsequent visualization in the University of California Santa Cruz (UCSC) Genome Browser as custom annotation tracks. However, summarizing these tracks can be a daunting task, particularly if there are a large number of binding sites or the binding sites are distributed widely across the genome. We have developed ChIPpeakAnno as a Bioconductor package within the statistical programming environment R to facilitate batch annotation of enriched peaks identified from ChIP-seq, ChIP-chip, cap analysis of gene expression (CAGE) or any experiments resulting in a large number of enriched genomic regions. The binding sites annotated with ChIPpeakAnno can be viewed easily as a table, a pie chart or plotted in histogram form, i.e., the distribution of distances to the nearest genes for each set of peaks. In addition, we have implemented functionalities for determining the significance of overlap between replicates or binding sites among transcription factors within a complex, and for drawing Venn diagrams to visualize the extent of the overlap between replicates. Furthermore, the package includes functionalities to retrieve sequences flanking putative binding sites for PCR amplification, cloning, or motif discovery, and to identify Gene Ontology (GO) terms associated with adjacent genes. ChIPpeakAnno enables batch annotation of the binding sites identified from ChIP-seq, ChIP-chip, CAGE or any technology that results in a large number of enriched genomic regions within the statistical programming environment R. Allowing users to pass their own annotation data such as a different Chromatin immunoprecipitation (ChIP) preparation and a dataset from literature, or existing annotation packages, such as GenomicFeatures and BSgenome, provides flexibility. Tight integration to the biomaRt package enables up-to-date annotation retrieval from the BioMart database.

MeSH Terms
Binding Sites Chromatin Immunoprecipitation/methods Genome Oligonucleotide Array Sequence Analysis/methods Software
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Zhu Lihua J
Program in Gene Function and Expression, University of Massachusetts Medical School, Worcester, Massachusetts 01605, USA. julie.zhu@umassmed.edu
Gazin Claude
Lawson Nathan D
Pagès Hervé
Lin Simon M
Lapointe David S
Green Michael R
References (30)
30 references, click to expand
  1. Efficient yeast ChIP-Seq using multiplex short-read DNA sequencing.
    BMC Genomics. 2009 Jan 21;10:37 PMID: 19159457
  2. CEAS: cis-regulatory element annotation system.
    Bioinformatics. 2009 Oct 1;25(19):2605-6 PMID: 19689956
  3. rtracklayer: an R package for interfacing with genome browsers.
    Bioinformatics. 2009 Jul 15;25(14):1841-2 PMID: 19468054
  4. GenomeGraphs: integrated genomic data visualization with R.
    BMC Bioinformatics. 2009 Jan 06;10:2 PMID: 19123956
  5. Model-based analysis of ChIP-Seq (MACS).
    Genome Biol. 2008;9(9):R137 PMID: 18798982
  6. EnsMart: a generic system for fast and flexible access to biological data.
    Genome Res. 2004 Jan;14(1):160-9 PMID: 14707178
  7. CARPET: a web-based package for the analysis of ChIP-chip and expression tiling data.
    Bioinformatics. 2008 Dec 15;24(24):2918-20 PMID: 18945685
  8. Genome-wide profiles of STAT1 DNA association using chromatin immunoprecipitation and massively parallel sequencing.
    Nat Methods. 2007 Aug;4(8):651-7 PMID: 17558387
  9. Ensembl 2005.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D447-53 PMID: 15608235
  10. Statistics for ChIP-chip and DNase hypersensitivity experiments on NimbleGen arrays.
    Methods Enzymol. 2006;411:270-82 PMID: 16939795
  11. Probabilistic base calling of Solexa sequencing data.
    BMC Bioinformatics. 2008 Oct 13;9:431 PMID: 18851737
  12. An integrated software system for analyzing ChIP-chip and ChIP-seq data.
    Nat Biotechnol. 2008 Nov;26(11):1293-300 PMID: 18978777
  13. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  14. BayesPeak: Bayesian analysis of ChIP-seq data.
    BMC Bioinformatics. 2009 Sep 21;10:299 PMID: 19772557
  15. Ringo--an R/Bioconductor package for analyzing ChIP-chip readouts.
    BMC Bioinformatics. 2007 Jun 26;8:221 PMID: 17594472
  16. Fitting a mixture model by expectation maximization to discover motifs in biopolymers.
    Proc Int Conf Intell Syst Mol Biol. 1994;2:28-36 PMID: 7584402
  17. A feature-based approach to modeling protein-DNA interactions.
    PLoS Comput Biol. 2008 Aug 22;4(8):e1000154 PMID: 18725950
  18. PeakSeq enables systematic scoring of ChIP-seq experiments relative to controls.
    Nat Biotechnol. 2009 Jan;27(1):66-75 PMID: 19122651
  19. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data.
    Bioinformatics. 2010 Jan 1;26(1):139-40 PMID: 19910308
  20. ShortRead: a bioconductor package for input, quality assessment and exploration of high-throughput sequence data.
    Bioinformatics. 2009 Oct 1;25(19):2607-8 PMID: 19654119
  21. Bioconductor: open software development for computational biology and bioinformatics.
    Genome Biol. 2004;5(10):R80 PMID: 15461798
  22. Systematic evaluation of variability in ChIP-chip experiments using predefined DNA targets.
    Genome Res. 2008 Mar;18(3):393-403 PMID: 18258921
  23. DEGseq: an R package for identifying differentially expressed genes from RNA-seq data.
    Bioinformatics. 2010 Jan 1;26(1):136-8 PMID: 19855105
  24. Genome-wide mapping of in vivo protein-DNA interactions.
    Science. 2007 Jun 8;316(5830):1497-502 PMID: 17540862
  25. Modeling ChIP sequencing in silico with applications.
    PLoS Comput Biol. 2008 Aug 22;4(8):e1000158 PMID: 18725927
  26. FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technology.
    Bioinformatics. 2008 Aug 1;24(15):1729-30 PMID: 18599518
  27. Linear models and empirical bayes methods for assessing differential expression in microarray experiments.
    Stat Appl Genet Mol Biol. 2004;3:Article3 PMID: 16646809
  28. MAMMOT--a set of tools for the design, management and visualization of genomic tiling arrays.
    Bioinformatics. 2006 Apr 1;22(7):883-4 PMID: 16452111
  29. BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis.
    Bioinformatics. 2005 Aug 15;21(16):3439-40 PMID: 16082012
  30. Genome-wide analysis of transcription factor binding sites based on ChIP-Seq data.
    Nat Methods. 2008 Sep;5(9):829-34 PMID: 19160518
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2010-05-11
Epub
2010-00-11
Pages
237
Language
English
Region
England
NLM ID
100965194
PMCID
PMC3098059
Subset
IM
Grants
NIGMS NIH HHS · R01 GM033977 · United States
NHLBI NIH HHS · R01 HL093467 · United States
NHLBI NIH HHS · R01 HL093467-03 · United States
NHLBI NIH HHS · R01 HL093766 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com