Home LiteratureArticle Details
PMID: 20628599 Published · epublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S.

Evaluation of algorithm performance in ChIP-seq peak detection.

PloS one ·Vol. 5 ·No. 7 ·2010-07-08 ·Pages e11471

Wilbanks EG, Facciotti MT

Abstract

Next-generation DNA sequencing coupled with chromatin immunoprecipitation (ChIP-seq) is revolutionizing our ability to interrogate whole genome protein-DNA interactions. Identification of protein binding sites from ChIP-seq data has required novel computational tools, distinct from those used for the analysis of ChIP-Chip experiments. The growing popularity of ChIP-seq spurred the development of many different analytical programs (at last count, we noted 31 open source methods), each with some purported advantage. Given that the literature is dense and empirical benchmarking challenging, selecting an appropriate method for ChIP-seq analysis has become a daunting task. Herein we compare the performance of eleven different peak calling programs on common empirical, transcription factor datasets and measure their sensitivity, accuracy and usability. Our analysis provides an unbiased critical assessment of available technologies, and should assist researchers in choosing a suitable tool for handling ChIP-seq data.

MeSH Terms
Algorithms Chromatin Immunoprecipitation/methods Sequence Analysis, DNA/methods
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Wilbanks Elizabeth G
Graduate Group in Microbiology, University of California Davis, Davis, California, United States of America.
Facciotti Marc T
References (49)
49 references, click to expand
  1. An effective approach for identification of in vivo protein-DNA binding sites from paired-end ChIP-Seq data.
    BMC Bioinformatics. 2010 Feb 09;11:81 PMID: 20144209
  2. ChIPOTle: a user-friendly tool for the analysis of ChIP-chip data.
    Genome Biol. 2005;6(11):R97 PMID: 16277752
  3. F-Seq: a feature density estimator for high-throughput sequence tags.
    Bioinformatics. 2008 Nov 1;24(21):2537-8 PMID: 18784119
  4. BayesPeak: Bayesian analysis of ChIP-seq data.
    BMC Bioinformatics. 2009 Sep 21;10:299 PMID: 19772557
  5. TileMap: create chromosomal map of tiling array hybridizations.
    Bioinformatics. 2005 Sep 15;21(18):3629-36 PMID: 16046496
  6. Comparative study on ChIP-seq data: normalization and binding pattern characterization.
    Bioinformatics. 2009 Sep 15;25(18):2334-40 PMID: 19561022
  7. Integration of external signaling pathways with the core transcriptional network in embryonic stem cells.
    Cell. 2008 Jun 13;133(6):1106-17 PMID: 18555785
  8. ChromaSig: a probabilistic approach to finding common chromatin signatures in the human genome.
    PLoS Comput Biol. 2008 Oct;4(10):e1000201 PMID: 18927605
  9. Genome-wide profiles of STAT1 DNA association using chromatin immunoprecipitation and massively parallel sequencing.
    Nat Methods. 2007 Aug;4(8):651-7 PMID: 17558387
  10. FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technology.
    Bioinformatics. 2008 Aug 1;24(15):1729-30 PMID: 18599518
  11. Model-based analysis of tiling-arrays for ChIP-chip.
    Proc Natl Acad Sci U S A. 2006 Aug 15;103(33):12457-62 PMID: 16895995
  12. MEME SUITE: tools for motif discovery and searching.
    Nucleic Acids Res. 2009 Jul;37(Web Server issue):W202-8 PMID: 19458158
  13. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  14. An integrated software system for analyzing ChIP-chip and ChIP-seq data.
    Nat Biotechnol. 2008 Nov;26(11):1293-300 PMID: 18978777
  15. Sole-Search: an integrated analysis program for peak detection and functional annotation using ChIP-seq data.
    Nucleic Acids Res. 2010 Jan;38(3):e13 PMID: 19906703
  16. Core transcriptional regulatory circuitry in human embryonic stem cells.
    Cell. 2005 Sep 23;122(6):947-56 PMID: 16153702
  17. Genomic location analysis by ChIP-Seq.
    J Cell Biochem. 2009 May 1;107(1):11-8 PMID: 19173299
  18. ChIP-seq: advantages and challenges of a maturing technology.
    Nat Rev Genet. 2009 Oct;10(10):669-80 PMID: 19736561
  19. Next-generation gap.
    Nat Methods. 2009 Nov;6(11 Suppl):S2-5 PMID: 19844227
  20. Genome-wide identification of in vivo protein-DNA binding sites from ChIP-Seq data.
    Nucleic Acids Res. 2008 Sep;36(16):5221-31 PMID: 18684996
  21. Inherent signals in sequencing-based Chromatin-ImmunoPrecipitation control libraries.
    PLoS One. 2009;4(4):e5241 PMID: 19367334
  22. Design and analysis of ChIP-seq experiments for DNA-binding proteins.
    Nat Biotechnol. 2008 Dec;26(12):1351-9 PMID: 19029915
  23. PeakSeq enables systematic scoring of ChIP-seq experiments relative to controls.
    Nat Biotechnol. 2009 Jan;27(1):66-75 PMID: 19122651
  24. Comparative genomics modeling of the NRSF/REST repressor network: from single conserved sites to genome-wide repertoire.
    Genome Res. 2006 Oct;16(10):1208-21 PMID: 16963704
  25. HPeak: an HMM-based algorithm for defining read-enriched regions in ChIP-Seq data.
    BMC Bioinformatics. 2010 Jul 02;11:369 PMID: 20598134
  26. Genome-wide uH2A localization analysis highlights Bmi1-dependent deposition of the mark at repressed genes.
    PLoS Genet. 2009 Jun;5(6):e1000506 PMID: 19503595
  27. High-resolution profiling of histone methylations in the human genome.
    Cell. 2007 May 18;129(4):823-37 PMID: 17512414
  28. An HMM approach to genome-wide identification of differential histone modification sites from ChIP-seq data.
    Bioinformatics. 2008 Oct 15;24(20):2344-9 PMID: 18667444
  29. A practical comparison of methods for detecting transcription factor binding sites in ChIP-seq experiments.
    BMC Genomics. 2009 Dec 18;10:618 PMID: 20017957
  30. A signal-noise model for significance analysis of ChIP-seq with negative control.
    Bioinformatics. 2010 May 1;26(9):1199-204 PMID: 20371496
  31. Model-based analysis of ChIP-Seq (MACS).
    Genome Biol. 2008;9(9):R137 PMID: 18798982
  32. Empirical methods for controlling false positives and estimating confidence in ChIP-Seq peaks.
    BMC Bioinformatics. 2008 Dec 05;9:523 PMID: 19061503
  33. A high-resolution map of active promoters in the human genome.
    Nature. 2005 Aug 11;436(7052):876-80 PMID: 15988478
  34. The ets-related transcription factor GABP directs bidirectional transcription.
    PLoS Genet. 2007 Nov;3(11):e208 PMID: 18020712
  35. A clustering approach for identification of enriched domains from histone modification ChIP-Seq data.
    Bioinformatics. 2009 Aug 1;25(15):1952-8 PMID: 19505939
  36. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  37. A Poisson mixture model to identify changes in RNA polymerase II binding quantity using high-throughput sequencing technology.
    BMC Genomics. 2008 Sep 16;9 Suppl 2:S23 PMID: 18831789
  38. TRANSFAC: a database on transcription factors and their DNA binding sites.
    Nucleic Acids Res. 1996 Jan 1;24(1):238-41 PMID: 8594589
  39. Extracting transcription factor targets from ChIP-Seq data.
    Nucleic Acids Res. 2009 Sep;37(17):e113 PMID: 19553195
  40. Comparing genome-wide chromatin profiles using ChIP-chip or ChIP-seq.
    Bioinformatics. 2010 Apr 15;26(8):1000-6 PMID: 20208068
  41. Computation for ChIP-seq and RNA-seq studies.
    Nat Methods. 2009 Nov;6(11 Suppl):S22-32 PMID: 19844228
  42. High-resolution computational models of genome binding events.
    Nat Biotechnol. 2006 Aug;24(8):963-70 PMID: 16900145
  43. Genome-wide mapping of in vivo protein-DNA interactions.
    Science. 2007 Jun 8;316(5830):1497-502 PMID: 17540862
  44. Mapping accessible chromatin regions using Sono-Seq.
    Proc Natl Acad Sci U S A. 2009 Sep 1;106(35):14926-31 PMID: 19706456
  45. Modeling ChIP sequencing in silico with applications.
    PLoS Comput Biol. 2008 Aug 22;4(8):e1000158 PMID: 18725927
  46. BEDTools: a flexible suite of utilities for comparing genomic features.
    Bioinformatics. 2010 Mar 15;26(6):841-2 PMID: 20110278
  47. Genome-wide analysis of transcription factor binding sites based on ChIP-Seq data.
    Nat Methods. 2008 Sep;5(9):829-34 PMID: 19160518
  48. Model-based deconvolution of genome-wide DNA binding.
    Bioinformatics. 2008 Feb 1;24(3):396-403 PMID: 18056063
  49. A blind deconvolution approach to high-resolution mapping of transcription factor binding sites from ChIP-seq data.
    Genome Biol. 2009;10(12):R142 PMID: 20028542
Article Info
Journal
PloS one
Abbr.
PLoS One
ISSN
1932-6203
Published
2010-07-08
Epub
2010-00-08
Pages
e11471
Language
English
Region
United States
NLM ID
101285081
PMCID
PMC2900203
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com