Abstract
The growth of sequencing-based Chromatin Immuno-Precipitation studies call for a more in-depth understanding of the nature of the technology and of the resultant data to reduce false positives and false negatives. Control libraries are typically constructed to complement such studies in order to mitigate the effect of systematic biases that might be present in the data. In this study, we explored multiple control libraries to obtain better understanding of what they truly represent. First, we analyzed the genome-wide profiles of various sequencing-based libraries at a low resolution of 1 Mbp, and compared them with each other as well as against aCGH data. We found that copy number plays a major influence in both ChIP-enriched as well as control libraries. Following that, we inspected the repeat regions to assess the extent of mapping bias. Next, significantly tag-rich 5 kbp regions were identified and they were associated with various genomic landmarks. For instance, we discovered that gene boundaries were surprisingly enriched with sequenced tags. Further, profiles between different cell types were noticeably distinct although the cell types were somewhat related and similar. We found that control libraries bear traces of systematic biases. The biases can be attributed to genomic copy number, inherent sequencing bias, plausible mapping ambiguity, and cell-type specific chromatin structure. Our results suggest careful analysis of control libraries can reveal promising biological insights.
MeSH Terms
Animals
Base Sequence
Cell Line
Cells
Chromatin
Chromatin Immunoprecipitation/methods
Chromosome Mapping
Codon, Terminator
Gene Dosage
Gene Library
Genes
Genome
Genomics/methods
Mathematical Concepts
Mice
Repetitive Sequences, Nucleic Acid
Sequence Analysis, DNA/methods
Chemicals
Chromatin
Codon, Terminator
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Vega Vinsensius B
Computational and Mathematical Biology Group, Genome Institute of Singapore, Singapore, Singapore.
Cheung Edwin
Palanisamy Nallasivam
Sung Wing-Kin
References (14)
14 references, click to expand
-
Whole-genome sequencing and variant discovery in C. elegans.
Nat Methods. 2008 Feb;5(2):183-8
PMID: 18204455
-
Mapping the chromosomal targets of STAT1 by Sequence Tag Analysis of Genomic Enrichment (STAGE).
Genome Res. 2007 Jun;17(6):910-6
PMID: 17568006
-
Genome-wide maps of chromatin state in pluripotent and lineage-committed cells.
Nature. 2007 Aug 2;448(7153):553-60
PMID: 17603471
-
The UCSC Genome Browser Database.
Nucleic Acids Res. 2003 Jan 1;31(1):51-4
PMID: 12519945
-
Whole-genome cartography of estrogen receptor alpha binding sites.
PLoS Genet. 2007 Jun;3(6):e87
PMID: 17542648
-
Model-based analysis of ChIP-Seq (MACS).
Genome Biol. 2008;9(9):R137
PMID: 18798982
-
Niche-independent symmetrical self-renewal of a mammalian tissue stem cell.
PLoS Biol. 2005 Sep;3(9):e283
PMID: 16086633
-
Defining the CREB regulon: a genome-wide analysis of transcription factor regulatory regions.
Cell. 2004 Dec 29;119(7):1041-54
PMID: 15620361
-
Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
Nucleic Acids Res. 2008 Sep;36(16):e105
PMID: 18660515
-
Evolution of the mammalian transcription factor binding repertoire via transposable elements.
Genome Res. 2008 Nov;18(11):1752-62
PMID: 18682548
-
Dynamic regulation of nucleosome positioning in the human genome.
Cell. 2008 Mar 7;132(5):887-98
PMID: 18329373
-
Integration of external signaling pathways with the core transcriptional network in embryonic stem cells.
Cell. 2008 Jun 13;133(6):1106-17
PMID: 18555785
-
A global map of p53 transcription-factor binding sites in the human genome.
Cell. 2006 Jan 13;124(1):207-19
PMID: 16413492
-
Genome-wide mapping of in vivo protein-DNA interactions.
Science. 2007 Jun 8;316(5830):1497-502
PMID: 17540862