Abstract
The ENCODE project is an international consortium with a goal of cataloguing all the functional elements in the human genome. The ENCODE Data Coordination Center (DCC) at the University of California, Santa Cruz serves as the central repository for ENCODE data. In this role, the DCC offers a collection of high-throughput, genome-wide data generated with technologies such as ChIP-Seq, RNA-Seq, DNA digestion and others. This data helps illuminate transcription factor-binding sites, histone marks, chromatin accessibility, DNA methylation, RNA expression, RNA binding and other cell-state indicators. It includes sequences with quality scores, alignments, signals calculated from the alignments, and in most cases, element or peak calls calculated from the signal data. Each data set is available for visualization and download via the UCSC Genome Browser (http://genome.ucsc.edu/). ENCODE data can also be retrieved using a metadata system that captures the experimental parameters of each assay. The ENCODE web portal at UCSC (http://encodeproject.org/) provides information about the ENCODE data and links for access.
MeSH Terms
Databases, Genetic
Gene Expression Regulation
Genome, Human
Genomics
Humans
Internet
Software
User-Computer Interface
Authors & Affiliations
23 authors, click to expand affiliations / ORCID
Raney Brian J
Center for Biomolecular Science and Engineering, School of Engineering and Howard Hughes Medical Institute, University of California Santa Cruz, Santa Cruz, CA 95064, USA. braney@soe.ucsc.edu
Cline Melissa S
Rosenbloom Kate R
Dreszer Timothy R
Learned Katrina
Barber Galt P
Meyer Laurence R
Sloan Cricket A
Malladi Venkat S
Roskin Krishna M
Suh Bernard B
Hinrichs Angie S
Clawson Hiram
Zweig Ann S
Kirkup Vanessa
Fujita Pauline A
Rhead Brooke
Smith Kayla E
Pohl Andy
Kuhn Robert M
Karolchik Donna
Haussler David
Kent W James
References (16)
16 references, click to expand
-
Incorporating sequence information into the scoring function: a hidden Markov model for improved peptide identification.
Bioinformatics. 2008 Mar 1;24(5):674-81
PMID: 18187442
-
Global mapping of protein-DNA interactions in vivo by digital genomic footprinting.
Nat Methods. 2009 Apr;6(4):283-9
PMID: 19305407
-
BigWig and BigBed: enabling browsing of large distributed datasets.
Bioinformatics. 2010 Sep 1;26(17):2204-7
PMID: 20639541
-
Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project.
Nature. 2007 Jun 14;447(7146):799-816
PMID: 17571346
-
ENCODE whole-genome data in the UCSC Genome Browser.
Nucleic Acids Res. 2010 Jan;38(Database issue):D620-5
PMID: 19920125
-
Unlocking the secrets of the genome.
Nature. 2009 Jun 18;459(7249):927-30
PMID: 19536255
-
Evolutionarily conserved and diverged alternative splicing events show different expression and functional profiles.
Nucleic Acids Res. 2005 Sep 29;33(17):5659-66
PMID: 16195578
-
Advances in RIP-chip analysis : RNA-binding protein immunoprecipitation-microarray profiling.
Methods Mol Biol. 2008;419:93-108
PMID: 18369977
-
GENCODE: producing a reference annotation for ENCODE.
Genome Biol. 2006;7 Suppl 1:S4.1-9
PMID: 16925838
-
Construction of a genome-scale structural map at single-nucleotide resolution.
Genome Res. 2007 Jun;17(6):947-53
PMID: 17568010
-
Evolution at two levels in humans and chimpanzees.
Science. 1975 Apr 11;188(4184):107-16
PMID: 1090005
-
Conserved expression without conserved regulatory sequence: the more things change, the more they stay the same.
Trends Genet. 2010 Feb;26(2):66-74
PMID: 20083321
-
The UCSC Genome Browser Database: update 2009.
Nucleic Acids Res. 2009 Jan;37(Database issue):D755-61
PMID: 18996895
-
The Sequence Alignment/Map format and SAMtools.
Bioinformatics. 2009 Aug 15;25(16):2078-9
PMID: 19505943
-
Independent functions of viral protein and nucleic acid in growth of bacteriophage.
J Gen Physiol. 1952 May;36(1):39-56
PMID: 12981234
-
The 1000 Genomes Project: new opportunities for research and social challenges.
Genome Med. 2010 Jan 21;2(1):3
PMID: 20193048