Home LiteratureArticle Details
PMID: 12520036 Published · ppublish English Journal Article Research Support, U.S. Gov't, P.H.S.

PEP: Predictions for Entire Proteomes.

Nucleic acids research ·Vol. 31 ·No. 1 ·2003-01-01 ·Pages 410-3

Carter P, Liu J, Rost B

Abstract

PEP is a database of Predictions for Entire Proteomes. The database contains summaries of analyses of protein sequences from a range of organisms representing all three major kingdoms of life: eukaryotes, prokaryotes and archaea. All proteins publicly available for organisms were aligned against SWISS-PROT, TrEMBL and PDB. Additionally, the following annotations are provided: secondary structure, transmembrane helices, coiled coils, regions of low complexity, signal peptides, PROSITE motifs, nuclear localization signals and classes of cellular function. Proteins that contain long regions without regular secondary structure are also identified. We have produced a related database of structural domain-like fragments derived from PEP and clusters based on homology between all fragments. The PEP database, fragments and clusters are distributed freely as a set of flat files and have been integrated into SRS. The PEP group of databases can be accessed from: http://cubic.bioc.columbia.edu/pep.

MeSH Terms
Animals Archaeal Proteins/chemistry Cluster Analysis Databases, Protein Eukaryotic Cells Humans Prokaryotic Cells Protein Conformation Proteins/chemistry Proteome/chemistry,physiology Sequence Homology, Amino Acid User-Computer Interface
Chemicals
Archaeal Proteins Proteins Proteome
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Carter Phil
CUBIC, Department of Biochemistry and Molecular Biophysics, Columbia University, 650 West 168th Street BB217, New York, NY 10032, USA. carter@cubic.bioc.columbia.edu
Liu Jinfeng
Rost Burkhard
References (22)
22 references, click to expand
  1. Review: protein secondary structure prediction continues to rise.
    J Struct Biol. 2001 May-Jun;134(2-3):204-18 PMID: 11551180
  2. Finding nuclear localization signals.
    EMBO Rep. 2000 Nov;1(5):411-5 PMID: 11258480
  3. Structural genomics: an approach to the protein folding problem.
    Proc Natl Acad Sci U S A. 2001 Nov 20;98(24):13488-9 PMID: 11717420
  4. The FlyBase database of the Drosophila genome projects and community literature.
    Nucleic Acids Res. 2002 Jan 1;30(1):106-8 PMID: 11752267
  5. The PROSITE database, its status in 2002.
    Nucleic Acids Res. 2002 Jan 1;30(1):235-8 PMID: 11752303
  6. Target space for structural genomics revisited.
    Bioinformatics. 2002 Jul;18(7):922-33 PMID: 12117789
  7. Did evolution leap to create the protein universe?
    Curr Opin Struct Biol. 2002 Jun;12(3):409-16 PMID: 12127462
  8. Loopy proteins appear conserved in evolution.
    J Mol Biol. 2002 Sep 6;322(1):53-64 PMID: 12215414
  9. NLSdb: database of nuclear localization signals.
    Nucleic Acids Res. 2003 Jan 1;31(1):397-9 PMID: 12520032
  10. Database of homology-derived protein structures and the structural meaning of sequence alignment.
    Proteins. 1991;9(1):56-68 PMID: 2017436
  11. Transforming a set of biological flat file libraries to a fast access network.
    Comput Appl Biosci. 1993 Feb;9(1):59-64 PMID: 8435769
  12. Local alignment statistics.
    Methods Enzymol. 1996;266:460-80 PMID: 8743700
  13. Prediction and analysis of coiled-coil structures.
    Methods Enzymol. 1996;266:513-25 PMID: 8743703
  14. PHD: predicting one-dimensional protein structure by profile-based neural networks.
    Methods Enzymol. 1996;266:525-39 PMID: 8743704
  15. Analysis of compositionally biased regions in sequence databases.
    Methods Enzymol. 1996;266:554-71 PMID: 8743706
  16. Identification of prokaryotic and eukaryotic signal peptides and prediction of their cleavage sites.
    Protein Eng. 1997 Jan;10(1):1-6 PMID: 9051728
  17. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  18. EUCLID: automatic classification of proteins in functional classes by their database annotations.
    Bioinformatics. 1998;14(6):542-3 PMID: 9694995
  19. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):45-8 PMID: 10592178
  20. The Protein Data Bank.
    Nucleic Acids Res. 2000 Jan 1;28(1):235-42 PMID: 10592235
  21. WormBase: network access to the genome and biology of Caenorhabditis elegans.
    Nucleic Acids Res. 2001 Jan 1;29(1):82-6 PMID: 11125056
  22. Comparing function and structure between entire proteomes.
    Protein Sci. 2001 Oct;10(10):1970-9 PMID: 11567088
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2003-01-01
Pages
410-3
Language
English
Region
England
NLM ID
0411011
PMCID
PMC165549
Subset
IM
Grants
NIGMS NIH HHS · P50 GM062413 · United States
NIGMS NIH HHS · R01 GM063029 · United States
NIGMS NIH HHS · 1-P50-GM62413-01 · United States
NIGMS NIH HHS · R01-GM63029-01 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com