Home LiteratureArticle Details
PMID: 11752312 Published · ppublish English Journal Article

SUPERFAMILY: HMMs representing all proteins of known structure. SCOP sequence searches, alignments and genome assignments.

Nucleic acids research ·Vol. 30 ·No. 1 ·2002-01-01 ·Pages 268-72

Gough J, Chothia C

Abstract

The SUPERFAMILY database contains a library of hidden Markov models representing all proteins of known structure. The database is based on the SCOP 'superfamily' level of protein domain classification which groups together the most distantly related proteins which have a common evolutionary ancestor. There is a public server at http://supfam.org which provides three services: sequence searching, multiple alignments to sequences of known structure, and structural assignments to all complete genomes. Given an amino acid or nucleotide query sequence the server will return the domain architecture and SCOP classification. The server produces alignments of the query sequences with sequences of known structure, and includes multiple alignments of genome and PDB sequences. The structural assignments are carried out on all complete genomes (currently 59) covering approximately half of the soluble protein domains. The assignments, superfamily breakdown and statistics on them are available from the server. The database is currently used by this group and others for genome annotation, structural genomics, gene prediction and domain-based genomic studies.

MeSH Terms
Amino Acid Sequence Animals Base Sequence Databases, Protein Evolution, Molecular Genome Humans Information Storage and Retrieval Internet Molecular Sequence Data Protein Structure, Tertiary Proteins/chemistry,genetics Sequence Alignment
Chemicals
Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Gough Julian
MRC Laboratory of Molecular Biology, Hills Road, Cambridge CB2 2QH, UK. jgough@mrc-lmb.cam.ac.uk
Chothia Cyrus
References (10)
10 references, click to expand
  1. SMART: a web-based tool for the study of genetically mobile domains.
    Nucleic Acids Res. 2000 Jan 1;28(1):231-4 PMID: 10592234
  2. The Pfam protein families database.
    Nucleic Acids Res. 2000 Jan 1;28(1):263-6 PMID: 10592242
  3. InterPro--an integrated documentation resource for protein families, domains and functional sites.
    Bioinformatics. 2000 Dec;16(12):1145-50 PMID: 11159333
  4. Domain combinations in archaeal, eubacterial and eukaryotic proteomes.
    J Mol Biol. 2001 Jul 6;310(2):311-25 PMID: 11428892
  5. Hidden Markov models for detecting remote protein homologies.
    Bioinformatics. 1998;14(10):846-56 PMID: 9927713
  6. Assignment of homology to genome sequences using a library of hidden Markov models that represent all proteins of known structure.
    J Mol Biol. 2001 Nov 2;313(4):903-19 PMID: 11697912
  7. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  8. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  9. CATH--a hierarchic classification of protein domain structures.
    Structure. 1997 Aug 15;5(8):1093-108 PMID: 9309224
  10. The evolution and structural anatomy of the small molecule metabolic pathways in Escherichia coli.
    J Mol Biol. 2001 Aug 24;311(4):693-708 PMID: 11518524
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2002-01-01
Pages
268-72
Language
English
Region
England
NLM ID
0411011
PMCID
PMC99153
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com