Home LiteratureArticle Details
PMID: 9223186 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, P.H.S.

Pfam: a comprehensive database of protein domain families based on seed alignments.

Proteins ·Vol. 28 ·No. 3 ·1997-07-00 ·Pages 405-20

Sonnhammer EL, Eddy SR, Durbin R

Abstract

Databases of multiple sequence alignments are a valuable aid to protein sequence classification and analysis. One of the main challenges when constructing such a database is to simultaneously satisfy the conflicting demands of completeness on the one hand and quality of alignment and domain definitions on the other. The latter properties are best dealt with by manual approaches, whereas completeness in practice is only amenable to automatic methods. Herein we present a database based on hidden Markov model profiles (HMMs), which combines high quality and completeness. Our database, Pfam, consists of parts A and B. Pfam-A is curated and contains well-characterized protein domain families with high quality alignments, which are maintained by using manually checked seed alignments and HMMs to find and align all members. Pfam-B contains sequence families that were generated automatically by applying the Domainer algorithm to cluster and align the remaining protein sequences after removal of Pfam-A domains. By using Pfam, a large number of previously unannotated proteins from the Caenorhabditis elegans genome project were classified. We have also identified many novel family memberships in known proteins, including new kazal, Fibronectin type III, and response regulator receiver domains. Pfam-A families have permanent accession numbers and form a library of HMMs available for searching and automatic annotation of new protein sequences.

MeSH Terms
Amino Acid Sequence Databases, Factual Models, Chemical Molecular Sequence Data Multigene Family Plant Proteins/chemistry Protein Structure, Tertiary Seeds/chemistry Sequence Alignment Sequence Homology, Amino Acid
Chemicals
Plant Proteins
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Sonnhammer E L
Sanger Centre, Wellcome Trust Genome Campus, Hinxton, Cambridge, United Kingdom.
Eddy S R
Durbin R
Article Info
Journal
Proteins
Abbr.
Proteins
ISSN
0887-3585
Published
1997-07-00
Pages
405-20
Language
English
Region
United States
NLM ID
8700181
Subset
IM
Grants
NHGRI NIH HHS · HG01363 · United States
Wellcome Trust · United Kingdom
Databases
SWISSPROT
P35475, P38938, P46720, P46721, P48441, Q00910, Q01634
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com