Home LiteratureArticle Details
PMID: 19920124 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

The Pfam protein families database.

Nucleic acids research ·Vol. 38 ·No. Database issue ·2010-01-00 ·Pages D211-22

Finn RD, Mistry J, Tate J, Coggill P, Heger A, Pollington JE, Gavin OL, Gunasekaran P, Ceric G, Forslund K, Holm L, Sonnhammer EL, Eddy SR, Bateman A

Abstract

Pfam is a widely used database of protein families and domains. This article describes a set of major updates that we have implemented in the latest release (version 24.0). The most important change is that we now use HMMER3, the latest version of the popular profile hidden Markov model package. This software is approximately 100 times faster than HMMER2 and is more sensitive due to the routine use of the forward algorithm. The move to HMMER3 has necessitated numerous changes to Pfam that are described in detail. Pfam release 24.0 contains 11,912 families, of which a large number have been significantly updated during the past two years. Pfam is available via servers in the UK (http://pfam.sanger.ac.uk/), the USA (http://pfam.janelia.org/) and Sweden (http://pfam.sbc.su.se/).

MeSH Terms
Amino Acid Sequence Animals Computational Biology/methods,trends Databases, Nucleic Acid Databases, Protein Genome, Archaeal Genome, Fungal Humans Information Storage and Retrieval/methods Internet Molecular Sequence Data Protein Structure, Tertiary Sequence Alignment Sequence Homology, Amino Acid Software
Authors & Affiliations
14 authors, click to expand affiliations / ORCID
Finn Robert D
Wellcome Trust Sanger Institute, Wellcome Trust Genome Campus, Hinxton, Cambridgeshire CB10 1SA, UK. rdf@sanger.ac.uk
Mistry Jaina
Tate John
Coggill Penny
Heger Andreas
Pollington Joanne E
Gavin O Luke
Gunasekaran Prasad
Ceric Goran
Forslund Kristoffer
Holm Liisa
Sonnhammer Erik L L
Eddy Sean R
Bateman Alex
References (24)
24 references, click to expand
  1. Profile Comparer: a program for scoring and aligning profile hidden Markov models.
    Bioinformatics. 2008 Nov 15;24(22):2630-1 PMID: 18845584
  2. Remediation of the protein data bank archive.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D426-33 PMID: 18073189
  3. Predicting active site residue annotations in the Pfam database.
    BMC Bioinformatics. 2007 Aug 09;8:298 PMID: 17688688
  4. The Universal Protein Resource (UniProt): an expanding universe of protein information.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D187-91 PMID: 16381842
  5. ProDom and ProDom-CG: tools for protein domain analysis and whole genome comparisons.
    Nucleic Acids Res. 2000 Jan 1;28(1):267-9 PMID: 10592243
  6. QuickTree: building huge Neighbour-Joining trees of protein sequences.
    Bioinformatics. 2002 Nov;18(11):1546-7 PMID: 12424131
  7. BLAT--the BLAST-like alignment tool.
    Genome Res. 2002 Apr;12(4):656-64 PMID: 11932250
  8. The Protein Feature Ontology: a tool for the unification of protein feature annotations.
    Bioinformatics. 2008 Dec 1;24(23):2767-72 PMID: 18936051
  9. Rfam: updates to the RNA families database.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D136-40 PMID: 18953034
  10. HMM Logos for visualization of protein families.
    BMC Bioinformatics. 2004 Jan 21;5:7 PMID: 14736340
  11. FastTree: computing large minimum evolution trees with profiles instead of a distance matrix.
    Mol Biol Evol. 2009 Jul;26(7):1641-50 PMID: 19377059
  12. ADDA: a domain database with global coverage of the protein universe.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D188-91 PMID: 15608174
  13. Pfam: a domain-centric method for analyzing proteins and proteomes.
    Methods Mol Biol. 2007;396:43-58 PMID: 18025685
  14. Protein homology detection by HMM-HMM comparison.
    Bioinformatics. 2005 Apr 1;21(7):951-60 PMID: 15531603
  15. PairsDB atlas of protein sequence space.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D276-80 PMID: 17986464
  16. InterPro: the integrative protein signature database.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D211-5 PMID: 18940856
  17. RSDB: representative protein sequence databases have high information content.
    Bioinformatics. 2000 May;16(5):458-64 PMID: 10871268
  18. Pfam: clans, web tools and services.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D247-51 PMID: 16381856
  19. TreeFam: 2008 Update.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D735-40 PMID: 18056084
  20. The Pfam protein families database.
    Nucleic Acids Res. 2000 Jan 1;28(1):263-6 PMID: 10592242
  21. SCOOP: a simple method for identification of novel protein superfamily relationships.
    Bioinformatics. 2007 Apr 1;23(7):809-14 PMID: 17277330
  22. BioLit: integrating biological literature with databases.
    Nucleic Acids Res. 2008 Jul 1;36(Web Server issue):W385-9 PMID: 18515836
  23. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  24. Data growth and its impact on the SCOP database: new developments.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D419-25 PMID: 18000004
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2010-01-00
Epub
2009-00-17
Pages
D211-22
Language
English
Region
England
NLM ID
0411011
PMCID
PMC2808889
Subset
IM
Grants
Medical Research Council · MC_U137761446 · United Kingdom
Howard Hughes Medical Institute · United States
Wellcome Trust · 087656 · United Kingdom
Wellcome Trust · WT077044/Z/05/Z · United Kingdom
Biotechnology and Biological Sciences Research Council · BB/F010435/1 · United Kingdom
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com