Home LiteratureArticle Details
PMID: 17986464 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

PairsDB atlas of protein sequence space.

Nucleic acids research ·Vol. 36 ·No. Database issue ·2008-01-00 ·Pages D276-80

Heger A, Korpelainen E, Hupponen T, Mattila K, Ollikainen V, Holm L

Abstract

Sequence similarity/database searching is a cornerstone of molecular biology. PairsDB is a database intended to make exploring protein sequences and their similarity relationships quick and easy. Behind PairsDB is a comprehensive collection of protein sequences and BLAST and PSI-BLAST alignments between them. Instead of running BLAST or PSI-BLAST individually on each request, results are retrieved instantaneously from a database of pre-computed alignments. Filtering options allow you to find a set of sequences satisfying a set of criteria-for example, all human proteins with solved structure and without transmembrane segments. PairsDB is continually updated and covers all sequences in Uniprot. The data is stored in a MySQL relational database. Data files will be made available for download at ftp://nic.funet.fi/pub/sci/molbio. PairsDB can also be accessed interactively at http://pairsdb.csc.fi. PairsDB data is a valuable platform to build various downstream automated analysis pipelines. For example, the graph of all-against-all similarity relationships is the starting point for clustering protein families, delineating domains, improving alignment accuracy by consistency measures, and defining orthologous genes. Moreover, query-anchored stacked sequence alignments, profiles and consensus sequences are useful in studies of sequence conservation patterns for clues about possible functional sites.

MeSH Terms
Amino Acid Sequence Conserved Sequence Databases, Protein Humans Internet Sequence Alignment Sequence Analysis, Protein User-Computer Interface
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Heger Andreas
MRC Functional Genetics Unit, University of Oxford, UK.
Korpelainen Eija
Hupponen Taavi
Mattila Kimmo
Ollikainen Vesa
Holm Liisa
References (25)
25 references, click to expand
  1. Predicting coiled coils from protein sequences.
    Science. 1991 May 24;252(5009):1162-4 PMID: 2031185
  2. Domains, motifs and clusters in the protein universe.
    Curr Opin Chem Biol. 2003 Feb;7(1):5-11 PMID: 12547420
  3. Exhaustive enumeration of protein domain families.
    J Mol Biol. 2003 May 2;328(3):749-67 PMID: 12706730
  4. Bayesian search of functionally divergent protein subgroups and their function specific residues.
    Bioinformatics. 2006 Oct 15;22(20):2466-74 PMID: 16870932
  5. Accurate detection of very sparse sequence motifs.
    J Comput Biol. 2004;11(5):843-57 PMID: 15700405
  6. ProbCons: Probabilistic consistency-based multiple sequence alignment.
    Genome Res. 2005 Feb;15(2):330-40 PMID: 15687296
  7. Detecting putative orthologs.
    Bioinformatics. 2003 Sep 1;19(13):1710-1 PMID: 15593400
  8. Removing near-neighbour redundancy from large protein sequence collections.
    Bioinformatics. 1998 Jun;14(5):423-9 PMID: 9682055
  9. Sensitive pattern discovery with 'fuzzy' alignments of distantly related proteins.
    Bioinformatics. 2003;19 Suppl 1:i130-7 PMID: 12855449
  10. Predicting transmembrane protein topology with a hidden Markov model: application to complete genomes.
    J Mol Biol. 2001 Jan 19;305(3):567-80 PMID: 11152613
  11. New developments in the InterPro database.
    Nucleic Acids Res. 2007 Jan;35(Database issue):D224-8 PMID: 17202162
  12. ADDA: a domain database with global coverage of the protein universe.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D188-91 PMID: 15608174
  13. COFFEE: an objective function for multiple sequence alignments.
    Bioinformatics. 1998 Jun;14(5):407-22 PMID: 9682054
  14. The HSSP database of protein structure-sequence alignments.
    Nucleic Acids Res. 1994 Sep;22(17):3597-9 PMID: 7937066
  15. T-Coffee: A novel method for fast and accurate multiple sequence alignment.
    J Mol Biol. 2000 Sep 8;302(1):205-17 PMID: 10964570
  16. The global trace graph, a novel paradigm for searching protein sequence databases.
    Bioinformatics. 2007 Sep 15;23(18):2361-7 PMID: 17823134
  17. Towards a covering set of protein family profiles.
    Prog Biophys Mol Biol. 2000;73(5):321-37 PMID: 11063778
  18. CAST: an iterative algorithm for the complexity analysis of sequence tracts. Complexity analysis of sequence tracts.
    Bioinformatics. 2000 Oct;16(10):915-22 PMID: 11120681
  19. Clustering of highly homologous sequences to reduce the size of large protein databases.
    Bioinformatics. 2001 Mar;17(3):282-3 PMID: 11294794
  20. The CATH domain structure database: new protocols and classification levels give a more comprehensive resource for exploring evolution.
    Nucleic Acids Res. 2007 Jan;35(Database issue):D291-7 PMID: 17135200
  21. CDD: a database of conserved domain alignments with links to domain three-dimensional structure.
    Nucleic Acids Res. 2002 Jan 1;30(1):281-3 PMID: 11752315
  22. RSDB: representative protein sequence databases have high information content.
    Bioinformatics. 2000 May;16(5):458-64 PMID: 10871268
  23. Pfam: clans, web tools and services.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D247-51 PMID: 16381856
  24. SCOP database in 2004: refinements integrate structure and sequence family data.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D226-9 PMID: 14681400
  25. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2008-01-00
Epub
2007-00-05
Pages
D276-80
Language
English
Region
England
NLM ID
0411011
PMCID
PMC2238971
Subset
IM
Grants
Medical Research Council · MC_U137761446 · United Kingdom
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com