Home LiteratureArticle Details
PMID: 9682055 Published · ppublish English Comparative Study Journal Article

Removing near-neighbour redundancy from large protein sequence collections.

Bioinformatics (Oxford, England) ·Vol. 14 ·No. 5 ·1998-06-00 ·Pages 423-9

Holm L, Sander C

Abstract

To maximize the chances of biological discovery, homology searching must use an up-to-date collection of sequences. However, the available sequence databases are growing rapidly and are partially redundant in content. This leads to increasing strain on CPU resources and decreasing density of first-hand annotation. These problems are addressed by clustering closely similar sequences to yield a covering of sequence space by a representative subset of sequences. No pair of sequences in the representative set has >90% mutual sequence identity. The representative set is derived by an exhaustive search for close similarities in the sequence database in which the need for explicit sequence alignment is significantly reduced by applying deca- and pentapeptide composition filters. The algorithm was applied to the union of the Swissprot, Swissnew, Trembl, Tremblnew, Genbank, PIR, Wormpep and PDB databases. The all-against-all comparison required to generate a representative set at 90% sequence identity was accomplished in 2 days CPU time, and the removal of fragments and close similarities yielded a size reduction of 46%, from 260 000 unique sequences to 140 000 representative sequences. The practical implications are (i) faster homology searches using, for example, Fasta or Blast, and (ii) unified annotation for all sequences clustered around a representative. As tens of thousands of sequence searches are performed daily world-wide, appropriate use of the non-redundant database can lead to major savings in computer resources, without loss of efficacy. A regularly updated non-redundant protein sequence database (nrdb90), a server for homology searches against nrdb90, and a Perl script (nrdb90.pl) implementing the algorithm are available for academic use from http://www.embl-ebi.ac. uk/holm/nrdb90. holm@embl-ebi.ac.uk

MeSH Terms
Algorithms Animals Computational Biology Connectin Databases, Factual Fungal Proteins/genetics Genome, Fungal Humans Muscle Proteins/genetics Protein Kinases/genetics Proteins/classification,genetics Saccharomyces cerevisiae/genetics Sequence Alignment/methods,statistics & numerical data Software
Chemicals
Connectin Fungal Proteins Muscle Proteins Proteins TTN protein, human Protein Kinases
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Holm L
EMBL-EBI, Cambridge CB10 1SD, UK.
Sander C
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4803
Published
1998-06-00
Pages
423-9
Language
English
Region
England
NLM ID
9808944
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com