Home LiteratureArticle Details
PMID: 12401134 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, P.H.S.

SeqHound: biological sequence and structure database as a platform for bioinformatics research.

BMC bioinformatics ·Vol. 3 ·2002-10-25 ·Pages 32

Michalickova K, Bader GD, Dumontier M, Lieu H, Betel D, Isserlin R, Hogue CW

Abstract

SeqHound has been developed as an integrated biological sequence, taxonomy, annotation and 3-D structure database system. It provides a high-performance server platform for bioinformatics research in a locally-hosted environment. SeqHound is based on the National Center for Biotechnology Information data model and programming tools. It offers daily updated contents of all Entrez sequence databases in addition to 3-D structural data and information about sequence redundancies, sequence neighbours, taxonomy, complete genomes, functional annotation including Gene Ontology terms and literature links to PubMed. SeqHound is accessible via a web server through a Perl, C or C++ remote API or an optimized local API. It provides functionality necessary to retrieve specialized subsets of sequences, structures and structural domains. Sequences may be retrieved in FASTA, GenBank, ASN.1 and XML formats. Structures are available in ASN.1, XML and PDB formats. Emphasis has been placed on complete genomes, taxonomy, domain and functional annotation as well as 3-D structural functionality in the API, while fielded text indexing functionality remains under development. SeqHound also offers a streamlined WWW interface for simple web-user queries. The system has proven useful in several published bioinformatics projects such as the BIND database and offers a cost-effective infrastructure for research. SeqHound will continue to develop and be provided as a service of the Blueprint Initiative at the Samuel Lunenfeld Research Institute. The source code and examples are available under the terms of the GNU public license at the Sourceforge site http://sourceforge.net/projects/slritools/ in the SLRI Toolkit.

MeSH Terms
Amino Acid Sequence Base Sequence Computational Biology/methods Databases, Genetic/classification Information Storage and Retrieval/methods Internet Models, Genetic Models, Molecular Molecular Sequence Data Software Structure-Activity Relationship
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Michalickova Katerina
Department of Biochemistry, University of Toronto, Toronto, Ontario, Canada M5S 1A8. katerina@mshri.on.ca
Bader Gary D
Dumontier Michel
Lieu Hao
Betel Doron
Isserlin Ruth
Hogue Christopher W V
References (22)
22 references, click to expand
  1. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):45-8 PMID: 10592178
  2. The Protein Data Bank.
    Nucleic Acids Res. 2000 Jan 1;28(1):235-42 PMID: 10592235
  3. BIND--a data specification for storing and describing biomolecular interactions, molecular complexes and pathways.
    Bioinformatics. 2000 May;16(5):465-77 PMID: 10871269
  4. RefSeq and LocusLink: NCBI gene-centered resources.
    Nucleic Acids Res. 2001 Jan 1;29(1):137-40 PMID: 11125071
  5. BIND--The Biomolecular Interaction Network Database.
    Nucleic Acids Res. 2001 Jan 1;29(1):242-5 PMID: 11125103
  6. Creating the gene ontology resource: design and implementation.
    Genome Res. 2001 Aug;11(8):1425-33 PMID: 11483584
  7. GenBank.
    Nucleic Acids Res. 2002 Jan 1;30(1):17-20 PMID: 11752243
  8. The EMBL Nucleotide Sequence Database.
    Nucleic Acids Res. 2002 Jan 1;30(1):21-6 PMID: 11752244
  9. The Protein Information Resource: an integrated public resource of functional annotation of proteins.
    Nucleic Acids Res. 2002 Jan 1;30(1):35-7 PMID: 11752247
  10. Recent improvements to the SMART domain-based sequence annotation resource.
    Nucleic Acids Res. 2002 Jan 1;30(1):242-4 PMID: 11752305
  11. MMDB: Entrez's 3D-structure database.
    Nucleic Acids Res. 2002 Jan 1;30(1):249-52 PMID: 11752307
  12. The Pfam protein families database.
    Nucleic Acids Res. 2002 Jan 1;30(1):276-80 PMID: 11752314
  13. CDD: a database of conserved domain alignments with links to domain three-dimensional structure.
    Nucleic Acids Res. 2002 Jan 1;30(1):281-3 PMID: 11752315
  14. NBLAST: a cluster variant of BLAST for NxN comparisons.
    BMC Bioinformatics. 2002 May 8;3:13 PMID: 12019022
  15. Kangaroo--a pattern-matching program for biological sequences.
    BMC Bioinformatics. 2002 Jul 31;3:20 PMID: 12150718
  16. CLUSTAL: a package for performing multiple sequence alignment on a microcomputer.
    Gene. 1988 Dec 15;73(1):237-44 PMID: 3243435
  17. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  18. dbEST--database for "expressed sequence tags".
    Nat Genet. 1993 Aug;4(4):332-3 PMID: 8401577
  19. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  20. Entrez: molecular biology database and retrieval system.
    Methods Enzymol. 1996;266:141-62 PMID: 8743683
  21. The NCBI data model.
    Methods Biochem Anal. 1998;39:121-44 PMID: 9707929
  22. Kleisli: a new tool for data integration in biology.
    Trends Biotechnol. 1999 Sep;17(9):351-5 PMID: 10461180
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2002-10-25
Epub
2002-00-25
Pages
32
Language
English
Region
England
NLM ID
100965194
PMCID
PMC138791
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com