Abstract
SeqHound has been developed as an integrated biological sequence, taxonomy, annotation and 3-D structure database system. It provides a high-performance server platform for bioinformatics research in a locally-hosted environment. SeqHound is based on the National Center for Biotechnology Information data model and programming tools. It offers daily updated contents of all Entrez sequence databases in addition to 3-D structural data and information about sequence redundancies, sequence neighbours, taxonomy, complete genomes, functional annotation including Gene Ontology terms and literature links to PubMed. SeqHound is accessible via a web server through a Perl, C or C++ remote API or an optimized local API. It provides functionality necessary to retrieve specialized subsets of sequences, structures and structural domains. Sequences may be retrieved in FASTA, GenBank, ASN.1 and XML formats. Structures are available in ASN.1, XML and PDB formats. Emphasis has been placed on complete genomes, taxonomy, domain and functional annotation as well as 3-D structural functionality in the API, while fielded text indexing functionality remains under development. SeqHound also offers a streamlined WWW interface for simple web-user queries. The system has proven useful in several published bioinformatics projects such as the BIND database and offers a cost-effective infrastructure for research. SeqHound will continue to develop and be provided as a service of the Blueprint Initiative at the Samuel Lunenfeld Research Institute. The source code and examples are available under the terms of the GNU public license at the Sourceforge site http://sourceforge.net/projects/slritools/ in the SLRI Toolkit.
MeSH Terms
Amino Acid Sequence
Base Sequence
Computational Biology/methods
Databases, Genetic/classification
Information Storage and Retrieval/methods
Internet
Models, Genetic
Models, Molecular
Molecular Sequence Data
Software
Structure-Activity Relationship
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Michalickova Katerina
Department of Biochemistry, University of Toronto, Toronto, Ontario, Canada M5S 1A8. katerina@mshri.on.ca
Bader Gary D
Dumontier Michel
Lieu Hao
Betel Doron
Isserlin Ruth
Hogue Christopher W V
References (22)
22 references, click to expand
-
The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
Nucleic Acids Res. 2000 Jan 1;28(1):45-8
PMID: 10592178
-
The Protein Data Bank.
Nucleic Acids Res. 2000 Jan 1;28(1):235-42
PMID: 10592235
-
BIND--a data specification for storing and describing biomolecular interactions, molecular complexes and pathways.
Bioinformatics. 2000 May;16(5):465-77
PMID: 10871269
-
RefSeq and LocusLink: NCBI gene-centered resources.
Nucleic Acids Res. 2001 Jan 1;29(1):137-40
PMID: 11125071
-
BIND--The Biomolecular Interaction Network Database.
Nucleic Acids Res. 2001 Jan 1;29(1):242-5
PMID: 11125103
-
Creating the gene ontology resource: design and implementation.
Genome Res. 2001 Aug;11(8):1425-33
PMID: 11483584
-
GenBank.
Nucleic Acids Res. 2002 Jan 1;30(1):17-20
PMID: 11752243
-
The EMBL Nucleotide Sequence Database.
Nucleic Acids Res. 2002 Jan 1;30(1):21-6
PMID: 11752244
-
The Protein Information Resource: an integrated public resource of functional annotation of proteins.
Nucleic Acids Res. 2002 Jan 1;30(1):35-7
PMID: 11752247
-
Recent improvements to the SMART domain-based sequence annotation resource.
Nucleic Acids Res. 2002 Jan 1;30(1):242-4
PMID: 11752305
-
MMDB: Entrez's 3D-structure database.
Nucleic Acids Res. 2002 Jan 1;30(1):249-52
PMID: 11752307
-
The Pfam protein families database.
Nucleic Acids Res. 2002 Jan 1;30(1):276-80
PMID: 11752314
-
CDD: a database of conserved domain alignments with links to domain three-dimensional structure.
Nucleic Acids Res. 2002 Jan 1;30(1):281-3
PMID: 11752315
-
NBLAST: a cluster variant of BLAST for NxN comparisons.
BMC Bioinformatics. 2002 May 8;3:13
PMID: 12019022
-
Kangaroo--a pattern-matching program for biological sequences.
BMC Bioinformatics. 2002 Jul 31;3:20
PMID: 12150718
-
CLUSTAL: a package for performing multiple sequence alignment on a microcomputer.
Gene. 1988 Dec 15;73(1):237-44
PMID: 3243435
-
Basic local alignment search tool.
J Mol Biol. 1990 Oct 5;215(3):403-10
PMID: 2231712
-
dbEST--database for "expressed sequence tags".
Nat Genet. 1993 Aug;4(4):332-3
PMID: 8401577
-
CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
Nucleic Acids Res. 1994 Nov 11;22(22):4673-80
PMID: 7984417
-
Entrez: molecular biology database and retrieval system.
Methods Enzymol. 1996;266:141-62
PMID: 8743683
-
The NCBI data model.
Methods Biochem Anal. 1998;39:121-44
PMID: 9707929
-
Kleisli: a new tool for data integration in biology.
Trends Biotechnol. 1999 Sep;17(9):351-5
PMID: 10461180