Abstract
Expressed sequence tags (ESTs) are generated and deposited in the public domain, as redundant, unannotated, single-pass reactions, with virtually no biological content. PipeOnline automatically analyses and transforms large collections of raw DNA-sequence data from chromatograms or FASTA files by calling the quality of bases, screening and removing vector sequences, assembling and rewriting consensus sequences of redundant input files into a unigene EST data set and finally through translation, amino acid sequence similarity searches, annotation of public databases and functional data. PipeOnline generates an annotated database, retaining the processed unigene sequence, clone/file history, alignments with similar sequences, and proposed functional classification, if available. Functional annotation is automatic and based on a novel method that relies on homology of amino acid sequence multiplicity within GenBank records. Records are examined through a function ordered browser or keyword queries with automated export of results. PipeOnline offers customization for individual projects (MyPipeOnline), automated updating and alert service. PipeOnline is available at http://stress-genomics.org.
MeSH Terms
Automation/methods
Computational Biology/methods
Consensus Sequence/genetics
Conserved Sequence/genetics
Databases, Genetic
Databases, Protein
Expressed Sequence Tags
Genomics/methods
Internet
Open Reading Frames
Sequence Alignment
Sequence Homology, Amino Acid
Software
Authors & Affiliations
10 authors, click to expand affiliations / ORCID
Ayoubi Patricia
Department of Microbiology and Molecular Genetics and. School of Mechanical and Aerospace Engineering, Oklahoma State University, Stillwater, OK 74078, USA.
Jin Xiaojing
Leite Saul
Liu Xianghui
Martajaja Jeson
Abduraham Abdurashid
Wan Qiaolan
Yan Wei
Misawa Eduardo
Prade Rolf A
References (26)
26 references, click to expand
-
Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
Nat Genet. 2000 May;25(1):25-9
PMID: 10802651
-
Automated genome sequence analysis and annotation.
Bioinformatics. 1999 May;15(5):391-412
PMID: 10366660
-
EST databases as multi-conditional gene expression datasets.
Pac Symp Biocomput. 2000;:430-42
PMID: 10902191
-
Recent developments and future directions in computational genomics.
FEBS Lett. 2000 Aug 25;480(1):42-8
PMID: 10967327
-
Mendel-GFDb and Mendel-ESTS: databases of plant gene families and ESTs annotated with gene family numbers and gene family names.
Nucleic Acids Res. 2001 Jan 1;29(1):120-2
PMID: 11125066
-
The TIGR Gene Indices: analysis of gene transcript sequences in highly sampled eukaryotic species.
Nucleic Acids Res. 2001 Jan 1;29(1):159-64
PMID: 11125077
-
STACK: Sequence Tag Alignment and Consensus Knowledgebase.
Nucleic Acids Res. 2001 Jan 1;29(1):234-8
PMID: 11125101
-
Assembly, annotation, and integration of UNIGENE clusters into the human genome draft.
Genome Res. 2001 May;11(5):904-18
PMID: 11337484
-
Genome analysis with gene-indexing databases.
Pharmacol Ther. 2001 Aug;91(2):115-32
PMID: 11728605
-
Cross-referencing eukaryotic genomes: TIGR Orthologous Gene Alignments (TOGA).
Genome Res. 2002 Mar;12(3):493-502
PMID: 11875039
-
Identification of common molecular subsequences.
J Mol Biol. 1981 Mar 25;147(1):195-7
PMID: 7265238
-
The COG database: a tool for genome-scale analysis of protein functions and evolution.
Nucleic Acids Res. 2000 Jan 1;28(1):33-6
PMID: 10592175
-
MIPS: a database for genomes and protein sequences.
Nucleic Acids Res. 2000 Jan 1;28(1):37-40
PMID: 10592176
-
The EcoCyc and MetaCyc databases.
Nucleic Acids Res. 2000 Jan 1;28(1):56-9
PMID: 10592180
-
WIT: integrated system for high-throughput genome sequence analysis and metabolic reconstruction.
Nucleic Acids Res. 2000 Jan 1;28(1):123-5
PMID: 10592199
-
The TIGR gene indices: reconstruction and representation of expressed gene sequences.
Nucleic Acids Res. 2000 Jan 1;28(1):141-5
PMID: 10592205
-
An improved algorithm for matching biological sequences.
J Mol Biol. 1982 Dec 15;162(3):705-8
PMID: 7166760
-
Basic local alignment search tool.
J Mol Biol. 1990 Oct 5;215(3):403-10
PMID: 2231712
-
Effective large-scale sequence similarity searches.
Methods Enzymol. 1996;266:212-27
PMID: 8743687
-
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
PMID: 9254694
-
The significance of digital gene expression profiles.
Genome Res. 1997 Oct;7(10):986-95
PMID: 9331369
-
MPW: the Metabolic Pathways Database.
Nucleic Acids Res. 1998 Jan 1;26(1):43-5
PMID: 9407141
-
Base-calling of automated sequencer traces using phred. I. Accuracy assessment.
Genome Res. 1998 Mar;8(3):175-85
PMID: 9521921
-
Automated sequence preprocessing in a large-scale sequencing environment.
Genome Res. 1998 Sep;8(9):975-84
PMID: 9750196
-
The ENZYME data bank in 1999.
Nucleic Acids Res. 1999 Jan 1;27(1):310-1
PMID: 9847212
-
An ontology for biological function based on molecular interactions.
Bioinformatics. 2000 Mar;16(3):269-85
PMID: 10869020