Home LiteratureArticle Details
PMID: 15096276 Published · epublish English Comparative Study Journal Article Research Support, Non-U.S. Gov't

Pegasys: software for executing and integrating analyses of biological sequences.

BMC bioinformatics ·Vol. 5 ·2004-04-19 ·Pages 40

Shah SP, He DY, Sawkins JN, Druce JC, Quon G, Lett D, Zheng GX, Xu T, Ouellette BF

Abstract

We present Pegasys--a flexible, modular and customizable software system that facilitates the execution and data integration from heterogeneous biological sequence analysis tools. The Pegasys system includes numerous tools for pair-wise and multiple sequence alignment, ab initio gene prediction, RNA gene detection, masking repetitive sequences in genomic DNA as well as filters for database formatting and processing raw output from various analysis tools. We introduce a novel data structure for creating workflows of sequence analyses and a unified data model to store its results. The software allows users to dynamically create analysis workflows at run-time by manipulating a graphical user interface. All non-serial dependent analyses are executed in parallel on a compute cluster for efficiency of data generation. The uniform data model and backend relational database management system of Pegasys allow for results of heterogeneous programs included in the workflow to be integrated and exported into General Feature Format for further analyses in GFF-dependent tools, or GAME XML for import into the Apollo genome editor. The modularity of the design allows for new tools to be added to the system with little programmer overhead. The database application programming interface allows programmatic access to the data stored in the backend through SQL queries. The Pegasys system enables biologists and bioinformaticians to create and manage sequence analysis workflows. The software is released under the Open Source GNU General Public License. All source code and documentation is available for download at http://bioinformatics.ubc.ca/pegasys/.

MeSH Terms
Computational Biology/methods,trends Computer Graphics DNA/genetics Databases, Genetic Game Theory Genetic Heterogeneity Humans Programming Languages Sequence Alignment/methods,trends Sequence Analysis, DNA/methods Software/trends Software Design User-Computer Interface
Chemicals
DNA
Authors & Affiliations
9 authors, click to expand affiliations / ORCID
Shah Sohrab P
UBC Bioinformatics Centre, University of British Columbia, Vancouver, British Columbia, Canada. sohrab@bioinformatics.ubc.ca
He David Y M
Sawkins Jessica N
Druce Jeffrey C
Quon Gerald
Lett Drew
Zheng Grace X Y
Xu Tao
Ouellette B F Francis
References (27)
27 references, click to expand
  1. The generic genome browser: a building block for a model organism system database.
    Genome Res. 2002 Oct;12(10):1599-610 PMID: 12368253
  2. Improving gene recognition accuracy by combining predictions from two gene-finding programs.
    Bioinformatics. 2002 Aug;18(8):1034-45 PMID: 12176826
  3. Rfam: an RNA family database.
    Nucleic Acids Res. 2003 Jan 1;31(1):439-41 PMID: 12520045
  4. An integrated computational pipeline and database to support whole-genome sequence annotation.
    Genome Biol. 2002;3(12):RESEARCH0081 PMID: 12537570
  5. Apollo: a sequence annotation editor.
    Genome Biol. 2002;3(12):RESEARCH0082 PMID: 12537571
  6. LAGAN and Multi-LAGAN: efficient tools for large-scale multiple alignment of genomic DNA.
    Genome Res. 2003 Apr;13(4):721-31 PMID: 12654723
  7. GenDB--an open source genome annotation system for prokaryote genomes.
    Nucleic Acids Res. 2003 Apr 15;31(8):2187-95 PMID: 12682369
  8. The discovery net system for high throughput bioinformatics.
    Bioinformatics. 2003;19 Suppl 1:i225-31 PMID: 12855463
  9. Biopipe: a flexible framework for protocol-based bioinformatics analysis.
    Genome Res. 2003 Aug;13(8):1904-15 PMID: 12869579
  10. The distributed annotation system.
    BMC Bioinformatics. 2001;2:7 PMID: 11667947
  11. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  12. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  13. tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence.
    Nucleic Acids Res. 1997 Mar 1;25(5):955-64 PMID: 9023104
  14. Prediction of complete gene structures in human genomic DNA.
    J Mol Biol. 1997 Apr 25;268(1):78-94 PMID: 9149143
  15. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  16. Two methods for improving performance of an HMM and their application for gene finding.
    Proc Int Conf Intell Syst Mol Biol. 1997;5:179-86 PMID: 9322033
  17. Microbial gene identification using interpolated Markov models.
    Nucleic Acids Res. 1998 Jan 15;26(2):544-8 PMID: 9421513
  18. A computer program for aligning a cDNA sequence with a genomic DNA sequence.
    Genome Res. 1998 Sep;8(9):967-74 PMID: 9750195
  19. Improved microbial gene identification with GLIMMER.
    Nucleic Acids Res. 1999 Dec 1;27(23):4636-41 PMID: 10556321
  20. Gene prediction and gene classes in Arabidopsis thaliana.
    J Biotechnol. 2000 Mar 31;78(3):293-9 PMID: 10751690
  21. EMBOSS: the European Molecular Biology Open Software Suite.
    Trends Genet. 2000 Jun;16(6):276-7 PMID: 10827456
  22. MaskerAid: a performance enhancement to RepeatMasker.
    Bioinformatics. 2000 Nov;16(11):1040-1 PMID: 11159316
  23. GeneSplicer: a new computational method for splice site prediction.
    Nucleic Acids Res. 2001 Mar 1;29(5):1185-90 PMID: 11222768
  24. Computational inference of homologous gene structures in the human genome.
    Genome Res. 2001 May;11(5):803-16 PMID: 11337476
  25. Integrating genomic homology into gene structure prediction.
    Bioinformatics. 2001;17 Suppl 1:S140-8 PMID: 11473003
  26. The Ensembl genome database project.
    Nucleic Acids Res. 2002 Jan 1;30(1):38-41 PMID: 11752248
  27. The TIGR rice genome annotation resource: annotating the rice genome and creating resources for plant biologists.
    Nucleic Acids Res. 2003 Jan 1;31(1):229-33 PMID: 12519988
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2004-04-19
Epub
2004-00-19
Pages
40
Language
English
Region
England
NLM ID
100965194
PMCID
PMC406494
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com