Home LiteratureArticle Details
PMID: 12368254 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, P.H.S.

The Bioperl toolkit: Perl modules for the life sciences.

Genome research ·Vol. 12 ·No. 10 ·2002-10-00 ·Pages 1611-8

Stajich JE, Block D, Boulez K, Brenner SE, Chervitz SA, Dagdigian C, Fuellen G, Gilbert JG, Korf I, Lapp H, Lehväslaiho H, Matsalla C, Mungall CJ, Osborne BI, Pocock MR, Schattner P, Senger M, Stein LD, Stupka E, Wilkinson MD, Birney E

Abstract

The Bioperl project is an international open-source collaboration of biologists, bioinformaticians, and computer scientists that has evolved over the past 7 yr into the most comprehensive library of Perl modules available for managing and manipulating life-science information. Bioperl provides an easy-to-use, stable, and consistent programming interface for bioinformatics application programmers. The Bioperl modules have been successfully and repeatedly used to reduce otherwise complex tasks to only a few lines of code. The Bioperl object model has been proven to be flexible enough to support enterprise-level applications such as EnsEMBL, while maintaining an easy learning curve for novice Perl programmers. Bioperl is capable of executing analyses and processing results from programs such as BLAST, ClustalW, or the EMBOSS suite. Interoperation with modules written in Python and Java is supported through the evolving BioCORBA bridge. Bioperl provides access to data stores such as GenBank and SwissProt via a flexible series of sequence input/output modules, and to the emerging common sequence data storage format of the Open Bioinformatics Database Access project. This study describes the overall architecture of the toolkit, the problem domains that it addresses, and gives specific examples of how the toolkit can be used to solve common life-sciences problems. We conclude with a discussion of how the open-source nature of the project has contributed to the development effort.

MeSH Terms
Algorithms Animals Biological Science Disciplines/methods,trends Computational Biology/methods,trends Computer Graphics Database Management Systems Databases, Genetic Humans Internet Online Systems Software Software Design Systems Integration
Authors & Affiliations
21 authors, click to expand affiliations / ORCID
Stajich Jason E
University Program in Genetics, Duke University, Durham, North Carolina 27710, USA. jason.stajich@duke.edu
Block David
Boulez Kris
Brenner Steven E
Chervitz Stephen A
Dagdigian Chris
Fuellen Georg
Gilbert James G R
Korf Ian
Lapp Hilmar
Lehväslaiho Heikki
Matsalla Chad
Mungall Chris J
Osborne Brian I
Pocock Matthew R
Schattner Peter
Senger Martin
Stein Lincoln D
Stupka Elia
Wilkinson Mark D
Birney Ewan
References (15)
15 references, click to expand
  1. Database resources of the National Center for Biotechnology Information: 2002 update.
    Nucleic Acids Res. 2002 Jan 1;30(1):13-6 PMID: 11752242
  2. XML, bioinformatics and data integration.
    Bioinformatics. 2001 Feb;17(2):115-25 PMID: 11238067
  3. Accessing and distributing EMBL data using CORBA (common object request broker architecture).
    Genome Biol. 2000;1(5):RESEARCH0010 PMID: 11178259
  4. RHdb: the Radiation Hybrid database.
    Nucleic Acids Res. 2001 Jan 1;29(1):165-6 PMID: 11125078
  5. T-Coffee: A novel method for fast and accurate multiple sequence alignment.
    J Mol Biol. 2000 Sep 8;302(1):205-17 PMID: 10964570
  6. EMBOSS: the European Molecular Biology Open Software Suite.
    Trends Genet. 2000 Jun;16(6):276-7 PMID: 10827456
  7. Improved tools for biological sequence comparison.
    Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8 PMID: 3162770
  8. Prediction of complete gene structures in human genomic DNA.
    J Mol Biol. 1997 Apr 25;268(1):78-94 PMID: 9149143
  9. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  10. The distributed annotation system.
    BMC Bioinformatics. 2001;2:7 PMID: 11667947
  11. Genquire: genome annotation browser/editor.
    Bioinformatics. 2002 Oct;18(10):1398-9 PMID: 12376386
  12. TFBS: Computational framework for transcription factor binding site analysis.
    Bioinformatics. 2002 Aug;18(8):1135-6 PMID: 12176838
  13. On the sequencing of the human genome.
    Proc Natl Acad Sci U S A. 2002 Mar 19;99(6):3712-6 PMID: 11880605
  14. The Ensembl genome database project.
    Nucleic Acids Res. 2002 Jan 1;30(1):38-41 PMID: 11752248
  15. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
2002-10-00
Pages
1611-8
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC187536
Subset
IM
Grants
NIGMS NIH HHS · T32 GM007754 · United States
NHGRI NIH HHS · U41 HG000739 · United States
NHGRI NIH HHS · 1 K32 HG00056 · United States
NHGRI NIH HHS · P41 HG002223 · United States
NHGRI NIH HHS · K22 HG000064 · United States
NHGRI NIH HHS · K22 HG000056 · United States
NHGRI NIH HHS · K22 HG-00064-01 · United States
NHGRI NIH HHS · P41 HG000739 · United States
NHGRI NIH HHS · P41HG02223 · United States
NHGRI NIH HHS · HG00739 · United States
NIGMS NIH HHS · T32 GM07754-22 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com