Home LiteratureArticle Details
PMID: 19208194 Published · epublish English Journal Article Research Support, N.I.H., Extramural

Pharmspresso: a text mining tool for extraction of pharmacogenomic concepts and relationships from full text.

BMC bioinformatics ·Vol. 10 Suppl 2 ·2009-02-05 ·Pages S6

Garten Y, Altman RB

Abstract

Pharmacogenomics studies the relationship between genetic variation and the variation in drug response phenotypes. The field is rapidly gaining importance: it promises drugs targeted to particular subpopulations based on genetic background. The pharmacogenomics literature has expanded rapidly, but is dispersed in many journals. It is challenging, therefore, to identify important associations between drugs and molecular entities--particularly genes and gene variants, and thus these critical connections are often lost. Text mining techniques can allow us to convert the free-style text to a computable, searchable format in which pharmacogenomic concepts (such as genes, drugs, polymorphisms, and diseases) are identified, and important links between these concepts are recorded. Availability of full text articles as input into text mining engines is key, as literature abstracts often do not contain sufficient information to identify these pharmacogenomic associations. Thus, building on a tool called Textpresso, we have created the Pharmspresso tool to assist in identifying important pharmacogenomic facts in full text articles. Pharmspresso parses text to find references to human genes, polymorphisms, drugs and diseases and their relationships. It presents these as a series of marked-up text fragments, in which key concepts are visually highlighted. To evaluate Pharmspresso, we used a gold standard of 45 human-curated articles. Pharmspresso identified 78%, 61%, and 74% of target gene, polymorphism, and drug concepts, respectively. Pharmspresso is a text analysis tool that extracts pharmacogenomic concepts from the literature automatically and thus captures our current understanding of gene-drug interactions in a computable form. We have made Pharmspresso available at http://pharmspresso.stanford.edu.

MeSH Terms
Computational Biology/methods Databases, Genetic Information Storage and Retrieval/methods Internet Pharmacogenetics/methods Software
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Garten Yael
Biomedical Informatics Training Program, Stanford University, Stanford, CA, USA. ygarten@stanford.edu
Altman Russ B
References (17)
17 references, click to expand
  1. Textpresso: an ontology-based information retrieval and extraction system for biological literature.
    PLoS Biol. 2004 Nov;2(11):e309 PMID: 15383839
  2. Extracting semantic predications from Medline citations for pharmacogenomics.
    Pac Symp Biocomput. 2007;:209-20 PMID: 17990493
  3. Supporting the curation of biological databases with reusable text mining.
    Genome Inform. 2005;16(2):32-44 PMID: 16901087
  4. Extracting and characterizing gene-drug relationships from the literature.
    Pharmacogenetics. 2004 Sep;14(9):577-86 PMID: 15475731
  5. Text detective: a rule-based system for gene annotation in biomedical texts.
    BMC Bioinformatics. 2005;6 Suppl 1:S10 PMID: 15960822
  6. Inferring pathways from gene lists using a literature-derived network of biological relationships.
    Bioinformatics. 2005 Mar;21(6):788-93 PMID: 15509611
  7. Predicting the insurgence of human genetic diseases associated to single point protein mutations with support vector machines and evolutionary information.
    Bioinformatics. 2006 Nov 15;22(22):2729-34 PMID: 16895930
  8. A statistical approach to scanning the biomedical literature for pharmacogenetics knowledge.
    J Am Med Inform Assoc. 2005 Mar-Apr;12(2):121-9 PMID: 15561790
  9. Ticlopidine as a selective mechanism-based inhibitor of human cytochrome P450 2C19.
    Biochemistry. 2001 Oct 9;40(40):12112-22 PMID: 11580286
  10. An automated procedure to identify biomedical articles that contain cancer-associated gene variants.
    Hum Mutat. 2006 Sep;27(9):957-64 PMID: 16865690
  11. GENIES: a natural-language processing system for the extraction of molecular pathways from journal articles.
    Bioinformatics. 2001;17 Suppl 1:S74-82 PMID: 11472995
  12. Automatic extraction of protein point mutations using a graph bigram association.
    PLoS Comput Biol. 2007 Feb 2;3(2):e16 PMID: 17274683
  13. Text mining for metabolic pathways, signaling cascades, and protein networks.
    Sci STKE. 2005 May 10;2005(283):pe21 PMID: 15886388
  14. MutationFinder: a high-performance system for extracting point mutation mentions from text.
    Bioinformatics. 2007 Jul 15;23(14):1862-5 PMID: 17495998
  15. Relemed: sentence-level search engine with relevance score for the MEDLINE database of biomedical articles.
    BMC Med Inform Decis Mak. 2007 Jan 10;7:1 PMID: 17214888
  16. Automated extraction of mutation data from the literature: application of MuteXt to G protein-coupled receptors and nuclear hormone receptors.
    Bioinformatics. 2004 Mar 1;20(4):557-68 PMID: 14990452
  17. GeneWays: a system for extracting, analyzing, visualizing, and integrating molecular pathway data.
    J Biomed Inform. 2004 Feb;37(1):43-53 PMID: 15016385
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2009-02-05
Epub
2009-00-05
Pages
S6
Language
English
Region
England
NLM ID
100965194
PMCID
PMC2646239
Subset
IM
Grants
NIGMS NIH HHS · R24 GM061374 · United States
NIGMS NIH HHS · GM61374 · United States
NLM NIH HHS · T15 LM007033 · United States
NIGMS NIH HHS · U01 GM061374 · United States
NLM NIH HHS · LM007033 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com