Abstract
Pharmacogenomics studies the relationship between genetic variation and the variation in drug response phenotypes. The field is rapidly gaining importance: it promises drugs targeted to particular subpopulations based on genetic background. The pharmacogenomics literature has expanded rapidly, but is dispersed in many journals. It is challenging, therefore, to identify important associations between drugs and molecular entities--particularly genes and gene variants, and thus these critical connections are often lost. Text mining techniques can allow us to convert the free-style text to a computable, searchable format in which pharmacogenomic concepts (such as genes, drugs, polymorphisms, and diseases) are identified, and important links between these concepts are recorded. Availability of full text articles as input into text mining engines is key, as literature abstracts often do not contain sufficient information to identify these pharmacogenomic associations. Thus, building on a tool called Textpresso, we have created the Pharmspresso tool to assist in identifying important pharmacogenomic facts in full text articles. Pharmspresso parses text to find references to human genes, polymorphisms, drugs and diseases and their relationships. It presents these as a series of marked-up text fragments, in which key concepts are visually highlighted. To evaluate Pharmspresso, we used a gold standard of 45 human-curated articles. Pharmspresso identified 78%, 61%, and 74% of target gene, polymorphism, and drug concepts, respectively. Pharmspresso is a text analysis tool that extracts pharmacogenomic concepts from the literature automatically and thus captures our current understanding of gene-drug interactions in a computable form. We have made Pharmspresso available at http://pharmspresso.stanford.edu.
MeSH Terms
Computational Biology/methods
Databases, Genetic
Information Storage and Retrieval/methods
Internet
Pharmacogenetics/methods
Software
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Garten Yael
Biomedical Informatics Training Program, Stanford University, Stanford, CA, USA. ygarten@stanford.edu
Altman Russ B
References (17)
17 references, click to expand
-
Textpresso: an ontology-based information retrieval and extraction system for biological literature.
PLoS Biol. 2004 Nov;2(11):e309
PMID: 15383839
-
Extracting semantic predications from Medline citations for pharmacogenomics.
Pac Symp Biocomput. 2007;:209-20
PMID: 17990493
-
Supporting the curation of biological databases with reusable text mining.
Genome Inform. 2005;16(2):32-44
PMID: 16901087
-
Extracting and characterizing gene-drug relationships from the literature.
Pharmacogenetics. 2004 Sep;14(9):577-86
PMID: 15475731
-
Text detective: a rule-based system for gene annotation in biomedical texts.
BMC Bioinformatics. 2005;6 Suppl 1:S10
PMID: 15960822
-
Inferring pathways from gene lists using a literature-derived network of biological relationships.
Bioinformatics. 2005 Mar;21(6):788-93
PMID: 15509611
-
Predicting the insurgence of human genetic diseases associated to single point protein mutations with support vector machines and evolutionary information.
Bioinformatics. 2006 Nov 15;22(22):2729-34
PMID: 16895930
-
A statistical approach to scanning the biomedical literature for pharmacogenetics knowledge.
J Am Med Inform Assoc. 2005 Mar-Apr;12(2):121-9
PMID: 15561790
-
Ticlopidine as a selective mechanism-based inhibitor of human cytochrome P450 2C19.
Biochemistry. 2001 Oct 9;40(40):12112-22
PMID: 11580286
-
An automated procedure to identify biomedical articles that contain cancer-associated gene variants.
Hum Mutat. 2006 Sep;27(9):957-64
PMID: 16865690
-
GENIES: a natural-language processing system for the extraction of molecular pathways from journal articles.
Bioinformatics. 2001;17 Suppl 1:S74-82
PMID: 11472995
-
Automatic extraction of protein point mutations using a graph bigram association.
PLoS Comput Biol. 2007 Feb 2;3(2):e16
PMID: 17274683
-
Text mining for metabolic pathways, signaling cascades, and protein networks.
Sci STKE. 2005 May 10;2005(283):pe21
PMID: 15886388
-
MutationFinder: a high-performance system for extracting point mutation mentions from text.
Bioinformatics. 2007 Jul 15;23(14):1862-5
PMID: 17495998
-
Relemed: sentence-level search engine with relevance score for the MEDLINE database of biomedical articles.
BMC Med Inform Decis Mak. 2007 Jan 10;7:1
PMID: 17214888
-
Automated extraction of mutation data from the literature: application of MuteXt to G protein-coupled receptors and nuclear hormone receptors.
Bioinformatics. 2004 Mar 1;20(4):557-68
PMID: 14990452
-
GeneWays: a system for extracting, analyzing, visualizing, and integrating molecular pathway data.
J Biomed Inform. 2004 Feb;37(1):43-53
PMID: 15016385