Home LiteratureArticle Details
PMID: 20723615 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

Using text to build semantic networks for pharmacogenomics.

Journal of biomedical informatics ·Vol. 43 ·No. 6 ·2010-12-00 ·Pages 1009-19

Coulet A, Shah NH, Garten Y, Musen M, Altman RB

Abstract

Most pharmacogenomics knowledge is contained in the text of published studies, and is thus not available for automated computation. Natural Language Processing (NLP) techniques for extracting relationships in specific domains often rely on hand-built rules and domain-specific ontologies to achieve good performance. In a new and evolving field such as pharmacogenomics (PGx), rules and ontologies may not be available. Recent progress in syntactic NLP parsing in the context of a large corpus of pharmacogenomics text provides new opportunities for automated relationship extraction. We describe an ontology of PGx relationships built starting from a lexicon of key pharmacogenomic entities and a syntactic parse of more than 87 million sentences from 17 million MEDLINE abstracts. We used the syntactic structure of PGx statements to systematically extract commonly occurring relationships and to map them to a common schema. Our extracted relationships have a 70-87.7% precision and involve not only key PGx entities such as genes, drugs, and phenotypes (e.g., VKORC1, warfarin, clotting disorder), but also critical entities that are frequently modified by these key entities (e.g., VKORC1 polymorphism, warfarin response, clotting disorder treatment). The result of our analysis is a network of 40,000 relationships between more than 200 entity types with clear semantics. This network is used to guide the curation of PGx knowledge and provide a computable resource for knowledge discovery.

MeSH Terms
Databases, Factual MEDLINE Natural Language Processing Pharmacogenetics/methods Semantics Terminology as Topic United States
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Coulet Adrien
Department of Medicine, 300 Pasteur Drive, Room S101, Mail Code 5110, Stanford University, Stanford, CA 94305, USA.
Shah Nigam H
Garten Yael
Musen Mark
Altman Russ B
References (15)
15 references, click to expand
  1. Pharmspresso: a text mining tool for extraction of pharmacogenomic concepts and relationships from full text.
    BMC Bioinformatics. 2009 Feb 05;10 Suppl 2:S6 PMID: 19208194
  2. Extracting semantic predications from Medline citations for pharmacogenomics.
    Pac Symp Biocomput. 2007;:209-20 PMID: 17990493
  3. Automatic extraction of biological information from scientific text: protein-protein interactions.
    Proc Int Conf Intell Syst Mol Biol. 1999;:60-7 PMID: 10786287
  4. GENIES: a natural-language processing system for the extraction of molecular pathways from journal articles.
    Bioinformatics. 2001;17 Suppl 1:S74-82 PMID: 11472995
  5. RelEx--relation extraction using dependency parse trees.
    Bioinformatics. 2007 Feb 1;23(3):365-71 PMID: 17142812
  6. Building disease-specific drug-protein connectivity maps from molecular interaction networks and PubMed abstracts.
    PLoS Comput Biol. 2009 Jul;5(7):e1000450 PMID: 19649302
  7. Empirical distributional semantics: methods and biomedical applications.
    J Biomed Inform. 2009 Apr;42(2):390-405 PMID: 19232399
  8. Towards pharmacogenomics knowledge discovery with the semantic web.
    Brief Bioinform. 2009 Mar;10(2):153-63 PMID: 19240125
  9. OpenDMAP: an open source, ontology-driven concept analysis engine, with applications to capturing knowledge regarding protein transport, protein interactions and cell-type-specific gene expression.
    BMC Bioinformatics. 2008 Jan 31;9:78 PMID: 18237434
  10. Unsupervised method for automatic construction of a disease dictionary from a large free text collection.
    AMIA Annu Symp Proc. 2008 Nov 06;:820-4 PMID: 18999169
  11. Querying parse tree database of Medline text to synthesize user-specific biomolecular networks.
    Pac Symp Biocomput. 2009;:87-98 PMID: 19209697
  12. Integrating genotype and phenotype information: an overview of the PharmGKB project. Pharmacogenetics Research Network and Knowledge Base.
    Pharmacogenomics J. 2001;1(3):167-70 PMID: 11908751
  13. Extraction of regulatory gene/protein networks from Medline.
    Bioinformatics. 2006 Mar 15;22(6):645-50 PMID: 16046493
  14. Evaluation of text-mining systems for biology: overview of the Second BioCreative community challenge.
    Genome Biol. 2008;9 Suppl 2:S1 PMID: 18834487
  15. Semantic relations asserting the etiology of genetic diseases.
    AMIA Annu Symp Proc. 2003;:554-8 PMID: 14728234
Article Info
Journal
Journal of biomedical informatics
Abbr.
J Biomed Inform
ISSN
1532-0480
Published
2010-12-00
Epub
2010-00-17
Pages
1009-19
Language
English
Region
United States
NLM ID
100970413
PMCID
PMC2991587
Subset
IM
Grants
NIGMS NIH HHS · R24 GM061374 · United States
NLM NIH HHS · R01 LM005652 · United States
NHGRI NIH HHS · U54 HG004028 · United States
NIGMS NIH HHS · R24 GM061374-12 · United States
NIGMS NIH HHS · U01 GM061374 · United States
NLM NIH HHS · LM-05652 · United States
NIGMS NIH HHS · GM61374 · United States
NLM NIH HHS · T15 LM007033 · United States
NHGRI NIH HHS · U54HG004028 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com