Home LiteratureArticle Details
PMID: 15561790 Published · ppublish English Journal Article Research Support, U.S. Gov't, P.H.S.

A statistical approach to scanning the biomedical literature for pharmacogenetics knowledge.

Journal of the American Medical Informatics Association : JAMIA ·Vol. 12 ·No. 2 ·2005-00-00 ·Pages 121-9

Rubin DL, Thorn CF, Klein TE, Altman RB

Abstract

Biomedical databases summarize current scientific knowledge, but they generally require years of laborious curation effort to build, focusing on identifying pertinent literature and data in the voluminous biomedical literature. It is difficult to manually extract useful information embedded in the large volumes of literature, and automated intelligent text analysis tools are becoming increasingly essential to assist in these curation activities. The goal of the authors was to develop an automated method to identify articles in Medline citations that contain pharmacogenetics data pertaining to gene-drug relationships. The authors built and evaluated several candidate statistical models that characterize pharmacogenetics articles in terms of word usage and the profile of Medical Subject Headings (MeSH) used in those articles. The best-performing model was used to scan the entire Medline article database (11 million articles) to identify candidate pharmacogenetics articles. A sampling of the articles identified from scanning Medline was reviewed by a pharmacologist to assess the precision of the method. The authors' approach identified 4,892 pharmacogenetics articles in the literature with 92% precision. Their automated method took a fraction of the time to acquire these articles compared with the time expected to be taken to accumulate them manually. The authors have built a Web resource (http://pharmdemo.stanford.edu/pharmdb/main.spy) to provide access to their results. A statistical classification approach can screen the primary literature to pharmacogenetics articles with high precision. Such methods may assist curators in acquiring pertinent literature in building biomedical databases.

MeSH Terms
Bayes Theorem Databases, Bibliographic/classification Information Storage and Retrieval/methods Logistic Models MEDLINE Medical Subject Headings Models, Statistical Pharmacogenetics/statistics & numerical data
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Rubin Daniel L
Section of Medical Informatics, MSOB X-215, Stanford, CA 94305, USA. rubin@smi.stanford.edu
Thorn Caroline F
Klein Teri E
Altman Russ B
References (13)
13 references, click to expand
  1. Accomplishments and challenges in literature data mining for biology.
    Bioinformatics. 2002 Dec;18(12):1553-61 PMID: 12490438
  2. Pharmacogenomics: translating functional genomics into rational therapeutics.
    Science. 1999 Oct 15;286(5439):487-91 PMID: 10521338
  3. Medical subject headings used to search the biomedical literature.
    J Am Med Inform Assoc. 2001 Jul-Aug;8(4):317-23 PMID: 11418538
  4. Pharmacogenomics: the inherited basis for interindividual differences in drug response.
    Annu Rev Genomics Hum Genet. 2001;2:9-39 PMID: 11701642
  5. PharmGKB: the Pharmacogenetics Knowledge Base.
    Nucleic Acids Res. 2002 Jan 1;30(1):163-5 PMID: 11752281
  6. Integrating genotype and phenotype information: an overview of the PharmGKB project. Pharmacogenetics Research Network and Knowledge Base.
    Pharmacogenomics J. 2001;1(3):167-70 PMID: 11908751
  7. Statistical features of human exons and their flanking regions.
    Hum Mol Genet. 1998 May;7(5):919-32 PMID: 9536098
  8. Evaluation of text data mining for database curation: lessons learned from the KDD Challenge Cup.
    Bioinformatics. 2003;19 Suppl 1:i331-9 PMID: 12855478
  9. GAPSCORE: finding gene and protein names one word at a time.
    Bioinformatics. 2004 Jan 22;20(2):216-25 PMID: 14734313
  10. Mining the biomedical literature in the genomic era: an overview.
    J Comput Biol. 2003;10(6):821-55 PMID: 14980013
  11. A resource to acquire and summarize pharmacogenetics knowledge in the literature.
    Stud Health Technol Inform. 2004;107(Pt 2):793-7 PMID: 15360921
  12. Extracting and characterizing gene-drug relationships from the literature.
    Pharmacogenetics. 2004 Sep;14(9):577-86 PMID: 15475731
  13. Selection of MEDLINE contents, the development of its thesaurus, and the indexing process.
    Med Inform (Lond). 1978 Sep;3(3):237-54 PMID: 359959
Article Info
Journal
Journal of the American Medical Informatics Association : JAMIA
Abbr.
J Am Med Inform Assoc
ISSN
1067-5027
Published
2005-00-00
Epub
2004-00-23
Pages
121-9
Language
English
Region
England
NLM ID
9430800
PMCID
PMC551544
Subset
IM
Grants
NIGMS NIH HHS · U01 GM061374 · United States
NIGMS NIH HHS · U01GM61374 · United States
Corrections
ErratumIn
-
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com