Home LiteratureArticle Details
PMID: 14704350 Published · epublish English Comparative Study Journal Article Research Support, Non-U.S. Gov't

Automatic extraction of mutations from Medline and cross-validation with OMIM.

Nucleic acids research ·Vol. 32 ·No. 1 ·2004-00-00 ·Pages 135-42

Rebholz-Schuhmann D, Marcel S, Albert S, Tolle R, Casari G, Kirsch H

Abstract

Mutations help us to understand the molecular origins of diseases. Researchers, therefore, both publish and seek disease-relevant mutations in public databases and in scientific literature, e.g. Medline. The retrieval tends to be time-consuming and incomplete. Automated screening of the literature is more efficient. We developed extraction methods (called MEMA) that scan Medline abstracts for mutations. MEMA identified 24,351 singleton mutations in conjunction with a HUGO gene name out of 16,728 abstracts. From a sample of 100 abstracts we estimated the recall for the identification of mutation-gene pairs to 35% at a precision of 93%. Recall for the mutation detection alone was >67% with a precision rate of >96%. This shows that our system produces reliable data. The subset consisting of protein sequence mutations (PSMs) from MEMA was compared to the entries in OMIM (20,503 entries versus 6699, respectively). We found 1826 PSM-gene pairs to be in common to both datasets (cross-validated). This is 27% of all PSM-gene pairs in OMIM and 91% of those pairs from OMIM which co-occur in at least one Medline abstract. We conclude that Medline covers a large portion of the mutations known to OMIM. Another large portion could be artificially produced mutations from mutagenesis experiments. Access to the database of extracted mutation-gene pairs is available through the web pages of the EBI (refer to http://www.ebi. ac.uk/rebholz/index.html).

MeSH Terms
Animals Automation Databases, Genetic Genetics, Medical/methods Humans Internet MEDLINE Mutation Point Mutation Polymorphism, Genetic Proteins/genetics Reproducibility of Results Sensitivity and Specificity Software Vocabulary
Chemicals
Proteins
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Rebholz-Schuhmann Dietrich
European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton CB10 1SD, UK. rebholz@ebi.ac.uk
Marcel Stephane
Albert Sylvie
Tolle Ralf
Casari Georg
Kirsch Harald
References (16)
16 references, click to expand
  1. Association of genes to genetically inherited diseases using data mining.
    Nat Genet. 2002 Jul;31(3):316-9 PMID: 12006977
  2. Guidelines for human gene nomenclature (1997). HUGO Nomenclature Committee.
    Genomics. 1997 Oct 15;45(2):468-71 PMID: 9344684
  3. A map of human genome sequence variation containing 1.42 million single nucleotide polymorphisms.
    Nature. 2001 Feb 15;409(6822):928-33 PMID: 11237013
  4. Detecting Gene Symbols and Names in Biological Texts: A First Step toward Pertinent Information Extraction.
    Genome Inform Ser Workshop Genome Inform. 1998;9:72-80 PMID: 11072323
  5. Automated extraction of information in molecular biology.
    FEBS Lett. 2000 Jun 30;476(1-2):12-7 PMID: 10878241
  6. Computer-assisted generation of a protein-interaction database for nuclear receptors.
    Mol Endocrinol. 2003 Aug;17(8):1555-67 PMID: 12738764
  7. EDGAR: extraction of drugs, genes and relations from the biomedical literature.
    Pac Symp Biocomput. 2000;:517-28 PMID: 10902199
  8. Disambiguating proteins, genes, and RNA in text: a machine learning approach.
    Bioinformatics. 2001;17 Suppl 1:S97-106 PMID: 11472998
  9. Information extraction in molecular biology.
    Brief Bioinform. 2002 Jun;3(2):154-65 PMID: 12139435
  10. Mining literature for protein-protein interactions.
    Bioinformatics. 2001 Apr;17(4):359-63 PMID: 11301305
  11. Automatic annotation for biological sequences by extraction of keywords from MEDLINE abstracts. Development of a prototype system.
    Proc Int Conf Intell Syst Mol Biol. 1997;5:25-32 PMID: 9322011
  12. Online Mendelian Inheritance in Man (OMIM).
    Hum Mutat. 2000;15(1):57-61 PMID: 10612823
  13. Extracting synonymous gene and protein terms from biological literature.
    Bioinformatics. 2003;19 Suppl 1:i340-9 PMID: 12855479
  14. Information technology tools for efficient SNP studies.
    Am J Pharmacogenomics. 2001;1(4):303-14 PMID: 12083962
  15. Automatic extraction of biological information from scientific text: protein-protein interactions.
    Proc Int Conf Intell Syst Mol Biol. 1999;:60-7 PMID: 10786287
  16. Molecular pathology of human haemoglobin.
    Nature. 1968 Aug 31;219(5157):902-9 PMID: 5691676
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2004-00-00
Epub
2004-00-02
Pages
135-42
Language
English
Region
England
NLM ID
0411011
PMCID
PMC373272
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com