Home LiteratureArticle Details
PMID: 15847682 Published · epublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, P.H.S.

Using co-occurrence network structure to extract synonymous gene and protein names from MEDLINE abstracts.

BMC bioinformatics ·Vol. 6 ·2005-04-22 ·Pages 103

Cohen AM, Hersh WR, Dubay C, Spackman K

Abstract

Text-mining can assist biomedical researchers in reducing information overload by extracting useful knowledge from large collections of text. We developed a novel text-mining method based on analyzing the network structure created by symbol co-occurrences as a way to extend the capabilities of knowledge extraction. The method was applied to the task of automatic gene and protein name synonym extraction. Performance was measured on a test set consisting of about 50,000 abstracts from one year of MEDLINE. Synonyms retrieved from curated genomics databases were used as a gold standard. The system obtained a maximum F-score of 22.21% (23.18% precision and 21.36% recall), with high efficiency in the use of seed pairs. The method performs comparably with other studied methods, does not rely on sophisticated named-entity recognition, and requires little initial seed knowledge.

MeSH Terms
Algorithms Artificial Intelligence Automation Computational Biology/methods Computer Graphics Computers Database Management Systems Databases, Bibliographic Databases, Genetic Gene Expression Regulation, Neoplastic Genome Humans Information Storage and Retrieval Information Systems MEDLINE Natural Language Processing Neoplasms/genetics Neural Networks, Computer Pattern Recognition, Automated Programming Languages Reproducibility of Results Software Software Design Terminology as Topic Vocabulary, Controlled
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Cohen A M
Department of Medical Informatics and Clinical Epidemiology, School of Medicine, Oregon Health & Science University, 3181 S,W, Sam Jackson Park Road, Portland, Oregon 97239-3098, USA. cohenaa@ohsu.edu
Hersh W R
Dubay C
Spackman K
References (28)
28 references, click to expand
  1. KEGG: kyoto encyclopedia of genes and genomes.
    Nucleic Acids Res. 2000 Jan 1;28(1):27-30 PMID: 10592173
  2. Detecting Gene Symbols and Names in Biological Texts: A First Step toward Pertinent Information Extraction.
    Genome Inform Ser Workshop Genome Inform. 1998;9:72-80 PMID: 11072323
  3. The structure of scientific collaboration networks.
    Proc Natl Acad Sci U S A. 2001 Jan 16;98(2):404-9 PMID: 11149952
  4. A literature network of human genes for high-throughput analysis of gene expression.
    Nat Genet. 2001 May;28(1):21-8 PMID: 11326270
  5. Automatic extraction of acronym-meaning pairs from MEDLINE databases.
    Stud Health Technol Inform. 2001;84(Pt 1):371-5 PMID: 11604766
  6. Genew: the human gene nomenclature database.
    Nucleic Acids Res. 2002 Jan 1;30(1):169-71 PMID: 11752283
  7. The HUGO Gene Nomenclature Committee (HGNC).
    Hum Genet. 2001 Dec;109(6):678-80 PMID: 11810281
  8. MeSHmap: a text mining tool for MEDLINE.
    Proc AMIA Symp. 2001;:642-6 PMID: 11825264
  9. Guidelines for human gene nomenclature.
    Genomics. 2002 Apr;79(4):464-70 PMID: 11944974
  10. Tagging gene and protein names in biomedical text.
    Bioinformatics. 2002 Aug;18(8):1124-32 PMID: 12176836
  11. Creating an online dictionary of abbreviations from MEDLINE.
    J Am Med Inform Assoc. 2002 Nov-Dec;9(6):612-20 PMID: 12386112
  12. Automatic extraction of gene and protein synonyms from MEDLINE and journal articles.
    Proc AMIA Symp. 2002;:919-23 PMID: 12463959
  13. An alternate translation initiation site circumvents an amino-terminal DAX1 nonsense mutation leading to a mild form of X-linked adrenal hypoplasia congenita.
    J Clin Endocrinol Metab. 2003 Jan;88(1):417-23 PMID: 12519885
  14. The FlyBase database of the Drosophila genome projects and community literature.
    Nucleic Acids Res. 2003 Jan 1;31(1):172-5 PMID: 12519974
  15. Playing biology's name game: identifying protein names in scientific text.
    Pac Symp Biocomput. 2003;:403-14 PMID: 12603045
  16. Mining terminological knowledge in large biomedical corpora.
    Pac Symp Biocomput. 2003;:415-26 PMID: 12603046
  17. p21 expression predicts outcome in p53-null ovarian carcinoma.
    Clin Cancer Res. 2003 Mar;9(3):1028-32 PMID: 12631602
  18. PTEN decreases in vivo vascularization of experimental gliomas in spite of proangiogenic stimuli.
    Cancer Res. 2003 May 1;63(9):2300-5 PMID: 12727853
  19. Rutabaga by any other name: extracting biological names.
    J Biomed Inform. 2002 Aug;35(4):247-59 PMID: 12755519
  20. Extracting synonymous gene and protein terms from biological literature.
    Bioinformatics. 2003;19 Suppl 1:i340-9 PMID: 12855479
  21. DAX1 and its network partners: exploring complexity in development.
    Mol Genet Metab. 2003 Sep-Oct;80(1-2):81-120 PMID: 14567960
  22. NOD2/CARD15 variants are associated with lower weight at diagnosis in children with Crohn's disease.
    Am J Gastroenterol. 2003 Nov;98(11):2479-84 PMID: 14638352
  23. Gene indexing: characterization and analysis of NLM's GeneRIFs.
    AMIA Annu Symp Proc. 2003;:460-4 PMID: 14728215
  24. Interference of BCR-ABL1 kinase activity with antigen receptor signaling in B cell precursor leukemia cells.
    Cell Cycle. 2004 Jul;3(7):858-60 PMID: 15254401
  25. Mining MEDLINE for implicit links between dietary substances and diseases.
    Bioinformatics. 2004 Aug 4;20 Suppl 1:i290-6 PMID: 15262811
  26. Medical literature as a potential source of new knowledge.
    Bull Med Libr Assoc. 1990 Jan;78(1):29-37 PMID: 2403828
  27. GeneNet: a gene network database and its automated visualization.
    Bioinformatics. 1998;14(6):529-37 PMID: 9694992
  28. Using ARROWSMITH: a computer-assisted approach to formulating and assessing scientific hypotheses.
    Comput Methods Programs Biomed. 1998 Nov;57(3):149-53 PMID: 9822851
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2005-04-22
Epub
2005-00-22
Pages
103
Language
English
Region
England
NLM ID
100965194
PMCID
PMC1090552
Subset
IM
Grants
NLM NIH HHS · T15 LM007088 · United States
NLM NIH HHS · 2 T15 LM07088-11 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com