Abstract
Text-mining can assist biomedical researchers in reducing information overload by extracting useful knowledge from large collections of text. We developed a novel text-mining method based on analyzing the network structure created by symbol co-occurrences as a way to extend the capabilities of knowledge extraction. The method was applied to the task of automatic gene and protein name synonym extraction. Performance was measured on a test set consisting of about 50,000 abstracts from one year of MEDLINE. Synonyms retrieved from curated genomics databases were used as a gold standard. The system obtained a maximum F-score of 22.21% (23.18% precision and 21.36% recall), with high efficiency in the use of seed pairs. The method performs comparably with other studied methods, does not rely on sophisticated named-entity recognition, and requires little initial seed knowledge.
MeSH Terms
Algorithms
Artificial Intelligence
Automation
Computational Biology/methods
Computer Graphics
Computers
Database Management Systems
Databases, Bibliographic
Databases, Genetic
Gene Expression Regulation, Neoplastic
Genome
Humans
Information Storage and Retrieval
Information Systems
MEDLINE
Natural Language Processing
Neoplasms/genetics
Neural Networks, Computer
Pattern Recognition, Automated
Programming Languages
Reproducibility of Results
Software
Software Design
Terminology as Topic
Vocabulary, Controlled
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Cohen A M
Department of Medical Informatics and Clinical Epidemiology, School of Medicine, Oregon Health & Science University, 3181 S,W, Sam Jackson Park Road, Portland, Oregon 97239-3098, USA. cohenaa@ohsu.edu
Hersh W R
Dubay C
Spackman K
References (28)
28 references, click to expand
-
KEGG: kyoto encyclopedia of genes and genomes.
Nucleic Acids Res. 2000 Jan 1;28(1):27-30
PMID: 10592173
-
Detecting Gene Symbols and Names in Biological Texts: A First Step toward Pertinent Information Extraction.
Genome Inform Ser Workshop Genome Inform. 1998;9:72-80
PMID: 11072323
-
The structure of scientific collaboration networks.
Proc Natl Acad Sci U S A. 2001 Jan 16;98(2):404-9
PMID: 11149952
-
A literature network of human genes for high-throughput analysis of gene expression.
Nat Genet. 2001 May;28(1):21-8
PMID: 11326270
-
Automatic extraction of acronym-meaning pairs from MEDLINE databases.
Stud Health Technol Inform. 2001;84(Pt 1):371-5
PMID: 11604766
-
Genew: the human gene nomenclature database.
Nucleic Acids Res. 2002 Jan 1;30(1):169-71
PMID: 11752283
-
The HUGO Gene Nomenclature Committee (HGNC).
Hum Genet. 2001 Dec;109(6):678-80
PMID: 11810281
-
MeSHmap: a text mining tool for MEDLINE.
Proc AMIA Symp. 2001;:642-6
PMID: 11825264
-
Guidelines for human gene nomenclature.
Genomics. 2002 Apr;79(4):464-70
PMID: 11944974
-
Tagging gene and protein names in biomedical text.
Bioinformatics. 2002 Aug;18(8):1124-32
PMID: 12176836
-
Creating an online dictionary of abbreviations from MEDLINE.
J Am Med Inform Assoc. 2002 Nov-Dec;9(6):612-20
PMID: 12386112
-
Automatic extraction of gene and protein synonyms from MEDLINE and journal articles.
Proc AMIA Symp. 2002;:919-23
PMID: 12463959
-
An alternate translation initiation site circumvents an amino-terminal DAX1 nonsense mutation leading to a mild form of X-linked adrenal hypoplasia congenita.
J Clin Endocrinol Metab. 2003 Jan;88(1):417-23
PMID: 12519885
-
The FlyBase database of the Drosophila genome projects and community literature.
Nucleic Acids Res. 2003 Jan 1;31(1):172-5
PMID: 12519974
-
Playing biology's name game: identifying protein names in scientific text.
Pac Symp Biocomput. 2003;:403-14
PMID: 12603045
-
Mining terminological knowledge in large biomedical corpora.
Pac Symp Biocomput. 2003;:415-26
PMID: 12603046
-
p21 expression predicts outcome in p53-null ovarian carcinoma.
Clin Cancer Res. 2003 Mar;9(3):1028-32
PMID: 12631602
-
PTEN decreases in vivo vascularization of experimental gliomas in spite of proangiogenic stimuli.
Cancer Res. 2003 May 1;63(9):2300-5
PMID: 12727853
-
Rutabaga by any other name: extracting biological names.
J Biomed Inform. 2002 Aug;35(4):247-59
PMID: 12755519
-
Extracting synonymous gene and protein terms from biological literature.
Bioinformatics. 2003;19 Suppl 1:i340-9
PMID: 12855479
-
DAX1 and its network partners: exploring complexity in development.
Mol Genet Metab. 2003 Sep-Oct;80(1-2):81-120
PMID: 14567960
-
NOD2/CARD15 variants are associated with lower weight at diagnosis in children with Crohn's disease.
Am J Gastroenterol. 2003 Nov;98(11):2479-84
PMID: 14638352
-
Gene indexing: characterization and analysis of NLM's GeneRIFs.
AMIA Annu Symp Proc. 2003;:460-4
PMID: 14728215
-
Interference of BCR-ABL1 kinase activity with antigen receptor signaling in B cell precursor leukemia cells.
Cell Cycle. 2004 Jul;3(7):858-60
PMID: 15254401
-
Mining MEDLINE for implicit links between dietary substances and diseases.
Bioinformatics. 2004 Aug 4;20 Suppl 1:i290-6
PMID: 15262811
-
Medical literature as a potential source of new knowledge.
Bull Med Libr Assoc. 1990 Jan;78(1):29-37
PMID: 2403828
-
GeneNet: a gene network database and its automated visualization.
Bioinformatics. 1998;14(6):529-37
PMID: 9694992
-
Using ARROWSMITH: a computer-assisted approach to formulating and assessing scientific hypotheses.
Comput Methods Programs Biomed. 1998 Nov;57(3):149-53
PMID: 9822851