Home LiteratureArticle Details
PMID: 24564403 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Literature classification for semi-automated updating of biological knowledgebases.

BMC genomics ·Vol. 14 Suppl 5 ·2013-00-00 ·Pages S14

Olsen L, Johan Kudahl U, Winther O, Brusic V

Abstract

As the output of biological assays increase in resolution and volume, the body of specialized biological data, such as functional annotations of gene and protein sequences, enables extraction of higher-level knowledge needed for practical application in bioinformatics. Whereas common types of biological data, such as sequence data, are extensively stored in biological databases, functional annotations, such as immunological epitopes, are found primarily in semi-structured formats or free text embedded in primary scientific literature. We defined and applied a machine learning approach for literature classification to support updating of TANTIGEN, a knowledgebase of tumor T-cell antigens. Abstracts from PubMed were downloaded and classified as either "relevant" or "irrelevant" for database update. Training and five-fold cross-validation of a k-NN classifier on 310 abstracts yielded classification accuracy of 0.95, thus showing significant value in support of data extraction from the literature. We here propose a conceptual framework for semi-automated extraction of epitope data embedded in scientific literature using principles from text mining and machine learning. The addition of such data will aid in the transition of biological databases to knowledgebases.

MeSH Terms
Artificial Intelligence Data Mining/methods Epitopes, T-Lymphocyte Humans Information Storage and Retrieval/methods Knowledge Bases PubMed
Chemicals
Epitopes, T-Lymphocyte
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Olsen Lars
Johan Kudahl Ulrich
Winther Ole
Brusic Vladimir
References (23)
23 references, click to expand
  1. The Catalogue of Somatic Mutations in Cancer (COSMIC).
    Curr Protoc Hum Genet. 2008 Apr;Chapter 10:Unit 10.11 PMID: 18428421
  2. T cell defined tumor antigens.
    Curr Opin Immunol. 1997 Oct;9(5):684-93 PMID: 9368778
  3. FLAVIdB: A data mining system for knowledge discovery in flaviviruses with direct applications in immunology and vaccinology.
    Immunome Res. 2011;7(3): PMID: 25544857
  4. Textmining in support of knowledge discovery for vaccine development.
    Methods. 2004 Dec;34(4):488-95 PMID: 15542375
  5. GenBank.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D48-53 PMID: 22144687
  6. NetMHC-3.0: accurate web accessible predictions of human, mouse and monkey MHC class I affinities for peptides of length 8-11.
    Nucleic Acids Res. 2008 Jul 1;36(Web Server issue):W509-12 PMID: 18463140
  7. A listing of human tumor antigens recognized by T cells: March 2004 update.
    Cancer Immunol Immunother. 2005 Mar;54(3):187-207 PMID: 15309328
  8. The 2013 Nucleic Acids Research Database Issue and the online molecular biology database collection.
    Nucleic Acids Res. 2013 Jan;41(Database issue):D1-7 PMID: 23203983
  9. Supporting the curation of biological databases with reusable text mining.
    Genome Inform. 2005;16(2):32-44 PMID: 16901087
  10. The immune epitope database 2.0.
    Nucleic Acids Res. 2010 Jan;38(Database issue):D854-62 PMID: 19906713
  11. A listing of human tumor antigens recognized by T cells.
    Cancer Immunol Immunother. 2001 Mar;50(1):3-15 PMID: 11315507
  12. Parallel detection of antigen-specific T cell responses by combinatorial encoding of MHC multimers.
    Nat Protoc. 2012 Apr 12;7(5):891-902 PMID: 22498709
  13. Information technologies for vaccine research.
    Expert Rev Vaccines. 2005 Jun;4(3):407-17 PMID: 16026252
  14. UniProt Knowledgebase: a hub of integrated protein data.
    Database (Oxford). 2011 Mar 29;2011:bar009 PMID: 21447597
  15. Recent developments in the MAFFT multiple sequence alignment program.
    Brief Bioinform. 2008 Jul;9(4):286-98 PMID: 18372315
  16. PubMed and beyond: a survey of web tools for searching biomedical literature.
    Database (Oxford). 2011 Jan 18;2011:baq036 PMID: 21245076
  17. SYFPEITHI: database for searching and T-cell epitope prediction.
    Methods Mol Biol. 2007;409:75-93 PMID: 18449993
  18. PubFinder: a tool for improving retrieval rate of relevant PubMed abstracts.
    Nucleic Acids Res. 2005 Jul 1;33(Web Server issue):W774-8 PMID: 15980583
  19. Identification of human MHC class I binding peptides using the iTOPIA- epitope discovery system.
    Methods Mol Biol. 2009;524:361-7 PMID: 19377958
  20. Influenza research database: an integrated bioinformatics resource for influenza research and surveillance.
    Influenza Other Respir Viruses. 2012 Nov;6(6):404-16 PMID: 22260278
  21. Prediction of MHC class II binding affinity using SMM-align, a novel stabilization matrix alignment method.
    BMC Bioinformatics. 2007 Jul 04;8:238 PMID: 17608956
  22. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  23. Linked data and provenance in biological data webs.
    Brief Bioinform. 2009 Mar;10(2):139-52 PMID: 19060306
Article Info
Journal
BMC genomics
Abbr.
BMC Genomics
ISSN
1471-2164
Published
2013-00-00
Epub
2013-00-16
Pages
S14
Language
English
Region
England
NLM ID
100965258
PMCID
PMC3852072
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com