Exploring supervised and unsupervised methods to detect topics in biomedical text
- PMID: 16539745
- PMCID: PMC1472693
- DOI: 10.1186/1471-2105-7-140
Exploring supervised and unsupervised methods to detect topics in biomedical text
Abstract
Background: Topic detection is a task that automatically identifies topics (e.g., "biochemistry" and "protein structure") in scientific articles based on information content. Topic detection will benefit many other natural language processing tasks including information retrieval, text summarization and question answering; and is a necessary step towards the building of an information system that provides an efficient way for biologists to seek information from an ocean of literature.
Results: We have explored the methods of Topic Spotting, a task of text categorization that applies the supervised machine-learning technique naïve Bayes to assign automatically a document into one or more predefined topics; and Topic Clustering, which apply unsupervised hierarchical clustering algorithms to aggregate documents into clusters such that each cluster represents a topic. We have applied our methods to detect topics of more than fifteen thousand of articles that represent over sixteen thousand entries in the Online Mendelian Inheritance in Man (OMIM) database. We have explored bag of words as the features. Additionally, we have explored semantic features; namely, the Medical Subject Headings (MeSH) that are assigned to the MEDLINE records, and the Unified Medical Language System (UMLS) semantic types that correspond to the MeSH terms, in addition to bag of words, to facilitate the tasks of topic detection. Our results indicate that incorporating the MeSH terms and the UMLS semantic types as additional features enhances the performance of topic detection and the naïve Bayes has the highest accuracy, 66.4%, for predicting the topic of an OMIM article as one of the total twenty-five topics.
Conclusion: Our results indicate that the supervised topic spotting methods outperformed the unsupervised topic clustering; on the other hand, the unsupervised topic clustering methods have the advantages of being robust and applicable in real world settings.
Figures




Similar articles
-
Using MEDLINE as a knowledge source for disambiguating abbreviations and acronyms in full-text biomedical journal articles.J Biomed Inform. 2007 Apr;40(2):150-9. doi: 10.1016/j.jbi.2006.06.001. Epub 2006 Jun 7. J Biomed Inform. 2007. PMID: 16843731
-
Automatic assignment of biomedical categories: toward a generic approach.Bioinformatics. 2006 Mar 15;22(6):658-64. doi: 10.1093/bioinformatics/bti783. Epub 2005 Nov 15. Bioinformatics. 2006. PMID: 16287934
-
Enhancing MEDLINE document clustering by incorporating MeSH semantic similarity.Bioinformatics. 2009 Aug 1;25(15):1944-51. doi: 10.1093/bioinformatics/btp338. Epub 2009 Jun 3. Bioinformatics. 2009. PMID: 19497938
-
Status of text-mining techniques applied to biomedical text.Drug Discov Today. 2006 Apr;11(7-8):315-25. doi: 10.1016/j.drudis.2006.02.011. Drug Discov Today. 2006. PMID: 16580973 Review.
-
Recent advances in natural language processing for biomedical applications.Int J Med Inform. 2006 Jun;75(6):413-7. doi: 10.1016/j.ijmedinf.2005.06.008. Epub 2005 Aug 31. Int J Med Inform. 2006. PMID: 16139564 Review.
Cited by
-
Towards answering biological questions with experimental evidence: automatically identifying text that summarize image content in full-text articles.AMIA Annu Symp Proc. 2006;2006:834-8. AMIA Annu Symp Proc. 2006. PMID: 17238458 Free PMC article.
-
AskHERMES: An online question answering system for complex clinical questions.J Biomed Inform. 2011 Apr;44(2):277-88. doi: 10.1016/j.jbi.2011.01.004. Epub 2011 Jan 21. J Biomed Inform. 2011. PMID: 21256977 Free PMC article.
-
Full text clustering and relationship network analysis of biomedical publications.PLoS One. 2014 Sep 24;9(9):e108847. doi: 10.1371/journal.pone.0108847. eCollection 2014. PLoS One. 2014. PMID: 25250864 Free PMC article.
-
Automated systems to identify relevant documents in product risk management.BMC Med Inform Decis Mak. 2012 Mar 2;12:13. doi: 10.1186/1472-6947-12-13. BMC Med Inform Decis Mak. 2012. PMID: 22380483 Free PMC article.
-
Clustering cliques for graph-based summarization of the biomedical research literature.BMC Bioinformatics. 2013 Jun 7;14:182. doi: 10.1186/1471-2105-14-182. BMC Bioinformatics. 2013. PMID: 23742159 Free PMC article.
References
-
- Smink LJ, Helton EM, Healy BC, Cavnor CC, Lam AC, Flamez D, Burren OS, Wang Y, Dolman GE, Burdick DB, Everett VH, Glusman G, Laneri D, Rowen L, Schuilenburg H, Walker NM, Mychaleckyj J, Wicker LS, Eizirik DL, Todd JA, Goodman N. T1DBase, a community web-based resource for type 1 diabetes research. Nucleic Acids Res. 2005;33:D544–9. doi: 10.1093/nar/gki095. - DOI - PMC - PubMed
-
- Yu H, Hatzivassiloglou V. Towards answering opinion questions: Separating facts from opinions and identifying the polarity of opinion sentences. 2003.
-
- Joachims T. Text categorization with support vector machines: Learning with many relevant features. 1998. pp. 137–142.
-
- Wilbur WJ. A thematic analysis of the AIDS literature. Pac Symp Biocomput. 2002:386–397. - PubMed
-
- Hearst M. The BioText project. A powerpoint presentation.
Publication types
MeSH terms
LinkOut - more resources
Full Text Sources
Other Literature Sources