Text classification models for the automatic detection of nonmedical prescription medication use from social media
- PMID: 33499852
- PMCID: PMC7835447
- DOI: 10.1186/s12911-021-01394-0
Text classification models for the automatic detection of nonmedical prescription medication use from social media
Abstract
Background: Prescription medication (PM) misuse/abuse has emerged as a national crisis in the United States, and social media has been suggested as a potential resource for performing active monitoring. However, automating a social media-based monitoring system is challenging-requiring advanced natural language processing (NLP) and machine learning methods. In this paper, we describe the development and evaluation of automatic text classification models for detecting self-reports of PM abuse from Twitter.
Methods: We experimented with state-of-the-art bi-directional transformer-based language models, which utilize tweet-level representations that enable transfer learning (e.g., BERT, RoBERTa, XLNet, AlBERT, and DistilBERT), proposed fusion-based approaches, and compared the developed models with several traditional machine learning, including deep learning, approaches. Using a public dataset, we evaluated the performances of the classifiers on their abilities to classify the non-majority "abuse/misuse" class.
Results: Our proposed fusion-based model performs significantly better than the best traditional model (F1-score [95% CI]: 0.67 [0.64-0.69] vs. 0.45 [0.42-0.48]). We illustrate, via experimentation using varying training set sizes, that the transformer-based models are more stable and require less annotated data compared to the other models. The significant improvements achieved by our best-performing classification model over past approaches makes it suitable for automated continuous monitoring of nonmedical PM use from Twitter.
Conclusions: BERT, BERT-like and fusion-based models outperform traditional machine learning and deep learning models, achieving substantial improvements over many years of past research on the topic of prescription medication misuse/abuse classification from social media, which had been shown to be a complex task due to the unique ways in which information about nonmedical use is presented. Several challenges associated with the lack of context and the nature of social media language need to be overcome to further improve BERT and BERT-like models. These experimental driven challenges are represented as potential future research directions.
Keywords: Machine learning; Natural language processing; Prescription medication misuse; Social media.
Conflict of interest statement
The authors declare no competing interests.
Figures






Similar articles
-
Promoting Reproducible Research for Characterizing Nonmedical Use of Medications Through Data Annotation: Description of a Twitter Corpus and Guidelines.J Med Internet Res. 2020 Feb 26;22(2):e15861. doi: 10.2196/15861. J Med Internet Res. 2020. PMID: 32130117 Free PMC article.
-
Mining social media for prescription medication abuse monitoring: a review and proposal for a data-centric framework.J Am Med Inform Assoc. 2020 Feb 1;27(2):315-329. doi: 10.1093/jamia/ocz162. J Am Med Inform Assoc. 2020. PMID: 31584645 Free PMC article. Review.
-
Detecting Potentially Harmful and Protective Suicide-Related Content on Twitter: Machine Learning Approach.J Med Internet Res. 2022 Aug 17;24(8):e34705. doi: 10.2196/34705. J Med Internet Res. 2022. PMID: 35976193 Free PMC article.
-
Machine Learning and Natural Language Processing for Geolocation-Centric Monitoring and Characterization of Opioid-Related Social Media Chatter.JAMA Netw Open. 2019 Nov 1;2(11):e1914672. doi: 10.1001/jamanetworkopen.2019.14672. JAMA Netw Open. 2019. PMID: 31693125 Free PMC article.
-
Natural Language Processing Applications in the Clinical Neurosciences: A Machine Learning Augmented Systematic Review.Acta Neurochir Suppl. 2022;134:277-289. doi: 10.1007/978-3-030-85292-4_32. Acta Neurochir Suppl. 2022. PMID: 34862552
Cited by
-
Automatic Detection of Twitter Users Who Express Chronic Stress Experiences via Supervised Machine Learning and Natural Language Processing.Comput Inform Nurs. 2023 Sep 1;41(9):717-724. doi: 10.1097/CIN.0000000000000985. Comput Inform Nurs. 2023. PMID: 36445331 Free PMC article.
-
Discerning conversational context in online health communities for personalized digital behavior change solutions using Pragmatics to Reveal Intent in Social Media (PRISM) framework.J Biomed Inform. 2023 Apr;140:104324. doi: 10.1016/j.jbi.2023.104324. Epub 2023 Feb 24. J Biomed Inform. 2023. PMID: 36842490 Free PMC article.
-
Can accurate demographic information about people who use prescription medications nonmedically be derived from Twitter?Proc Natl Acad Sci U S A. 2023 Feb 21;120(8):e2207391120. doi: 10.1073/pnas.2207391120. Epub 2023 Feb 14. Proc Natl Acad Sci U S A. 2023. PMID: 36787355 Free PMC article.
-
Large-Scale Social Media Analysis Reveals Emotions Associated with Nonmedical Prescription Drug Use.Health Data Sci. 2022;2022:9851989. doi: 10.34133/2022/9851989. Epub 2022 Apr 27. Health Data Sci. 2022. PMID: 37621877 Free PMC article.
-
Task-Specific Transformer-Based Language Models in Health Care: Scoping Review.JMIR Med Inform. 2024 Nov 18;12:e49724. doi: 10.2196/49724. JMIR Med Inform. 2024. PMID: 39556827 Free PMC article.
References
-
- National Institute on Drug Abuse. Misuse of Prescription Drugs. 2018 Dec.
-
- Schepis TS. The prescription drug abuse epidemic : incidence, treatment, prevention, and policy. 1st ed. Praeger; 2018.
-
- Hedegaard H, Miniño AM, Warner M. Drug Overdose Deaths in the United States, 1999–2018 Key findings Data from the National Vital Statistics System, Mortality. 2020 Jan.
-
- Centers for Disease Control and Prevention. Wide-ranging online data for epidemiologic research (WONDER). 2020.
-
- What States Need to Know about PDMPs | Drug Overdose | CDC Injury Center.
Publication types
MeSH terms
Substances
Grants and funding
LinkOut - more resources
Full Text Sources
Other Literature Sources
Miscellaneous