Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2021 Jan 26;21(1):27.
doi: 10.1186/s12911-021-01394-0.

Text classification models for the automatic detection of nonmedical prescription medication use from social media

Affiliations

Text classification models for the automatic detection of nonmedical prescription medication use from social media

Mohammed Ali Al-Garadi et al. BMC Med Inform Decis Mak. .

Abstract

Background: Prescription medication (PM) misuse/abuse has emerged as a national crisis in the United States, and social media has been suggested as a potential resource for performing active monitoring. However, automating a social media-based monitoring system is challenging-requiring advanced natural language processing (NLP) and machine learning methods. In this paper, we describe the development and evaluation of automatic text classification models for detecting self-reports of PM abuse from Twitter.

Methods: We experimented with state-of-the-art bi-directional transformer-based language models, which utilize tweet-level representations that enable transfer learning (e.g., BERT, RoBERTa, XLNet, AlBERT, and DistilBERT), proposed fusion-based approaches, and compared the developed models with several traditional machine learning, including deep learning, approaches. Using a public dataset, we evaluated the performances of the classifiers on their abilities to classify the non-majority "abuse/misuse" class.

Results: Our proposed fusion-based model performs significantly better than the best traditional model (F1-score [95% CI]: 0.67 [0.64-0.69] vs. 0.45 [0.42-0.48]). We illustrate, via experimentation using varying training set sizes, that the transformer-based models are more stable and require less annotated data compared to the other models. The significant improvements achieved by our best-performing classification model over past approaches makes it suitable for automated continuous monitoring of nonmedical PM use from Twitter.

Conclusions: BERT, BERT-like and fusion-based models outperform traditional machine learning and deep learning models, achieving substantial improvements over many years of past research on the topic of prescription medication misuse/abuse classification from social media, which had been shown to be a complex task due to the unique ways in which information about nonmedical use is presented. Several challenges associated with the lack of context and the nature of social media language need to be overcome to further improve BERT and BERT-like models. These experimental driven challenges are represented as potential future research directions.

Keywords: Machine learning; Natural language processing; Prescription medication misuse; Social media.

PubMed Disclaimer

Conflict of interest statement

The authors declare no competing interests.

Figures

Fig. 1
Fig. 1
Class distribution in the training, validation, and test datasets
Fig. 2
Fig. 2
Sample sequence embeddings showing how the term “drugs” is represented differently when appearing in different sequences
Fig. 3
Fig. 3
Comparison between the best models from each learning category (A F1-score Class A)
Fig. 4
Fig. 4
Learning curve at different amount of training data used for training
Fig. 5
Fig. 5
20 TF-IDF bigram word clouds for automatically classified tweets from each of the 4 categories
Fig. 6
Fig. 6
Confusion Matrix Fusion 1 based Model (0 = Class A, 1 = Class C, 2 = Class M, 3 = Class U)

Similar articles

Cited by

References

    1. National Institute on Drug Abuse. Misuse of Prescription Drugs. 2018 Dec.
    1. Schepis TS. The prescription drug abuse epidemic : incidence, treatment, prevention, and policy. 1st ed. Praeger; 2018.
    1. Hedegaard H, Miniño AM, Warner M. Drug Overdose Deaths in the United States, 1999–2018 Key findings Data from the National Vital Statistics System, Mortality. 2020 Jan.
    1. Centers for Disease Control and Prevention. Wide-ranging online data for epidemiologic research (WONDER). 2020.
    1. What States Need to Know about PDMPs | Drug Overdose | CDC Injury Center.

Publication types

Substances

LinkOut - more resources