. 2024 Mar 15:26:e47923.

doi: 10.2196/47923.

Methods and Annotated Data Sets Used to Predict the Gender and Age of Twitter Users: Scoping Review

Karen O'Connor¹, Su Golder², Davy Weissenbacher³, Ari Z Klein¹, Arjun Magge¹, Graciela Gonzalez-Hernandez³

Affiliations

¹ Department of Biostatistics, Epidemiology and Informatics, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, United States.
² Department of Health Sciences, University of York, York, United Kingdom.
³ Department of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA, United States.

PMID: 38488839
PMCID: PMC10980991
DOI: 10.2196/47923

Methods and Annotated Data Sets Used to Predict the Gender and Age of Twitter Users: Scoping Review

Karen O'Connor et al. J Med Internet Res. 2024.

. 2024 Mar 15:26:e47923.

doi: 10.2196/47923.

Authors

Karen O'Connor¹, Su Golder², Davy Weissenbacher³, Ari Z Klein¹, Arjun Magge¹, Graciela Gonzalez-Hernandez³

Affiliations

¹ Department of Biostatistics, Epidemiology and Informatics, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, United States.
² Department of Health Sciences, University of York, York, United Kingdom.
³ Department of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA, United States.

PMID: 38488839
PMCID: PMC10980991
DOI: 10.2196/47923

Abstract

Background: Patient health data collected from a variety of nontraditional resources, commonly referred to as real-world data, can be a key information source for health and social science research. Social media platforms, such as Twitter (Twitter, Inc), offer vast amounts of real-world data. An important aspect of incorporating social media data in scientific research is identifying the demographic characteristics of the users who posted those data. Age and gender are considered key demographics for assessing the representativeness of the sample and enable researchers to study subgroups and disparities effectively. However, deciphering the age and gender of social media users poses challenges.

Objective: This scoping review aims to summarize the existing literature on the prediction of the age and gender of Twitter users and provide an overview of the methods used.

Methods: We searched 15 electronic databases and carried out reference checking to identify relevant studies that met our inclusion criteria: studies that predicted the age or gender of Twitter users using computational methods. The screening process was performed independently by 2 researchers to ensure the accuracy and reliability of the included studies.

Results: Of the initial 684 studies retrieved, 74 (10.8%) studies met our inclusion criteria. Among these 74 studies, 42 (57%) focused on predicting gender, 8 (11%) focused on predicting age, and 24 (32%) predicted a combination of both age and gender. Gender prediction was predominantly approached as a binary classification task, with the reported performance of the methods ranging from 0.58 to 0.96 F₁-score or 0.51 to 0.97 accuracy. Age prediction approaches varied in terms of classification groups, with a higher range of reported performance, ranging from 0.31 to 0.94 F₁-score or 0.43 to 0.86 accuracy. The heterogeneous nature of the studies and the reporting of dissimilar performance metrics made it challenging to quantitatively synthesize results and draw definitive conclusions.

Conclusions: Our review found that although automated methods for predicting the age and gender of Twitter users have evolved to incorporate techniques such as deep neural networks, a significant proportion of the attempts rely on traditional machine learning methods, suggesting that there is potential to improve the performance of these tasks by using more advanced methods. Gender prediction has generally achieved a higher reported performance than age prediction. However, the lack of standardized reporting of performance metrics or standard annotated corpora to evaluate the methods used hinders any meaningful comparison of the approaches. Potential biases stemming from the collection and labeling of data used in the studies was identified as a problem, emphasizing the need for careful consideration and mitigation of biases in future studies. This scoping review provides valuable insights into the methods used for predicting the age and gender of Twitter users, along with the challenges and considerations associated with these methods.

Keywords: Twitter; age; age prediction; demographics; gender; gender prediction; machine learning; neural network; prediction; real-world data; social media.

©Karen O'Connor, Su Golder, Davy Weissenbacher, Ari Z Klein, Arjun Magge, Graciela Gonzalez-Hernandez. Originally published in the Journal of Medical Internet Research (https://www.jmir.org), 15.03.2024.

PubMed Disclaimer

Conflict of interest statement

Conflicts of Interest: None declared.

Figures

**Figure 1**
PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) flow diagram of the included studies.

**Figure 2**
Summary of recommendations for best practices in the collection of training data, and the development and dissemination of age or gender machine learning prediction models.

See this image and copyright information in PMC

Cited by

Leveraging Natural Language Processing and Machine Learning Methods for Adverse Drug Event Detection in Electronic Health/Medical Records: A Scoping Review.
Golder S, Xu D, O'Connor K, Wang Y, Batra M, Hernandez GG. Golder S, et al. Drug Saf. 2025 Apr;48(4):321-337. doi: 10.1007/s40264-024-01505-6. Epub 2025 Jan 9. Drug Saf. 2025. PMID: 39786481 Free PMC article.

References

1. Real-world evidence. U.S. Food and Drug Administration. [2023-03-30]. https://www.fda.gov/science-research/science-and-research-special-topics... .
1. Alessa A, Faezipour M. A review of influenza detection and prediction through social networking sites. Theor Biol Med Model. 2018 Feb 01;15(1):2. doi: 10.1186/s12976-017-0074-5. https://tbiomed.biomedcentral.com/articles/10.1186/s12976-017-0074-5 10.1186/s12976-017-0074-5 - DOI - DOI - PMC - PubMed
1. Bisanzio D, Kraemer MU, Bogoch II, Brewer T, Brownstein JS, Reithinger R. Use of Twitter social media activity as a proxy for human mobility to predict the spatiotemporal spread of COVID-19 at global scale. Geospat Health. 2020 Jun 15;15(1) doi: 10.4081/gh.2020.882. doi: 10.4081/gh.2020.882. - DOI - DOI - PubMed
1. Magge A, Tutubalina E, Miftahutdinov Z, Alimova I, Dirkson A, Verberne S, Weissenbacher D, Gonzalez-Hernandez G. DeepADEMiner: a deep learning pharmacovigilance pipeline for extraction and normalization of adverse drug event mentions on Twitter. J Am Med Inform Assoc. 2021 Sep 18;28(10):2184–92. doi: 10.1093/jamia/ocab114. https://europepmc.org/abstract/MED/34270701 6322900 - DOI - PMC - PubMed
1. Nikfarjam A, Sarker A, O'Connor K, Ginn R, Gonzalez G. Pharmacovigilance from social media: mining adverse drug reaction mentions using sequence labeling with word embedding cluster features. J Am Med Inform Assoc. 2015 May;22(3):671–81. doi: 10.1093/jamia/ocu041. https://europepmc.org/abstract/MED/25755127 ocu041 - DOI - PMC - PubMed

Publication types

Actions
Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions

Grants and funding

R01 LM011176/LM/NLM NIH HHS/United States

LinkOut - more resources

Full Text Sources
Miscellaneous
- NCI CPTAC Assay Portal

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Methods and Annotated Data Sets Used to Predict the Gender and Age of Twitter Users: Scoping Review

Affiliations

Methods and Annotated Data Sets Used to Predict the Gender and Age of Twitter Users: Scoping Review

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

Similar articles

Cited by

References

Publication types

MeSH terms

Grants and funding

LinkOut - more resources

Full Text Sources

Miscellaneous

Abstract

Conflict of interest statement

Figures

Similar articles

Cited by

References

Publication types

MeSH terms

Related information

Grants and funding

LinkOut - more resources

Full Text Sources

Miscellaneous