Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2025 Dec 15.
doi: 10.1038/s41562-025-02360-w. Online ahead of print.

Multimodal large language models can make context-sensitive hate speech evaluations aligned with human judgement

Affiliations

Multimodal large language models can make context-sensitive hate speech evaluations aligned with human judgement

Thomas Davidson. Nat Hum Behav. .

Abstract

Multimodal large language models (MLLMs) could enhance the accuracy of automated content moderation by integrating contextual information. This study examines how MLLMs evaluate hate speech through a series of conjoint experiments. Models are provided with a hate speech policy and shown simulated social media posts that systematically vary in slur usage, user demographics and other attributes. The decisions from MLLMs are benchmarked against judgements by human participants (n = 1,854). The results demonstrate that larger, more advanced models can make context-sensitive evaluations that are closely aligned with human judgement. However, pervasive demographic and lexical biases remain, particularly among smaller models. Further analyses show that context sensitivity can be amplified via prompting but not eliminated, and that some models are especially responsive to visual identity cues. These findings highlight the benefits and risks of using MLLMs for content moderation and demonstrate the utility of conjoint experiments for auditing artificial intelligence in complex, context-dependent applications.

PubMed Disclaimer

Conflict of interest statement

Competing interests: The author declares no competing interests.

References

    1. Grimmelmann, J. The virtues of moderation. Yale J. Law Technol. 17, 42–109 (2015).
    1. Klonick, K. The new governors: the people, rules, and processes governing online speech. Harv. Law Rev. 131, 1598–1670 (2018).
    1. Gillespie, T. Custodians of the Internet: Platforms, Content Moderation, and the Hidden Decisions That Shape Social Media (Yale Univ. Press, 2018).
    1. Roberts, S. T. Behind the Screen (Yale Univ. Press, 2019).
    1. Kaye, D. A. Speech Police: The Global Struggle to Govern the Internet (Columbia Global Reports, 2019).

LinkOut - more resources