PAPER DIGEST
Most Influential ACL 2021 Paper · 2026-03 edition

StereoSet: Measuring Stereotypical Bias in Pretrained Language Models

Moin Nadeem; Anna Bethke; Siva Reddy

Venue
Annual Meeting of the Association for Computational Linguistics (ACL) 2021
Recognition
Most Influential ACL 2021 Paper (Rank No. 3)
Edition
2026-03
Impact factor
8
Certificate ID
976e7e315eb006cc

Abstract

A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or African Americans are athletic. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on large real-world data, they are known to capture stereotypical biases. It is important to quantify to what extent these biases are present in them. Although this is a rapidly growing area of research, existing literature lacks in two important aspects: 1) they mainly evaluate bias of pretrained language models on a small set of artificial sentences, even though these models are trained on natural data 2) current evaluations focus on measuring bias without considering the language modeling ability of a model, which could lead to misleading trust on a model even if it is a poor language model. We address both these problems. We present StereoSet, a large-scale natural English dataset to measure stereotypical biases in four domains: gender, profession, race, and religion. We contrast both stereotypical bias and language modeling ability of popular models like BERT, GPT-2, RoBERTa, and XLnet. We show that these models exhibit strong stereotypical biases. Our data and code are available at https://stereoset.mit.edu.

Download PDF certificate