Abstract
Despite the growing reliance on fairness benchmarks to evaluate languagemodels, the datasets that underpin these benchmarks remain criticallyunderexamined. This survey addresses that overlooked foundation by offering acomprehensive analysis of the most widely used fairness datasets in languagemodel research. To ground this analysis, we characterize each dataset acrosskey dimensions, including provenance, demographic scope, annotation design, andintended use, revealing the assumptions and limitations baked into currentevaluation practices. Building on this foundation, we propose a unifiedevaluation framework that surfaces consistent patterns of demographicdisparities across benchmarks and scoring metrics. Applying this framework tosixteen popular datasets, we uncover overlooked biases that may distortconclusions about model fairness and offer guidance on selecting, combining,and interpreting these resources more effectively and responsibly. Our findingshighlight an urgent need for new benchmarks that capture a broader range ofsocial contexts and fairness notions. To support future research, we releaseall data, code, and results athttps://github.com/vanbanTruong/Fairness-in-Large-Language-Models/tree/main/datasets,fostering transparency and reproducibility in the evaluation of language modelfairness.