Cultural Fidelity in Large-Language Models: An Evaluation of Online Language Resources as a Driver of Model Performance in Value Representation

Abstract

The training data for LLMs embeds societal values, increasing theirfamiliarity with the language's culture. Our analysis found that 44% of thevariance in the ability of GPT-4o to reflect the societal values of a country,as measured by the World Values Survey, correlates with the availability ofdigital resources in that language. Notably, the error rate was more than fivetimes higher for the languages of the lowest resource compared to the languagesof the highest resource. For GPT-4-turbo, this correlation rose to 72%,suggesting efforts to improve the familiarity with the non-English languagebeyond the web-scraped data. Our study developed one of the largest and mostrobust datasets in this topic area with 21 country-language pairs, each ofwhich contain 94 survey questions verified by native speakers. Our resultshighlight the link between LLM performance and digital data availability intarget languages. Weaker performance in low-resource languages, especiallyprominent in the Global South, may worsen digital divides. We discussstrategies proposed to address this, including developing multilingual LLMsfrom the ground up and enhancing fine-tuning on diverse linguistic datasets, asseen in African language initiatives.

Quick Read (beta)

loading the full paper ...