SYNTOLOGY HomeExplorerAtlasCodeMethodologyAboutDevelopersFeedPricing
Paper · 2310.17586 · EMNLP · 2023

Global Voices, Local Biases: Socio-Cultural Prejudices across Languages

Antonios Anastasopoulos, Ziwei Zhu, Chahat Raj, Anjishnu Mukherjee

arXiv · PDF · Open in the Atlas

Code that ran

We lifted 7 functions out of this paper's own repositories and ran 1 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.

RepositoryRoleRan
iamshnoo/weathub canonical 1 of 7
FunctionStatusWhere it lives
embedding_variance Ran iamshnoo/weathub/src/compare_embeddings.py
pointer only (licence: GPL-3.0) · get_code("c8721af1c24e064f")
get_weat_google_human Not yet run iamshnoo/weathub/src/load_annotations.py
pointer only (licence: GPL-3.0) · get_code("78828d54d4541067")
load_hf_tokenizer_model Not yet run iamshnoo/weathub/src/encoding_utils.py
pointer only (licence: GPL-3.0) · get_code("d77adb941faf57ff")
process_annotation_file Not yet run iamshnoo/weathub/src/load_annotations.py
pointer only (licence: GPL-3.0) · get_code("500e9ab633684d53")
process_sheet_one Not yet run iamshnoo/weathub/src/load_annotations.py
pointer only (licence: GPL-3.0) · get_code("67f8b72607a187d2")
save_results_to_df Not yet run iamshnoo/weathub/src/run_weat.py
pointer only (licence: GPL-3.0) · get_code("1eab4d0c275953ef")
save_results_to_df Not yet run iamshnoo/weathub/src/valence_weat.py
pointer only (licence: GPL-3.0) · get_code("b8ea64c8aee135eb")

Repositories linked to this paper

Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.

Abstract

Human biases are ubiquitous but not uniform: disparities exist across linguistic, cultural, and societal borders. As large amounts of recent literature suggest, language models (LMs) trained on human data can reflect and often amplify the effects of these social biases. However, the vast majority of existing studies on bias are heavily skewed towards Western and European languages. In this work, we scale the Word Embedding Association Test (WEAT) to 24 languages, enabling broader studies and yielding interesting findings about LM bias. We additionally enhance this data with culturally relevant information for each language, capturing local contexts on a global scale. Further, to encompass more widely prevalent societal biases, we examine new bias dimensions across toxicity, ableism, and more. Moreover, we delve deeper into the Indian linguistic landscape, conducting a comprehensive regional bias analysis across six prevalent Indian languages. Finally, we highlight the significance of these social biases and the new dimensions through an extensive comparison of embedding methods, reinforcing the need to address them in pursuit of more equitable language models. 1 * Equal contribution 1 All code, data and results are available here: https:// github.com/iamshnoo/weathub.

For agents

The same record, over MCP at https://syntology.ai/mcp:

get_harvested_code_for_paper("2310.17586")
get_code_for_paper("2310.17586")
have("2310.17586")

Connect an agent — have() is free.