SYNTOLOGY HomeExplorerAtlasCodeMethodologyAboutDevelopersFeedPricing
Paper · 2410.10303 · NeurIPS · 2024

A Comparative Study of Translation Bias and Accuracy in Multilingual Large Language Models for Cross-Language Claim Verification

Gary Sun, Kevin Zhu, Jonathan Lu, Aryan Singhal, Veronica Shao, Ryan Ding

arXiv · PDF · Open in the Atlas

Code that ran

We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.

Repositories linked to this paper

Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.

Abstract

The rise of digital misinformation has heightened interest in using multilingual Large Language Models (LLMs) for fact-checking. This study systematically evaluates translation bias and the effectiveness of LLMs for cross-lingual claim verification across 15 languages from five language families: Romance, Slavic, Turkic, Indo-Aryan, and Kartvelian. Using the XFACT dataset to assess their impact on accuracy and bias, we investigate two distinct translation methods: pretranslation and self-translation. We use mBERT's performance on the English dataset as a baseline to compare language-specific accuracies. Our findings reveal that low-resource languages exhibit significantly lower accuracy in direct inference due to underrepresentation in the training data. Furthermore, larger models demonstrate superior performance in self-translation, improving translation accuracy and reducing bias. These results highlight the need for balanced multilingual training, especially in low-resource languages, to promote equitable access to reliable fact-checking tools and minimize the risk of spreading misinformation in different linguistic contexts. Multilingual Large Language Models (LLMs), such as GPT-4 and Llama 3.1, have shown remarkable capabilities in various languages and tasks [Ahuja et al., 2024]. Thus, there has been increasing interest in possible usages of LLMs for claim verification across languages [Panchendrarajan and Zubiaga, 2024]. However, recent studies have revealed significant disparities in their performance and bias in different languages [Xu et al., 2024, Huang et al., 2024]. This variability is especially concerning given the importance of claim verification in combating misinformation [Sundriyal et al., 2023]. The performance discrepancies observed in LLMs often favor resource-rich languages like English, French, and German over resource-poor languages such as Kannada and Occitan [Robinson et al., 2023, Bawden and Yvon, 2023, Quelle and Bovet, 2024]. These differences stem from variations in accuracy and translation quality between languages. Although LLMs demonstrate impressive average performance in a wide range of languages, Li et al. [2024] highlights persistent gaps between high-resource and low-resource languages, emphasizing the need for more balanced data collection and training approaches. Addressing misinformation for claim verification tasks is critical, as ineffective claim verification can spread false information between languages and vulnerable populations [Thorne and Vlachos, 2018]. Although advances in LLMs, such as Meta's Llama 3.1 models [Dubey et al., 2024], have improved multilingual capabilities, reliance on external translation methods in some contexts-especially by users or systems that use third-party services such as Google Translate or that rely on the LLM Preprint. Under review.

For agents

The same record, over MCP at https://syntology.ai/mcp:

get_harvested_code_for_paper("2410.10303")
get_code_for_paper("2410.10303")
have("2410.10303")

Connect an agent — have() is free.