Skip to main navigation Skip to search Skip to main content

Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili

Research output: Chapter in Book/Report/Conference proceedingConference contribution

2 Downloads (Pure)

Abstract

Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric. We investigate whether social biases generalise across languages by submitting 4,900 symmetric English–Swahili prompt pairs to GPT-5.2 and Gemini 2.5 Flash across nine demographic bias axes, yielding 19,600 completions evaluated for stereotype prevalence, sentiment, refusal behaviour, and cross-lingual semantic similarity. Our findings show that bias transforms rather than transfers: stereotype rates shifted by up to 12 percentage points on specific axes, Gemini’s neutral-sentiment rate doubled in Swahili, and GPT-5.2 refused 169 prompts in English and zero in Swahili, indicating safety mechanisms functionally anchored to English-language tokens. Over 55% of prompt pairs produced semantically dissimilar completions across both models. These reinforce the idea that English-only bias audits do not produce adequate coverage for multilingual deployment.
Original languageEnglish
Title of host publicationProceedings of the 1st Workshop on Multilinguality in the Era of Large Language Models (MeLLM 2026)
EditorsKaiyu Huang, Fengran Mo, Pinzhen Chen, Meng Jiang
Place of PublicationSan Diego, United States
PublisherAssociation for Computational Linguistics, ACL
Pages181-190
Number of pages10
ISBN (Print)9798891764309
DOIs
Publication statusE-pub ahead of print - 4 Jul 2026
Event1st Workshop on Multilinguality in the Era of Large Language Models - Grand Hyatt Manchester San Diego, San Diego, United States
Duration: 4 Jul 20264 Jul 2026
https://mellm.org/

Conference

Conference1st Workshop on Multilinguality in the Era of Large Language Models
Abbreviated titleMeLLM @ ACL 2026
Country/TerritoryUnited States
CitySan Diego
Period4/07/264/07/26
Internet address

Bibliographical note

Anthology ID: 2026.mellm-1.17

Fingerprint

Dive into the research topics of 'Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili'. Together they form a unique fingerprint.

Cite this