AI Safety Built in the West Fails Users Everywhere

The recent pause by OpenAI on model training due to safety concerns highlights a critical flaw in AI development: safety measures are largely designed in the West, failing users in other regions. While AI excels in solving complex problems in wealthy nations, it struggles with basic tasks in developing countries, where trust and safety teams are absent and cultural contexts are ignored. This disparity creates a new kind of AI divide, where English speakers are safer than those using low-resource languages.
The Safety Gap in Non-Western Languages
OpenAI's voluntary pause, prompted by models hacking websites, underscores the rapid pace of AI development. However, the real issue lies in the deployment risks that are often overlooked. Trust and safety teams, concentrated in Silicon Valley, fail to account for languages like Tigrinya, where machine translation errors can be life-threatening. For instance, 'smallpox' was rendered as 'syphilis' and 'intravenous antibiotics' as 'intravenous insecticides.' Such mistranslations in health queries can lead to incorrect diagnoses and treatment decisions, disproportionately affecting users in low- and middle-income countries.
Who Defines Safety?
Elizabeth Orembo of Research ICT Africa points out that the core question is who gets to define what constitutes a safety problem. Big tech firms focus on model risks like deception and cyber capabilities, but pay little attention to deployment risks such as discrimination, exclusion, and language failures. The evaluations they run assume reliable infrastructure and legal systems, which are often absent in developing nations. As a result, a model can pass every frontier safety evaluation and still produce unsafe outcomes when deployed in these contexts.
The New AI Divide
Dhanaraj Thakur from George Washington University Law School notes that guardrails may work well in English but fail in low-resource languages, leading to higher hallucination rates and poor translation. This creates a new kind of AI divide, where English speakers are safer than those who speak other languages. The consequences are immediate: AI-powered facial recognition and ID systems can deny wages, meals, or school attendance, and tools that mistranslate local terms can misidentify crops, affecting livelihoods.
Calls for Action and the Road Ahead
Governments are beginning to address these issues, with the Bletchley Declaration and India's AI summit focusing on safety. China has proposed mechanisms to manage AI risks in developing nations. However, Urvashi Aneja of Digital Futures Lab warns that without investment in safety infrastructure now, public trust will erode, deepening inequality. Over 1,300 AI employees have signed an open letter urging for time to address emerging risks. The story of Sumiya Khan, a 16-year-old in New Delhi who trusted ChatGPT's advice and delayed seeing a doctor for anemia, illustrates the real-world stakes. Her mother said, 'We trusted the AI because it sounded so convincing. That made us wait longer than we should have.'
Key Takeaways
- AI safety frameworks are designed in the West, ignoring non-Western languages and contexts, leading to harmful errors.
- Mistranslations in health queries can be life-threatening, as seen in Tigrinya and Hindi examples.
- The AI divide is widening: English speakers are safer than users of low-resource languages.
- Companies focus on model risks, neglecting deployment risks like discrimination and language failures.
- Without urgent investment in safety infrastructure, public trust in AI will erode in developing nations.
Source: Rest of World • 🌍
Keep Reading


