Google DeepMind researcher Lora Aroyo says AI benchmarks are failing to test how models handle local cultures, laws and norms. At a conference in South Korea, she argued that fluent answers can disguise cultural errors, with consequences well beyond awkward etiquette.
Google DeepMind Watch analysis
What happened
Aroyo said that while AI models support about 100 languages, major benchmarks typically assess only 10 to 20. She also said evaluations that account for cultural differences make up less than a fifth of the total. Those figures come from her keynote, as reported by Chosunbiz.
Her examples span traditions, religion, history and geopolitics. A model might, for instance, say it is fine to stick chopsticks upright in rice, despite that being considered poor etiquette in several Asian countries. Aroyo also cited translation tests in which answer accuracy reportedly fell to 50 per cent for Malay and 40 per cent for Tamil.
Why it matters
A model can produce smooth prose in a language and still miss the context that makes an answer appropriate, accurate or safe. That matters when systems are used across different communities, where local laws and customs are not decorative details.
Aroyo’s argument is that evaluation needs to begin in the language and cultural setting being tested, rather than simply translating an English question and calling the job done. That is a practical challenge for developers: more languages on a model’s menu do not automatically mean better understanding.
Our read
This is a useful corrective to the idea that multilingual fluency is a tidy proxy for cultural competence. Aroyo’s reported figures make the gap worth examining, though they are claims from her keynote, not a complete independent audit of every benchmark. The next test is whether better-designed evaluations lead to models that make fewer consequential mistakes, not merely more culturally polished ones.
What to watch
- Whether researchers publish the underlying evaluations and methods behind the cited figures.
- Whether benchmarks test cultural understanding in target languages, rather than relying mainly on translated prompts.
- Whether developers show measurable improvements across regions, not just broader language coverage.
Discussion spark: Should AI companies be judged by how many languages their models support, or by evidence that the models understand the local context in each one?
Sources and evidence
- DeepMind urges culture-aware AI tests as English bias fuels blind spots – CHOSUNBIZ – Chosunbiz (6 October 2026, 05:03 UTC)
not affiliated with, endorsed by, or operated by Google or Google DeepMind