ScopeBench: Evaluating Cultural Norm Scope Awareness in LLMs
Haojun Liang, Ziwei Zhao, Chenyu Li, and 8 more authors
In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), Main Conference, 2026
A cultural-scope safety benchmark that treats norm scope as a first-class safety dimension. We build a dual-anchored annotation protocol across 49 countries and 9 cultural clusters (3,323 norms), define three scope error types (Inflation, Deflation, Displacement), and evaluate 21 frontier LLMs across 7 model families.