Artificial Intelligence · 25.08.2026, 07:01 UTC
Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.AI ↗ |
| Veröffentlicht | 25.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.22335v1 Announce Type: cross Abstract: Bengali is the seventh-most-spoken language globally, yet LLM safety evaluation remains overwhelmingly English-centric. We introduce BanglaSafe, a benchmark of 879 Bengali prompts combining 309 natively authored prompts with 570 expert-reviewed prompts, spanning 17 culturally grounded harm categories and five prompting conditions that vary language, writing style, and authority framing. Evaluating 18 frontier LLMs, we find that over half of all responses are unsafe or partially unsafe (53.6%) while 14.7% contains strictly harmful content, and that the strongest observed effect is not the switch from English to Bengali but the choice of writing style within Bengali: the same harmful request phrased as a formal newspaper investigation succeeds 17 percentage points more often than the same request phrased as a casual message, with no adversarial engineering involved. We further show that existing safety classifiers struggle to reliably evaluate Bengali content, with even frontier models failing on nearly half of all cases.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Adversarial Entropy Inflation Against Gumbel-Based Inference Verification
- info The Emergence of Relevance Through Axiomatic Attention Patterns During LoRA Fine-Tuning
- info Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents
- info Mycelial Search: A Graph-Structured Metaheuristic for Continuous Optimisation