Artificial Intelligence · 01.09.2026, 17:02 UTC
MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.LG ↗ |
| Veröffentlicht | 01.09.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2603.02482v2 Announce Type: replace Abstract: Safety evaluation of multimodal large language models requires tracking not only whether an attack succeeds, but also how the interaction unfolds across turns and input modalities. We present MUSE (Multimodal Unified Safety Evaluation), an open-source, browser-based, run-centric platform for multimodal safety evaluation. MUSE treats each attack run as the persistent unit of execution, inspection, and analysis, preserving its configuration, multi-turn trajectory, delivered modalities and media, target responses, and safety judgments. A five-level response taxonomy further distinguishes full Compliance from Partial Compliance and refusal behavior, yielding hard ASR, soft ASR, and gray-zone width (GZW). Across 11,700 evaluations on six multimodal LLMs, direct text-only requests yield only 3.1% macro hard ASR and 4.4% soft ASR, while iterative attack procedures are substantially more effective. Attack effectiveness also varies substantially with the attacker backbone. In contrast, Inter-Turn Modality Switching (ITMS), evaluated as a controlled delivery-modality probe, does not consistently increase attack success. These results demonstrate the value of run-centric, fine-grained evaluation for characterizing multimodal safety behavior beyond a single binary success metric.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning
- info Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning
- info Large Language Models Systematically Favor Popular Options: Evidence and Mitigation Across MCQs
- info When Patients Cut In: Extending Clinical Conversational AI Safety to Interruptions