Artificial Intelligence · 28.08.2026, 08:02 UTC
A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.LG ↗ |
| Veröffentlicht | 28.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.26506v1 Announce Type: new Abstract: Model merging enables combining multiple fine-tuned models without additional training, but its safety implications remain poorly understood. Prior work primarily attributes merging risks to unsafe constituent models, implicitly assuming that merging individually aligned models preserves safety. In contrast, we show that model merging reveals a previously overlooked jailbreak risk rooted in the pretrained foundation model, even when all constituent models are individually safety-aligned. Motivated by this observation, we study a new threat setting where an attacker constructs jailbreak prompts that generalize across merged models sharing the same pretrained backbone, without access to the exact merging coefficients or constituent checkpoints. To exploit this phenomenon, we propose \textbf{Basin-Aware Jailbreak (BAJ)}, which formulates jailbreak generation as a min--max optimization over the merging space to produce transferable adversarial suffixes across merged model families. Experiments across diverse backbones and merging settings show that BAJ achieves consistently high transfer success rates and remains effective under existing defenses.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Hierarchical Channel Stacking: A Structured Decision Framework for AI-Generated Image Detection
- info Dynamical phase selection controls compute scaling in looped transformers
- info Hadamard Flattening and Gaussian Pooling Sketch for Least Squares with Coordinate-wise Guarantee
- info Sharp Minimax Regret for Infinite-Memory Logistic Prediction