Artificial Intelligence · 28.08.2026, 06:32 UTC
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.AI ↗ |
| Veröffentlicht | 28.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.27439v1 Announce Type: cross Abstract: LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-based retrieval. However, such retrieval can reuse misleading experiences due to retrieval bias and unclear tool credit, and full trajectories add context overhead while reducing interpretability. We propose RedEvoAgent, a black-box red-teaming agent that distills cross-case attack trajectories into a concise, human-readable attack skill. The attack skill adaptively evolves through tool-effectiveness profiling and Deciding-Tool Attribution for skill updates, and a validation ratchet that retains only updates improving validation performance. Experiments on multiple benchmarks, target models, and target execution harnesses show that RedEvoAgent outperforms fixed and agentic baselines, improves tool efficiency, and transfers across attacker models and target execution harnesses.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info No Plan, Yet Human: A Reactive Robotics Model Predicts Human Planning Failures on a Clinical Task
- info MedFabric: Gold Evidence Hides the Difficulty of Word-Level Medical Fabrication Detection
- info MambaCSP: Hybrid-Attention State Space Models for Hardware-Efficient Channel State Prediction
- info MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation