Artificial Intelligence · 25.08.2026, 11:46 UTC
SAVER: Selective Auditing of Verbal Evidence for Error Recovery in VLM Change Reasoning
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.CL ↗ |
| Veröffentlicht | 25.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.22857v1 Announce Type: new Abstract: Vision-language models (VLMs) frequently fail at visual change reasoning, even when their vision encoders contain sufficient information. We observe that correct VLM outputs tend to contain explicit verbal evidence (object names, colors, spatial locations) that supports the claimed change, while incorrect outputs often lack such evidence. We propose SAVER (Selective Auditing of Verbal Evidence for Error Recovery), a lightweight, rule-based method that parses VLM responses for this evidence and triggers structured reprompting only when evidence is missing or inconsistent. Across three change detection benchmarks and four VLMs, SAVER significantly improves accuracy on tasks where errors stem from the model failing to articulate what it saw (expression failures), with gains up to +25.8% on CLEVR-Change. The evidence patterns can also be generated by an LLM in a single call, matching the hand-tuned gate on CLEVR-Change. Ablation experiments confirm that the evidence gate, not reprompting alone, drives the improvement.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info DynHD: Hallucination Detection for Diffusion Large Language Models via Denoising Dynamics Deviation Learning
- info Dialects of Translationese Shape Language Model Learning
- info Adaptive Test-Time Compute Allocation for Block Diffusion Language Models in Complex Reasoning
- info LakeHopper: Knowledge-Aware Adaptation of Column Type Annotators across Data Lakes