Artificial Intelligence · 31.08.2026, 06:17 UTC
Prompts Without Evidence: How Neuroimaging Mentions Shift Clinical Vision-Language Model Predictions
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.LG ↗ |
| Veröffentlicht | 31.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2603.28387v3 Announce Type: replace-cross Abstract: Trustworthy clinical AI must use real evidence and avoid relying on surface-level artifacts. We evaluate 12 open-weight vision-language models (VLMs) on two clinical neuroimaging cohorts for binary classification of affective disorders and cognitive decline. Both cohorts include structural magnetic resonance imaging (MRI) acquired under their original research protocols. Prior work does not establish the included neuroimaging inputs as reliable stand-alone diagnostic evidence for the present tasks. Nevertheless, when neuroimaging context is introduced, smaller VLMs gain up to 0.66 F1 under the evaluated augmented conditions, becoming competitive with models an order of magnitude larger. Confidence estimation shows that most of the calibration improvement for the analyzed smaller models occurs after the MRI reference is added to the prompt, before any image is supplied. Our preliminary expert case study finds that faithfulness remains low in every condition examined, with the reviewed model introducing unverified clinical details. Finally, in our single-model intervention, preference alignment suppresses MRI-referencing behavior but reduces the augmented-condition advantage, leaving the underlying issue unresolved. These results caution against reading surface metric gains as evidence of true multimodal integration, with direct implications for clinical VLM deployment.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding
- info Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models
- info One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography
- info MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize