Artificial Intelligence · 25.08.2026, 12:31 UTC
All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.CL ↗ |
| Veröffentlicht | 25.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.22425v1 Announce Type: cross Abstract: Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified profile, in an acute crisis: a caregiver learning of a relative's dementia diagnosis. Auditors blind to the profile prompt recovered the specified bands from dialogue alone with high agreement on every instrument (ICC(2,4) = 0.91; 0.79-0.96 by instrument; band-score r = 0.78), as expected for the Big Five but equally for coping style, coping self-efficacy, resilience and reactance, which the lexical approach never covered. Such evaluation therefore reaches beyond the Five Factor Model to motivational, regulatory and self-appraisal dispositions. The four models were not distinguishable on emotion stabilisation and failed alike, sharing three modes: verbosity, a talk-to-listen ratio above one, and problem-solving before the situation had been explored.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization
- info Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- info Bayes with No Shame: Admissibility Geometries of Predictive Inference
- info MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs