Artificial Intelligence · 25.08.2026, 07:01 UTC
LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.AI ↗ |
| Veröffentlicht | 25.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.22295v1 Announce Type: cross Abstract: Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question difficulty, scenario, and underlying capability demands can vary substantially. Simple retrospective averages may confound model ability with item characteristics. In this paper, we study a model-based evaluation framework that combines multidimensional item response theory model with question contexts to predict LLM performance on unseen questions. The framework represents LLMs through latent capability profiles while using question content to inform item characteristics, allowing information to transfer beyond previously observed items. Empirically, we find that for within-scenario evaluation, incorporating question embeddings improves prediction relative to model-free baselines, and that multidimensional latent structure provides a richer description of capability variation than unidimensional alternatives. At the same time, our results reveal an important limitation that the generalizability does not necessarily translate into reliable prediction under cross-scenario shift. These findings suggest that context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation, while also highlighting cross-scenario generalization as a central open challenge.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Adversarial Entropy Inflation Against Gumbel-Based Inference Verification
- info The Emergence of Relevance Through Axiomatic Attention Patterns During LoRA Fine-Tuning
- info Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents
- info Mycelial Search: A Graph-Structured Metaheuristic for Continuous Optimisation