Artificial Intelligence · 25.08.2026, 08:46 UTC
GIM: Evaluating models via tasks that integrate multiple cognitive domains
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.AI ↗ |
| Veröffentlicht | 25.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2605.18663v2 Announce Type: replace Abstract: As LLM benchmarks saturate, the evaluation community has pursued two strategies to increase difficulty: escalating knowledge demands (GPQA, HLE) or removing knowledge entirely in favor of abstract reasoning (ARC-AGI). The first conflates memorization with capability; the second divorces reasoning from the practical contexts in which it matters. We take a different approach. The Grounded Integration Measure (GIM) is a benchmark of 820 original problems (615 public, 205 private) where difficulty comes from integration; individual problems require coordinating multiple cognitive operations (constraint satisfaction, state tracking, epistemic vigilance, audience calibration) over broadly accessible knowledge, so that reasoning stays grounded in realistic tasks without being gated on specialized expertise. Each problem is an original expert-authored composition, majority with rubric-decomposed scoring. We calibrate a judge-aware continuous response 2-parameter logistic (2PL) IRT model across 53 test-configurations (unique model x thinking-level pairs) and five calibrated judges, using 203,800 epoch-averaged prompt-judge cells derived from >1M raw judge-scored observations, producing robust ability estimates that correctly order test-configurations even when raw accuracy is distorted by errors, missing data, or judge leniency differences. Using this framework, we present a comprehensive leaderboard spanning 22 models and 47 reporting test-configurations, and conduct what is to our knowledge the most extensive published study of …
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels
- info GRC-ProbNet: Uncertainty-aware Feature Extraction for Cardiovascular Disease Classification
- info DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation
- info Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting