Artificial Intelligence · 25.08.2026, 04:02 UTC
Evaluating Multimodal Narrative Understanding of Popular Hollywood Films
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.AI ↗ |
| Veröffentlicht | 25.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.21430v1 Announce Type: new Abstract: Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning about film history and the evolution of narrative techniques. But the creation of stable benchmarks built around Hollywood films is complicated by copyright protections. In this work, we address these concerns directly, by building a new collection of Hollywood films defined by two criteria: box office popularity (where we publish the first large-scale, open collection of weekly box office earnings reported by Variety magazine from 1922-1979); and likely public domain status (by researching copyright registrations and renewals in the US Catalog of Copyright Entries). We build a new multimodal MCQ benchmark on top of this collection that focuses on narrative elements that directly evaluate the abilities of models to inform meaningful research on film narrative; we find that many vision-language models struggle on this task (with many performing at near-chance levels of accuracy), while audio-visual models (including those that use audio in captioning scenes) reach a maximum accuracy of 61.1%, well below human-level performance.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info A Survey Instrument to Assess Students' AI and Generative AI Knowledge
- info Model of Models: When Does Emitting a Specialist Beat Attending, Adapting, or Tuning?
- info A Social Media Analysis of Discourse on the Israel--Palestine Conflict on Telegram
- info Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems