Artificial Intelligence · 01.09.2026, 17:17 UTC
An Identifiability Theory of Masked Prediction: Mode Blindness and Mask Schedules
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.LG ↗ |
| Veröffentlicht | 01.09.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.01383v3 Announce Type: replace Abstract: Masked prediction learns by inferring missing variables from visible context. This raises a fundamental question: when does near-optimal conditional prediction determine the joint data law? We study this via an $\varepsilon$-identifiability modulus measuring the largest joint-law error compatible with masked-prediction excess risk at most $\varepsilon$. For slow-mixing data laws with separated global modes, we show that a model can assign substantially incorrect probabilities to entire data regimes while incurring exponentially small excess risk. An exact information decomposition reveals why: for a fixed mask, the prediction loss detects only the portion of the mode-weight mismatch that the visible context leaves unresolved. For small mode-weight perturbations, this sensitivity is proportional to residual mode uncertainty. Once averaged over masks, this residual uncertainty governs the objective's sensitivity to global mode frequencies, with low-visibility masks restoring mode-weight sensitivity and positive full-mask mass providing universal joint-law control under the joint conditional objective. We provide computational and empirical evidence for these predictions through exact calculations, controlled optimization experiments, and measurements on natural text. More broadly, our study suggests that a predictive objective can identify global distinctions only insofar as its conditioning structure leaves them unresolved.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Forecasting Weather-Driven Price Dynamics Across Sri Lankan Tea Market Catalogues
- info SatDL: Jointly Optimizing Data Redistribution and Training for Satellite-Based Distributed Learning
- info GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented Reasoning
- info A Comprehensive Review of Large Language Models for Nanophotonics: From Surrogate Modeling to Autonomous Design