Artificial Intelligence · 27.08.2026, 07:17 UTC
JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.LG ↗ |
| Veröffentlicht | 27.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.24044v2 Announce Type: replace Abstract: Latent world models plan by predicting how candidate actions advance learned latent dynamics. In self-predictive models, however, the encoder and predictor are optimized jointly and can co-adapt to latent transitions that are easy to predict but weakly constrained by the physical evolution of the scene. We introduce the cross-predictive JEPA (JEPA-x), which grounds visual latent dynamics in privileged physical trajectories. JEPA-x treats visual observations and physical states as corresponding views of the same action-conditioned trajectory, advances both through a shared predictor, and matches predictions from either view to future representations in both modalities. This requires the action-conditioned predictor to learn a common transition rule for both the visual and physical descriptions of the scene. The physical branch is used only during training, leaving no computational overhead at deployment. Empirical results show that JEPA-x reduces the rollout drift of a newly fitted predictor from $0.361$ to $0.104$ and increases mean control success from $53.6\%$ to $78.2\%$ on a multi-task suite spanning six evaluation subfamilies. We additionally show that making physical state decodable is not the load-bearing factor; rather, the gains arise from how cross-prediction shapes the geometry of the learned latent dynamics.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Natural Language Input, Semantic Track Representation, and LLM Inference: Making the Maritime Information Exchange Model Tractable
- info Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study
- info Beyond Local Surprise: Grounded Dialogue as Selective Belief Revision under Referential Uncertainty
- info VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following