Artificial Intelligence · 01.09.2026, 13:47 UTC
PathBridger: Subgoal Bridges for Offline Goal-Conditioned Reinforcement Learning
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.LG ↗ |
| Veröffentlicht | 01.09.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.29061v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (GCRL) aims to learn policies for reaching diverse goals entirely from fixed trajectory data. Long-horizon offline GCRL remains challenging because sparse goal-reaching signals must be propagated over many steps, while execution errors cannot be corrected through additional environment interaction. Existing methods address these challenges by improving long-range value estimation or reducing the effective decision horizon through subgoals, options, and action chunks. In several hierarchical methods, however, a selected subgoal specifies where to go, while the intervening state-space path remains implicit in an endpoint-conditioned low-level policy. To address this interface, we propose PathBridger, a hierarchical offline GCRL method that explicitly connects subgoal selection to short-horizon execution. PathBridger constructs a state-space bridge toward the selected intermediate endpoint and decodes it into a short executable action chunk using an inverse dynamics model. Experiments across the evaluated OGBench tasks demonstrate strong aggregate performance, with particularly large gains on the multi-object Cube manipulation tasks. Code: https://github.com/SChoish/PathBridger
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Reciprocity Separates Gradient Flow from Rotation in Conservative Physical Learning
- info PAC: Progress-Augmented Advantage Curriculum for Multi-Task Reinforcement Learning of LLMs
- info When the Martingale Never Stops Firing: Anytime-Valid Gating on Real Forecast Streams
- info Generative multi-domain transfer learning for fault detection in data-scarce wind turbines