Artificial Intelligence · 01.09.2026, 13:02 UTC
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.AI ↗ |
| Veröffentlicht | 01.09.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2607.19313v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{zero} learning signal. Providing privileged guidance during training, such as solution prefixes, can help overcome this learning cliff by steering the model towards {correct solutions with non-zero reward}. {We call these rollouts \textit{off-context}: they are generated from a training prompt that contains privileged guidance, while the target objective is defined by the original prompt without that guidance.} {We introduce} Off-Context GRPO (OC-GRPO), a minimally modified variant of GRPO that uses guided rollouts but applies an importance-corrected objective to steer the update back toward the original unguided objective, avoiding the mismatch that destabilizes uncorrected guided training. Empirically, our algorithm achieves a 3.8\% absolute improvement (13.7\% relative gain) over vanilla GRPO on average across standard mathematical reasoning benchmarks with negligible additional cost.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Designing for the Next Click: Bandits for Real-Time Page Layout
- info PathBridger: Subgoal Bridges for Offline Goal-Conditioned Reinforcement Learning
- info Mycelial Search: A Graph-Structured Metaheuristic for Continuous Optimisation
- info LITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space