Artificial Intelligence · 27.08.2026, 07:17 UTC
Towards a theory of inference-time alignment with unknown rewards
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.LG ↗ |
| Veröffentlicht | 27.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.15402v3 Announce Type: replace Abstract: Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation. Yet, alignment has remained poorly understood from a statistical learning perspective. We formulate inference-time alignment as a weak-to-strong learning problem, where a reference policy (weak model) is assumed to be fairly good and the goal is to produce a strong model that predicts a good response at test time with arbitrarily high probability. Our problem is formulated as learning from scratch --- everything is learned from data rather than assuming access to a good reward estimate, and thus differs from the existing inference-time alignment theory. Our framework shares similarity to the recent work of Joshi et al., (arXiv:2510.15464), where for each prompt, there could be multiple good responses. Our definition of the alignment learnability follows the standard PAC learning principle. We introduce a novel combinatorial dimension of the reward class which we call the alignment dimension, and show that it completely characterizes the alignment learnability --- a reward class is alignment learnable if and only if its alignment dimension is finite. The core of our learning procedure works by learning a pairwise comparator and then running a tournament over candidate responses. We believe that our results might shed light toward establishing a complete theoretical understanding of alignment.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Natural Language Input, Semantic Track Representation, and LLM Inference: Making the Maritime Information Exchange Model Tractable
- info Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence
- info MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation
- info Cross-Domain Transfer with Particle Physics Foundation Models: From Jets to Neutrino Interactions