Artificial Intelligence · 31.08.2026, 05:47 UTC
How Far Should Tokenization Go? Predictive Effectiveness and Relational Losslessness
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.LG ↗ |
| Veröffentlicht | 31.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2608.18025v2 Announce Type: replace Abstract: GPT-style models have achieved remarkable success with finite vocabularies of reusable tokens, making the token interface a central component of modern sequence modeling. Symbolic music appears naturally compatible with this paradigm: it consists of discrete note events and recurring structures such as chords, motifs, and phrases. However, when tokenization moves beyond language, the interface must be specified for each domain. Existing work offers many effective designs, but no unified criterion for deciding what tokenization should represent and how far it should go. Using predictive codelength as a common criterion, we formulate the Effectiveness--Losslessness Framework to define where tokenization should begin and where it should end. The Fact--Token Boundary marks where observation-determined structure should enter the token interface, through operations such as coordinate construction. Within this interface, the resulting carrier may be reversibly recoded without changing the represented facts. The Token--State Boundary marks where tokenization should stop: relations that depend on context should remain for model-state computation rather than being fixed in advance by the tokenizer. We validate the framework through controlled multi-seed symbolic-music experiments, with an independent-corpus replication of the temporal intervention. Making musical time explicit consistently reduces predictive code and also improves pitch and duration prediction, while tonal-frame canonicalization and pitch factorization provide …
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info SimpCue: Cue-Based Prompting for Multilingual Text Simplification
- info Beyond Global Scalars: Synergizing Token-Level Statistics and Deep Semantics for Adversarial AIGC Text Detection
- info Predicting Turn-Taking Outcomes in Multi-Party Conversation: Interpretable Modelling of Speech and Gaze Dynamics with Interpersonal Closeness
- info QUORUM: QUality-Optimized Routing Using Multiple annotators