Artificial Intelligence · 04.08.2026, 07:03 UTC
Priors Persist Through Suppression: A Stroop Paradigm for Lexical Override
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.CL ↗ |
| Veröffentlicht | 04.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2606.07555v4 Announce Type: replace Abstract: Glossaries, technical specifications, and system prompts routinely ask language models to use familiar words in unfamiliar ways. The instruction competes with what the word already means, and even when it wins, the pretrained prior keeps operating underneath. We test this with a Stroop-style paradigm: a prompt redefines a word (doctor now means forest), asks for a related word, and we score the new meaning against the word's pretrained associate (hospital) under matched neutral controls. Across 11 open-weight models from 1B to 9B parameters, the old meaning interferes in every model, remapping type, and prompt framing we test. After item-level controls, a model's ordinary preference for the old associate predicts the size of the interference, in the three document-relevant remapping types though not in antonyms. On antonym remapping, activation patching in five models locates the repair: restoring three prompt positions (where the word is redefined, where its new meaning appears, and where the question repeats it) recovers almost all of the effect (normalized recovery R in [0.92, 1.06]). The repair is asymmetric. The old meaning's logit falls under any perturbation of those positions, so pushing it down is not what separates a working override from a failing one; the new meaning survives only while the position carrying it is intact. What a local definition achieves is a protected new meaning, not a suppressed old one.