Artificial Intelligence · 21.08.2026, 06:16 UTC
Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.AI ↗ |
| Veröffentlicht | 21.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2605.25848v2 Announce Type: replace-cross Abstract: A concept probe is only as reliable as the layer it is taken from. Probing at a fixed late layer, or at the peak of a separation curve, ignores a structural feature of how concepts form: the probe direction rotates substantially during assembly and does not settle until after the Concept Allocation Zone (CAZ) in which it forms. We introduce Geometric Evolution Maps (GEMs). A GEM records a concept's directional trajectory across one CAZ segment, takes the settled direction at that segment's final layer, and locates the handoff layer immediately beyond it - the first layer at which that direction can be evaluated outside the window it was estimated on. The handoff layer is derived from the segment boundary, not detected by a rotation criterion. A concept typically occupies several segments; we call that set its atlas. This rotation is substantial and consistent across architectures and concept types - a property of the representation, not an estimation artifact. The extracted direction is causal: ablating it suppresses far more separation than ablating a random direction, though no single site carries the effect alone. Probing after rotation completes is more precise than probing during it; a depth-matched control shows this advantage comes from probing at greater depth, not any privilege specific to the handoff boundary. A small number of structured exceptions are documented rather than left unaccounted for. A concept's atlas - its full set of GEMs, not any single one of them - is therefore the unit this method …