Megadose AI progress, ranked and analyzed.

The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts

· ArXiv · AI/CL/LG ·
The paper argues that “concept directions” may be tracking answer probabilities, not concepts themselves.

The authors define “answer basins” as all continuations that produce the same answer, with each basin weighted by its total model probability. They claim the linear structures seen in probing and steering arise from differences in that answer-level probability measure. In their experiments, concept-consistent effects and reversals depend on whether concept labels align with the model’s answer measure. ArXiv · AI/CL/LG's note

score 4

Categories: Research