What AstroPT knows about galaxies, and what that can teach us about LLMs
AstroPT is being used as a ground-truth testbed for interpretability claims that are harder to verify in language models.
The paper probes a transformer trained on millions of galaxy images and tracks what its frozen representations reveal across checkpoints, layers, model sizes, and objectives. It finds galaxy properties become decodable in an order that matches known difficulty: pixel-direct quantities appear earlier and shallower, while inferred properties emerge later and deeper. The sequence holds across tested objectives, with capacity changing strength but not order. The authors argue astronomy can calibrate mechanistic interpretability methods before they are applied to LLMs with less ground truth. HF Daily Papers' note
The paper probes a transformer trained on millions of galaxy images and tracks what its frozen representations reveal across checkpoints, layers, model sizes, and objectives. It finds galaxy properties become decodable in an order that matches known difficulty: pixel-direct quantities appear earlier and shallower, while inferred properties emerge later and deeper. The sequence holds across tested objectives, with capacity changing strength but not order. The authors argue astronomy can calibrate mechanistic interpretability methods before they are applied to LLMs with less ground truth. HF Daily Papers' note
score 5