Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification
The paper claims model ancestry can be verified from weights alone, even after common checkpoint alterations.
The method removes a shared residual component that can make unrelated compatible checkpoints look connected. It then compares checkpoint-specific structure across residual blocks and calibrates a symmetric lineage score against independent models. In the reported benchmarks, it separates fine-tuned, LoRA-merged, pruned, and quantized descendants from independent and distilled models, with AUROC of 1.0. The authors also report robustness under function-preserving laundering and a LLaMA-2 case study identifying 3 related and 7 unrelated public checkpoints. HF Daily Papers' note
The method removes a shared residual component that can make unrelated compatible checkpoints look connected. It then compares checkpoint-specific structure across residual blocks and calibrates a symmetric lineage score against independent models. In the reported benchmarks, it separates fine-tuned, LoRA-merged, pruned, and quantized descendants from independent and distilled models, with AUROC of 1.0. The authors also report robustness under function-preserving laundering and a LLaMA-2 case study identifying 3 related and 7 unrelated public checkpoints. HF Daily Papers' note
score 4