Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective
The paper’s detector adjusts loss by predictive entropy to separate memorized training data from merely predictable text.
The authors argue that likelihood alone can falsely flag non-members when the model can predict them easily. Their Energy Transfer Detection method uses an inclined boundary in loss-entropy space, framed as a Helmholtz free-energy score. In experiments, ETD reports the best average detection results, with AUROC gains up to 3.5% and TPR@5%FPR gains up to 5.1%. Source: ArXiv · AI/CL/LG's note.
The authors argue that likelihood alone can falsely flag non-members when the model can predict them easily. Their Energy Transfer Detection method uses an inclined boundary in loss-entropy space, framed as a Helmholtz free-energy score. In experiments, ETD reports the best average detection results, with AUROC gains up to 3.5% and TPR@5%FPR gains up to 5.1%. Source: ArXiv · AI/CL/LG's note.
score 5