Megadose AI progress, ranked and analyzed.

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

· ArXiv · AI/CL/LG ·
The paper argues that influential training samples matter more when their answers are rewritten than when their weights are nudged.

The authors use influence functions to pick training examples, then replace those examples’ responses with behavior-aligned or behavior-opposed supervision while leaving prompts fixed. Across four open-weight LLMs, rewriting produced stronger and more durable shifts than reweighting on the same selected samples. The main testbed is epistemic abstention, with a similar qualitative contrast reported for safety refusal. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research