Megadose Built for builders and researchers.

Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution

· HF Daily Papers ·
Provenance recovered cleanly before rewriting, then failed as a guide to training value.

In financial-risk text, source attribution hit 98.7% on original generated passages, but dropped to 53.1% after paraphrasing and 29.0% after style rewriting. Generated-versus-human detection stayed near perfect against the tested human comparison set. A source-based selection rule and a reference-model scoring rule chose different generated examples, but the planned comparison did not find a stable difference in later model degradation. The paper’s point is narrower than provenance being useless: source identity and training fitness are separate measurements. HF Daily Papers' note

score 4

Categories: Research