Megadose AI progress, ranked and analyzed.

Rare Event Estimation via Iterative Unalignment

· ArXiv · AI/CL/LG ·
The paper proposes estimating catastrophic low-probability agent behaviors by deliberately perturbing a model into a better rare-event sampler.

The method builds an importance-sampling proposal by changing the original model’s weights, then searches that weight space with gradients. Its objective tries to amplify the target event while regularizing the sampler so the probability estimate stays stable. The authors test it on roughly 120M- and 2.6B-parameter models across more than 300 rare events, including probabilities as low as `10^-9`. In the most verifiable settings, they report more than `800x` compute-weighted efficiency over naive Monte Carlo for events below `10^-7`. ArXiv · AI/CL/LG's note

score 5

Categories: Research