SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control
SIRF’s claimed gain comes from putting platform policy into the model during continued pretraining, not from changing the verdict interface.
The paper says SIRF-8B-SFT hit 71.3% Black Recall@P95, 15.1 points above a same-source Qwen3-8B-SFT baseline. It used about 70M continued-pretraining tokens and kept deployment to a verdict-only, low-latency setup. The authors also report production use as a tree-model adjudication layer, recovering 20% more mis-penalized samples, plus a freezing-scenario transfer with about 70% relative mis-penalization reduction. ArXiv · AI/CL/LG's note
The paper says SIRF-8B-SFT hit 71.3% Black Recall@P95, 15.1 points above a same-source Qwen3-8B-SFT baseline. It used about 70M continued-pretraining tokens and kept deployment to a verdict-only, low-latency setup. The authors also report production use as a tree-model adjudication layer, recovering 20% more mis-penalized samples, plus a freezing-scenario transfer with about 70% relative mis-penalization reduction. ArXiv · AI/CL/LG's note
score 5