Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization
On ModerationBench, foundation models nearly tripled Bluesky’s deployed moderation F1 score on random posts.
The paper compares instruction-based policy reasoning with example-based precedent matching for vision-language moderation. Its benchmark uses 4,000 manually annotated Bluesky posts drawn from the wild. Both guidance styles reached similar peak effectiveness, with models scoring 0.60 F1 versus 0.22 for Bluesky’s system on random posts. The authors frame the result as evidence that foundation models can support more reliable, adaptable policy operationalization at scale. ArXiv · AI/CL/LG's note
The paper compares instruction-based policy reasoning with example-based precedent matching for vision-language moderation. Its benchmark uses 4,000 manually annotated Bluesky posts drawn from the wild. Both guidance styles reached similar peak effectiveness, with models scoring 0.60 F1 versus 0.22 for Bluesky’s system on random posts. The authors frame the result as evidence that foundation models can support more reliable, adaptable policy operationalization at scale. ArXiv · AI/CL/LG's note
score 5