Introducing Shieldstral.
Mistral released a 3B open-weights safety model that takes moderation policy as plain-language input at inference time.
Shieldstral treats moderation as a yes/no question-answering task across text, images, and text-plus-image inputs. The company says it matches or beats open guard models up to seven times larger on text safety and sets a new mark on multimodal moderation benchmarks. It returns calibrated yes/no probabilities from one forward pass, so teams can threshold or rank by confidence. The weights are released under Apache 2.0. Mistral AI's note
Shieldstral treats moderation as a yes/no question-answering task across text, images, and text-plus-image inputs. The company says it matches or beats open guard models up to seven times larger on text safety and sets a new mark on multimodal moderation benchmarks. It returns calibrated yes/no probabilities from one forward pass, so teams can threshold or rank by confidence. The weights are released under Apache 2.0. Mistral AI's note
score 6
Discussions
- hn · 261 points · 64 comments
- hn · 278 points · 69 comments
- hn · 286 points · 69 comments
- hn · 294 points · 70 comments
- hn · 297 points · 72 comments
- hn · 307 points · 74 comments
- hn · 319 points · 75 comments
- hn · 327 points · 79 comments
- hn · 329 points · 79 comments
- hn · 338 points · 82 comments
- hn · 344 points · 83 comments
- hn · 356 points · 85 comments
- hn · 361 points · 89 comments
- hn · 365 points · 90 comments
- hn · 377 points · 90 comments
- hn · 382 points · 94 comments
- hn · 390 points · 96 comments
- hn · 400 points · 98 comments
- hn · 409 points · 106 comments
- hn · 421 points · 109 comments
- hn · 426 points · 109 comments
- hn · 435 points · 113 comments
- hn · 441 points · 112 comments
- hn · 447 points · 112 comments
- hn · 449 points · 114 comments
- hn · 454 points · 116 comments
- hn · 456 points · 118 comments
- hn · 461 points · 121 comments
- hn · 462 points · 123 comments