Megadose AI progress, ranked and analyzed.

Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis

· ArXiv · AI/CL/LG ·
A hybrid pair of small open-weight models beat the best ungrounded frontier baseline on Meta’s malware-analysis benchmark.

The paper tests small-language-model orchestration on structured questions about malware detonation reports. Its best SLM setup, Qwen3-4B with Foundation-Sec-8B, reached 35.30% accuracy, above the strongest cyber-specialized baseline at 22.54% and the strongest ungrounded frontier baseline at 34.77%. Grounded Gemini still led overall when given the same evidence pipeline, at 38.22%. The work is slated for RAID 2026. ArXiv · AI/CL/LG's note

score 4

Categories: Research