SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents
The paper tests whether LLM agents can run the whole extreme-weather warning chain, not just isolated weather tasks.
The authors introduce SIREN-Bench, with 600 question-answer cases across 19 tasks and an end-to-end warning workflow. They report that existing weather-agent frameworks show major gaps on that benchmark. Their SIREN framework adds historical-case retrieval, skill distillation, predictive modeling, and access to mixed weather evidence and tools. In experiments, it beats weather-agent baselines on both single warning procedures and full warning chains. ArXiv · AI/CL/LG's note
The authors introduce SIREN-Bench, with 600 question-answer cases across 19 tasks and an end-to-end warning workflow. They report that existing weather-agent frameworks show major gaps on that benchmark. Their SIREN framework adds historical-case retrieval, skill distillation, predictive modeling, and access to mixed weather evidence and tools. In experiments, it beats weather-agent baselines on both single warning procedures and full warning chains. ArXiv · AI/CL/LG's note
score 4