Megadose AI progress, ranked and analyzed.

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

· ArXiv · AI/CL/LG ·
KOPA-Bench tests whether open-source agents can chain live Korean government API calls, and EDGE is the authors’ recipe for generating training data that actually executes.

The paper introduces 145 real-world tasks built around Korean open public APIs. Its EDGE method maps which tool outputs can feed later inputs, then keeps only links that work against live APIs. The authors say GRPO fine-tuning on that synthesized data lets a 9B model nearly match an untuned 27B model from the same family. They also report gains beyond KOPA-Bench on BFCL. ArXiv · AI/CL/LG's note

score 4

Categories: Research