Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe
Making open-source AI assistants better at chaining government services together
Open-source AI models struggle when they need to chain multiple steps across government APIs—calling one service, using its result to call another, and so on. Researchers created a benchmark of 145 real Korean government tasks to measure this gap, then built a technique that learns which tool combinations actually work by testing them live, generating training data that teaches smaller models to perform nearly as well as much larger ones.
As governments adopt open-source AI to protect citizen data, they need systems that can actually navigate their own services reliably. This work shows that a smaller, cheaper model can now handle complex multi-step government requests—making it practical for public agencies to deploy capable AI agents without buying expensive proprietary systems or hosting massive models.