Tool‑calling large language model (LLM) agents are increasingly deployed in enterprise applications, yet effective evaluation and optimization demand high‑quality, diverse task datasets that are often hard to obtain because of privacy and other constraints. Existing synthetic task generation approaches typically produce generic tasks, ignoring the agent’s internal state or database and failing to capture real‑world usage diversity.\
We introduce EdgeGen, a synthetic task generation framework. It extracts compliance rules from an agent’s specification and uses them to generate database‑grounded edge‑case tasks that intentionally violate those rules. When combined with conventional synthetic data, EdgeGen enables both finetuning and harness optimization of the agent, forming a fully automated closed‑loop pipeline that requires no human annotation.\
In experiments on the tau2bench airline domain, finetuning on EdgeGen‑generated data yields a consistent mean progress improvement ranging from 2% to 42%, whereas baseline methods cause degradation for some models. For harness optimization, EdgeGen achieves a mean progress gain of about 10% over human‑curated harnesses and roughly 30% over the base harness for the Gemma‑4‑e4b model.\
Review: EdgeGen’s rule‑driven edge‑case synthesis markedly enhances the robustness of tool‑calling agents on non‑happy‑path scenarios, offering a practical, fully automated solution for building safer, more reliable enterprise‑grade LLM systems.