The pipeline learns from what an agent gets wrong
ServiceNow’s AutoSynthData starts with a target model’s failures and a stronger teacher’s successful examples. It uses those gaps to create new tasks that exercise the missing capability, then validates whether each task is possible and whether the agent completed it. The authors say the system generated 2,000 validated tasks in roughly 18 hours.
A useful task includes an environment specification, a user request and a verifier. The verifier matters: synthetic prompts alone can be nonsensical, impossible or too easy, so a training set needs checks that distinguish success from a plausible-sounding answer.
Reported gains are specific to the benchmark
On EnterpriseOps Gym Hybrid, the team reports a 7.2-point Pass@1 improvement, which it describes as a 35% relative gain. It also reports an IT service management result rising from 18.77 to 27.18. These are results from the authors’ benchmark experiments, not evidence that the same gain transfers to every company or agent.
The method repeats the loop as the model improves: new failures point to the next training curriculum. That is a practical way to target data generation instead of producing a large, undirected pile of synthetic examples.
Why it matters for enterprise agents
Business workflows are shaped by the software, rules and data state in a particular organization. An agent can be broadly capable yet still fail on a local tool sequence or policy constraint. AutoSynthData is designed to generate targeted practice for those environment-specific gaps.
The open-source write-up explains the task design and validation process, giving builders a basis to inspect or adapt the approach. The key measure is not how many prompts it can generate, but whether the tasks expose real weaknesses and produce better performance on independent, relevant tests.
Source published October 2, 2026. Coverage is based on the maker’s announcement and demonstration.
