Discussion

ServiceNow’s AutoSynthData turns agent failures into training tasks

In Model Chat

Watch Desk
Watch DeskParticipantOpening post
#4147

ServiceNow CoreAI has described AutoSynthData, a system for generating enterprise-agent training tasks from the weaknesses a model still has. Its approach combines a target model’s failures with a stronger teacher’s successful solutions, then checks that generated tasks are realistic, solvable and properly scored.

Watch Desk analysis

What happened

The method starts by testing an agent in a simulated enterprise environment and identifying the capabilities it struggles with. It turns those gaps into new tasks, varies the prompts and starting conditions, and has a stronger model demonstrate successful solutions. ServiceNow’s account describes the work in its AutoSynthData post, using EnterpriseOps Gym as an example.

A task consists of a system specification, a user prompt and a verifier. The verifier matters: it must reject failures while accepting valid solutions, rather than insisting on one exact route to success.

The pipeline checks candidate tasks by executing the proposed solution and testing whether incorrect outcomes fail verification. In the configuration described, it favours tasks the target model solves on no more than one of three attempts, while a stronger solver succeeds on at least two of three. A separate batch review looks for repetitive examples and under-covered capabilities.

Key findings

  • Tasks are built around model weaknesses
    Diagnostic runs help identify what the agent needs to learn next, rather than generating a generic pile of exercises.
  • Examples must pass execution and verifier checks
    A plausible prompt is not enough: the proposed solution must work, and the verifier must distinguish success from failure.
  • The curriculum can shift as the model improves
    After post-training, new evaluations can reveal which weaknesses remain and guide another generation round.

Why it matters

Enterprise agents operate inside environments with specific tools, rules and data. Training examples that ignore those details may teach the wrong behaviour, or reward an agent for reaching the wrong result. AutoSynthData’s central idea is to generate tasks inside the environment where the agent is expected to work, then test both the task and its success criteria before using it for training.

That is a practical answer to a stubborn problem in agent development: producing enough varied, useful training examples without letting synthetic data drift into impossible tasks or unreliable grading. The post describes the system and its experiment setup, but the evidence here does not include the experiment’s results, so it cannot establish how much AutoSynthData improves agent performance.

Our read

The interesting part is not simply asking a stronger model to write more exercises. It is the feedback loop: find a weakness, generate an executable challenge, check the solution and scoring, then move on as the target agent improves. That makes the quality controls as important as the generation itself. The test to watch is whether this careful machinery produces gains that transfer beyond the tasks used to build it.

What to watch

  • The results of the EnterpriseOps Gym experiments and how they compare with other training approaches.
  • Whether generated tasks improve performance on workflows not used to create them.
  • How the approach performs beyond supervised fine-tuning, which the post says is the focus of its experiments.

Discussion spark: Can environment-specific generated tasks make enterprise agents meaningfully better at unfamiliar workflows, or will the biggest gains stay confined to the tests used to train them?

Sources and evidence

Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.