AutoSynthData: Generating Training Data for Enterprise Agents
ServiceNow CoreAI has developed AutoSynthData, a system designed to bridge the gap between a model's general capabilities and the specific requirements of enterprise environments. By utilizing a target model's failures and a stronger teacher model's successes, AutoSynthData automatically generates and validates synthetic training tasks to improve agent performance in stateful enterprise settings.
The Framework for Useful Agentic Tasks
To generate effective training data, AutoSynthData defines a task as a triplet consisting of a system specification, a user prompt, and a verifier. Each component must meet specific criteria to ensure the training signal is high-quality:
- System Specification: Defines constraints, instructions, and environment policies. It must be compatible with the environment's tools and state.
- User Prompt: Specifies the goal and constraints. To be useful for training, prompts must be feasible (solvable in the environment), realistic (resembling actual user requests), and difficult (exposing a current weakness of the agent).
- Verifier: Determines success. A valid verifier must be consistent with the prompt and specification, sound (rejecting failures), and complete (accepting all valid solutions).
AutoSynthData Pipeline Overview
AutoSynthData operates as a closed-loop system that identifies capability gaps and fills them with targeted synthetic data. The process follows these primary stages:
- Gap Identification: The target model is evaluated using diagnostic tasks. Patterns of failure are analyzed alongside a stronger teacher model's successful trajectories.
- Capability Specification: Findings are distilled into "capability specification cards" that describe the required workflow, tools, and success criteria without providing the original prompts or trajectories.
- Task Generation: The generator uses these cards to create new, varied tasks. This occurs in two phases:
- Target Phase: Creates a core set of vetted examples based on the capability specifications.
- Multiply Phase: Expands the dataset by creating novel variants of accepted target samples to increase scale while limiting drift.
- Validation and Repair: Every candidate task undergoes a rigorous quality-control loop, including solver evaluation, positive verification (confirming the solution works), and negative verification (confirming the verifier rejects incorrect outcomes).
- Critique and Repair: Failed candidates are sent to a critic for diagnosis and targeted repair rather than being discarded immediately.
Dataset Quality and Batch Management
High-quality synthetic data requires both sample-level and batch-level verification to avoid redundancy and ensure coverage.
Sample-Level Verification
Candidates are prioritized based on difficulty: the system favors tasks that the target model solves in no more than one of three trials, but the stronger solver solves in at least two of three.
Batch-Level Review
A meta-review process examines accepted and rejected samples to identify overrepresented task families or missing capability dimensions. The controller then adjusts the generation strategy to balance diversity and reduce redundancy within the generation budget.
Experimental Results in EnterpriseOps Gym
ServiceNow tested AutoSynthData using the EnterpriseOps Gym benchmark across two different domains, using Gemma-4-26B-A4B-it as the target model.
Hybrid Domain
Using Qwen3.8-27B as the teacher, AutoSynthData generated 2,000 synthetic samples in 18 hours. Fine-tuning Gemma on this data resulted in a 7.2 percentage point increase in mean Pass@1 (a 35% relative improvement) and increased verifier success from 63.01% to 68.55%. This closed 59% of the original Pass@1 gap between the target and reference models.
ITSM Domain
Using DeepSeek-V4.1-Flash as the teacher, the system generated 1,994 synthetic samples in 66 hours. This resulted in a mean Pass@1 increase from 18.77% to 27.18%.
Moving the Training Frontier
AutoSynthData treats synthetic data generation as a search for tasks at the target model's capability boundary. Because the useful training distribution shifts as the model improves, the pipeline is designed to be iterative. After post-training, the model is re-evaluated, and the remaining failures guide the next round of generation. While these experiments focused on Supervised Fine-Tuning (SFT), ServiceNow notes that this mechanism could similarly support reinforcement learning (RL) by generating tasks that challenge the current policy.