All AI videos

AI video

How To Test AI Agents With Simulations

Raindrop co-founder Ben Hylak explains how teams can test agent changes against simulated environments built from production traces before users encounter a regression.

Hamel Husain · 2026 featured video

Watch on YouTube

What this video shows

Hylak starts from a practical failure mode: a prompt or tool change can alter later decisions even when the initial input stays fixed. His proposed test runs the changed agent against a synthetic copy of the surrounding services, then compares tool calls, errors, step count, latency, cost, and final output with the earlier trajectory.

The method adds a useful pre-release check, but it depends on the quality of the simulated world and the selected traces. OpenAI describes a related deployment-simulation method for model behavior, while Raindrop sells the product shown here. Learnetto recommends keeping deterministic checks, adversarial cases, human review, and a bounded production rollout alongside simulation.

OpenAI documents a related deployment simulation method and explains why historical traffic may not represent future use. Read Raindrop's own simulation product description for the synthetic-service model shown in the talk.

What you will learn

  • Replay complete trajectories because a changed tool result can alter every later decision.
  • Compare tool errors, actions, cost, latency, and final outcomes instead of scoring only the final text.
  • Choose a small set of representative and high-risk traces before generating thousands of weak cases.
  • Treat production behavior as the final evidence and use simulation to reduce risk before a controlled release.

How to apply this safely

  1. Capture consented traces for one bounded workflow and redact data that the test does not need.
  2. List every service and side effect the agent can reach, then create isolated substitutes with realistic state transitions.
  3. Run the old and changed agent against the same scenarios and review material trajectory differences.
  4. Release to a small cohort with monitoring, rollback criteria, and a check for failures absent from the simulation.

Important limitations

  • Hylak co-founded Raindrop and demonstrates its product. The talk does not compare simulation fidelity or defect detection with independent tools.
  • Historical traces omit new workflows and rare attacks, while a synthetic service can behave differently from production under concurrency, permissions, or partial failure.

Sources to check

Continue learning on Learnetto

AI evals guide

Combine simulations with datasets, graders, traces, and release gates.