Cloud and provider evals

Microsoft Foundry evaluations

Microsoft Foundry evaluations run in the portal or SDK over datasets, traces, and simulated conversations. The evaluator catalog spans quality, similarity, RAG, risk and safety, agents, rubrics, and Azure OpenAI graders.

Maintainer
Microsoft
Deployment
Managed Microsoft Foundry service and SDK
Status
Current as of September 22, 2026

What it does

Microsoft Foundry evaluations run in the portal or SDK over datasets, traces, and simulated conversations. The evaluator catalog spans quality, similarity, RAG, risk and safety, agents, rubrics, and Azure OpenAI graders.

  • Evaluation over datasets, traces, existing conversations, and simulated conversations
  • Quality, similarity, RAG, safety, agent, rubric, and Azure OpenAI graders
  • Turn- and conversation-level results, custom evaluators, inspection, and comparison

Where it fits

Microsoft Foundry teams evaluating models, RAG systems, and single- or multi-turn agents for quality, safety, task completion, and tool use.

Use it after defining the behavior that matters, collecting representative cases, and deciding how each case will be graded. A tool can execute and organize an eval, but it cannot decide which failures are costly for your users.

Limitations and cautions

  • Azure AI Foundry was renamed Microsoft Foundry, while classic documentation still exists; match guidance to the portal generation.
  • Several evaluator families remain in public preview without an SLA, and judge calls consume Azure OpenAI quota.

How to evaluate the evaluator

Before using any score as a release gate, create a small evaluator-validation set that humans have labeled. Compare false passes, false failures, disagreement by failure type, run-to-run variance, latency, and cost. Keep deterministic checks for verifiable facts, calibrate model judges against domain experts, and review raw traces when aggregate scores move.

Primary sources