Cloud and provider evals
Microsoft Foundry evaluations
Microsoft Foundry evaluations run in the portal or SDK over datasets, traces, and simulated conversations. The evaluator catalog spans quality, similarity, RAG, risk and safety, agents, rubrics, and Azure OpenAI graders.
- Maintainer
- Microsoft
- Deployment
- Managed Microsoft Foundry service and SDK
- Status
- Current as of September 22, 2026
What it does
Microsoft Foundry evaluations run in the portal or SDK over datasets, traces, and simulated conversations. The evaluator catalog spans quality, similarity, RAG, risk and safety, agents, rubrics, and Azure OpenAI graders.
- Evaluation over datasets, traces, existing conversations, and simulated conversations
- Quality, similarity, RAG, safety, agent, rubric, and Azure OpenAI graders
- Turn- and conversation-level results, custom evaluators, inspection, and comparison
Where it fits
Microsoft Foundry teams evaluating models, RAG systems, and single- or multi-turn agents for quality, safety, task completion, and tool use.
Use it after defining the behavior that matters, collecting representative cases, and deciding how each case will be graded. A tool can execute and organize an eval, but it cannot decide which failures are costly for your users.
Limitations and cautions
- Azure AI Foundry was renamed Microsoft Foundry, while classic documentation still exists; match guidance to the portal generation.
- Several evaluator families remain in public preview without an SLA, and judge calls consume Azure OpenAI quota.
How to evaluate the evaluator
Before using any score as a release gate, create a small evaluator-validation set that humans have labeled. Compare false passes, false failures, disagreement by failure type, run-to-run variance, latency, and cost. Keep deterministic checks for verifiable facts, calibrate model judges against domain experts, and review raw traces when aggregate scores move.