What this video shows
Laurie Voss instruments a financial-research agent, reads its traces, and identifies concrete failures before choosing metrics. The workshop starts with a deterministic ticker check, then shows why a correctness judge rejected all 13 reports when it lacked the live web sources that the agent had used. A faithfulness evaluator with the research context produces more informative results.
The workshop then builds a focused actionability rubric, compares model judgments with human annotations, and explains meta-evaluation, precision, recall, held-out cases, and judge bias. Failing traces become a dataset, prompt changes run against the same cases, and online evaluations watch new production traffic. Arize supplies the tooling in the demo, while the method also works with other tracing and evaluation stacks.
Use Arize's eval documentation for the judge types demonstrated in the workshop. Read the OpenTelemetry trace concepts before instrumenting a provider-neutral execution path.
What you will learn
- Read traces before writing an evaluator because a final answer can hide failed tools, missing files, wasted searches, or unsafe intermediate actions.
- Each evaluator should test one named property and receive the evidence needed for that property, including retrieved sources when it judges faithfulness.
- A model judge needs its own evaluation against reviewed examples because polished explanations can still accompany wrong labels.
- Compare agent changes on the same dataset and keep held-out cases so repeated prompt editing does not tune only for the visible examples.
How to apply this safely
- Collect a small set of representative traces and label the failure categories before choosing scores or graders.
- Write the cheapest deterministic checks first, then add one focused model judge where code cannot express the requirement.
- Have reviewers label a calibration set, measure judge disagreements, and revise the rubric with examples that clarify the boundary.
- Run the old and changed agent on the same fixed cases, inspect regressions, and add sampled online checks only after the offline comparison is useful.
Important limitations
- The workshop uses Arize AX and its evaluators, so product-specific setup occupies part of the session. The trace-first method and fixed-case experiments do not require that platform.
- The financial agent and demonstration labels teach the workflow rather than prove a universal judge accuracy. Your own production criteria need domain reviewers and representative data.
Sources to check
- Arize AX eval documentation Official definitions for deterministic, model, agent, and hosted evaluators used in the workshop.
- OpenTelemetry traces Provider-neutral concepts for recording the spans and context behind one request.
Continue learning on Learnetto
AI evals guide
Compare datasets, graders, tracing, and production monitoring methods.
Best AI agent evaluation courses
Build a longer learning path around agent tests and traces.
AI engineering courses
Connect evals with deployment, model selection, cost, and latency.