Choose this when
Your AI feature already has users, stakeholders, or enough risk that mistakes matter.
AI learning path
Stop judging AI quality by vibes and start building repeatable checks.
Your AI feature already has users, stakeholders, or enough risk that mistakes matter.
You can define task examples, expected behavior, graders, traces, regressions, and review workflows.
Move on when quality discussions point to examples and metrics, not taste.
Do
Work through the material inside each step. Videos are embedded where they fit; tutorials and references sit next to the task they support.
Step 1
Turn real user tasks, edge cases, and failures into a small eval set.
Watch here
►
Introduces evaluation workflows and measurement for LLM apps.
Open here
Practitioners building product-specific evals from traces, domain-expert judgments, and observed failures instead of relying on generic benchmarks.
Open resourceStep 2
Combine exact checks, human review, model grading, and trace inspection.
Watch here
►
Use this when moving from examples to traces and debugging.
Open here
Product and engineering teams designing practical evaluations for multi-turn, tool-using agents.
Open resourceStep 3
Compare prompts, models, retrieval changes, and releases before users see them.
Watch here
►
Regression testing and adversarial checks for prompt and model changes.
Open here
Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required.
Open resourceEngineering teams that want LLM and agent evaluations to behave like software unit tests, with thresholds, assertions, and CI-friendly failures.
Open resourceEvaluating retrieval quality, grounded generation, agent tool use, and production-aligned test data for RAG applications.
Open resourceRigorous, reproducible model and agent capability or safety evaluations involving tools, multi-turn interaction, coding, or sandboxed environments.
Open resourceTeams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting.
Open resourceTeams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.
Open resourceCreate a 20-row eval set for one AI workflow and run two prompt versions against it.
Reference
Step 1
Practitioners building product-specific evals from traces, domain-expert judgments, and observed failures instead of relying on generic benchmarks.
Step 2
Product and engineering teams designing practical evaluations for multi-turn, tool-using agents.
Step 3
Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required.
Step 3
Engineering teams that want LLM and agent evaluations to behave like software unit tests, with thresholds, assertions, and CI-friendly failures.
Step 3
Evaluating retrieval quality, grounded generation, agent tool use, and production-aligned test data for RAG applications.
Step 3
Rigorous, reproducible model and agent capability or safety evaluations involving tools, multi-turn interaction, coding, or sandboxed environments.
Step 3
Teams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting.
Step 3
Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.
Intermediate to advanced
Read the evals guide and build a small test set for your own app.
View educator
Intermediate
Review the course outcomes and pair it with a real feature you can evaluate.
View educator
Intermediate to advanced
Use the book page and related essays as a production engineering path.
View educatorBeginner to intermediate
Read the public notes and examples before deciding whether the paid material matches your business.
View educator
Intermediate
Review the Maven syllabus and compare it to your current product workflow.
View educator
Beginner to intermediate
Browse the How I AI interviews and copy the workflows that match your role.
View educator