All learning paths

AI learning path

Evals and reliability

Stop judging AI quality by vibes and start building repeatable checks.

Best for
AI product teams
Level
Intermediate
Time
8-16 hours

Choose this when

Your AI feature already has users, stakeholders, or enough risk that mistakes matter.

You should be able to

You can define task examples, expected behavior, graders, traces, regressions, and review workflows.

Checkpoint

Move on when quality discussions point to examples and metrics, not taste.

Do

Learning sequence

Work through the material inside each step. Videos are embedded where they fit; tutorials and references sit next to the task they support.

Search more

Step 1

Collect examples

Turn real user tasks, edge cases, and failures into a small eval set.

  • Examples
  • Edge cases
  • Labels

Watch here

LLM evaluation with W&B video thumbnail ►

LLM evaluation with W&B

Weights & Biases

Introduces evaluation workflows and measurement for LLM apps.

Open here

LLM Evals: Everything You Need to Know

Application-evaluation learning resource · Hamel Husain and Shreya Shankar · Intermediate to advanced

Practitioners building product-specific evals from traces, domain-expert judgments, and observed failures instead of relying on generic benchmarks.

Open resource

Step 2

Choose graders

Combine exact checks, human review, model grading, and trace inspection.

  • Graders
  • Traces
  • Review

Watch here

AI evals with Phoenix video thumbnail ►

AI evals with Phoenix

Arize AI

Use this when moving from examples to traces and debugging.

Open here

Step 3

Run regressions

Compare prompts, models, retrieval changes, and releases before users see them.

  • Baselines
  • Regression tests
  • Release gates

Watch here

Promptfoo red teaming video thumbnail ►

Promptfoo red teaming

Promptfoo

Regression testing and adversarial checks for prompt and model changes.

Open here

Promptfoo

Open-source eval and red-team framework · OpenAI / Promptfoo · Intermediate to advanced

Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required.

Open resource

DeepEval

Pytest-style LLM evaluation framework · Confident AI · Intermediate to advanced

Engineering teams that want LLM and agent evaluations to behave like software unit tests, with thresholds, assertions, and CI-friendly failures.

Open resource

Ragas

RAG and AI-application evaluation library · Vibrant Labs · Intermediate to advanced

Evaluating retrieval quality, grounded generation, agent tool use, and production-aligned test data for RAG applications.

Open resource

Inspect AI

Model and agent evaluation harness · UK AI Security Institute · Intermediate to advanced

Rigorous, reproducible model and agent capability or safety evaluations involving tools, multi-turn interaction, coding, or sandboxed environments.

Open resource

Arize Phoenix

AI observability and evaluation platform · Arize AI · Intermediate to advanced

Teams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting.

Open resource

Braintrust

AI evaluation and observability platform · Braintrust Data · Intermediate to advanced

Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.

Open resource

Practice task

Create a 20-row eval set for one AI workflow and run two prompt versions against it.

Reference

All resources in this path

Search resources

Step 1

LLM Evals: Everything You Need to Know

Application-evaluation learning resource · Hamel Husain and Shreya Shankar · Intermediate to advanced

Practitioners building product-specific evals from traces, domain-expert judgments, and observed failures instead of relying on generic benchmarks.

Step 2

Demystifying evals for AI agents

Agent-evaluation learning resource · Anthropic · Intermediate to advanced

Product and engineering teams designing practical evaluations for multi-turn, tool-using agents.

Step 3

Promptfoo

Open-source eval and red-team framework · OpenAI / Promptfoo · Intermediate to advanced

Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required.

Step 3

DeepEval

Pytest-style LLM evaluation framework · Confident AI · Intermediate to advanced

Engineering teams that want LLM and agent evaluations to behave like software unit tests, with thresholds, assertions, and CI-friendly failures.

Step 3

Ragas

RAG and AI-application evaluation library · Vibrant Labs · Intermediate to advanced

Evaluating retrieval quality, grounded generation, agent tool use, and production-aligned test data for RAG applications.

Step 3

Inspect AI

Model and agent evaluation harness · UK AI Security Institute · Intermediate to advanced

Rigorous, reproducible model and agent capability or safety evaluations involving tools, multi-turn interaction, coding, or sandboxed environments.

Step 3

Arize Phoenix

AI observability and evaluation platform · Arize AI · Intermediate to advanced

Teams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting.

Step 3

Braintrust

AI evaluation and observability platform · Braintrust Data · Intermediate to advanced

Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.

Educators to follow

Chip Huyen profile photo

Chip Huyen

Intermediate to advanced

Use the book page and related essays as a production engineering path.

View educator
Josh Pigford profile photo

Josh Pigford

Beginner to intermediate

Read the public notes and examples before deciding whether the paid material matches your business.

View educator