Course
AI Evals for Engineers and PMs
Intermediate
Use this when you want Shreya Shankar's material for evals and related AI skills.
AI educator
AI Evals for Engineers and PMs
Useful if you need to judge whether an AI feature is actually improving.
Start with: Review the course outcomes and pair it with a real feature you can evaluate.
Course
Intermediate
Use this when you want Shreya Shankar's material for evals and related AI skills.
Engineers, PMs, AI product teams should start here when they need evals, llm reliability, and product quality. The strongest fit is a learner who wants material in these formats: course, essays.
Review the course outcomes and pair it with a real feature you can evaluate. After that, open one related resource below and write down the exact workflow, concept, or implementation pattern you want to apply.
Useful if you need to judge whether an AI feature is actually improving. Use this profile when you are comparing educators by topic, level, format, and practical usefulness rather than browsing random AI content.
Compare the skill coverage, the starting recommendation, the educator's own resources, and any videos when available. If you need evals, search the directory for that skill and shortlist three profiles before committing to a course, book, or playlist.
| Resource | Kind | Level | Use when |
|---|---|---|---|
|
AI SDK v6 Crash Course
Matt Pocock
|
Workshop | Intermediate | You want a structured AI SDK v6 course that covers model choice, text and object generation, UI streams, agents, persistence, context engineering, evals, and advanced app patterns. |
|
The AI Engineer Roadmap
Matt Pocock
|
Free tutorial | Beginner to intermediate | You want a guided path through core AI concepts, model selection, the AI engineering mindset, evals, and techniques for improving LLM-powered apps. |
|
LLM Evals
Hamel Husain
|
Guide | Intermediate | Your AI app needs quality checks before users see it. |
|
Evaluating AI Agents
DeepLearning.AI
|
Short course | Intermediate | You need to test, trace, and improve agent workflows instead of judging only single LLM responses. |
|
Building and Evaluating Advanced RAG Applications
DeepLearning.AI
|
Short course | Intermediate | You already know basic RAG and need better retrieval, evaluation, and production-quality patterns. |
|
AI Product Management Specialization
Duke University
|
Specialization | Beginner to intermediate | You want a structured product-management route for scoping, evaluating, and shipping AI products. |
|
Promptfoo
OpenAI / Promptfoo
|
Open-source eval and red-team framework | Intermediate to advanced | Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required. |
|
DeepEval
Confident AI
|
Pytest-style LLM evaluation framework | Intermediate to advanced | Engineering teams that want LLM and agent evaluations to behave like software unit tests, with thresholds, assertions, and CI-friendly failures. |
|
Ragas
Vibrant Labs
|
RAG and AI-application evaluation library | Intermediate to advanced | Evaluating retrieval quality, grounded generation, agent tool use, and production-aligned test data for RAG applications. |
|
Inspect AI
UK AI Security Institute
|
Model and agent evaluation harness | Intermediate to advanced | Rigorous, reproducible model and agent capability or safety evaluations involving tools, multi-turn interaction, coding, or sandboxed environments. |
|
Arize Phoenix
Arize AI
|
AI observability and evaluation platform | Intermediate to advanced | Teams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting. |
|
OpenEvals
LangChain
|
Reusable evaluator library | Intermediate to advanced | Python or TypeScript developers who want composable evaluator functions without adopting a full evaluation platform. |