AI directory search

Search across educators, skills, and resources.

Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.

14 matches for "reliability"

Video matches

Watch first when you want a fast feel for the topic before opening courses, docs, or profiles.

AI Evals for Engineers & PMs video thumbnail

AI Evals for Engineers & PMs

Hamel Husain and Shreya Shankar · evals, product, llm reliability

Educators

Learning paths

Resources

Evaluating AI Agents

Short course · DeepLearning.AI · Intermediate

You need to test, trace, and improve agent workflows instead of judging only single LLM responses.

agent evals, evals, agents, reliability, tracing

How to Evaluate AI Agents

Evaluation guide · Hugging Face Community · Intermediate

You need a practical evaluation plan built around representative tasks, controlled environments, observable traces, outcome and constraint metrics, repeated trials, and production failures turned into regression tests.

hugging face, agents, evals, evaluation harnesses, trajectory metrics

OpenAI production best practices

Guide · OpenAI · Intermediate

You are moving from experiments to production and need the official OpenAI guidance on latency, retries, rate limits, safety, monitoring, and operational rollout.

openai, production, reliability, latency, cost

OpenAI eval design guide

Guide · OpenAI · Intermediate

You need practical guidance for designing representative eval datasets, choosing graders, and turning model testing into an engineering loop instead of ad hoc spot checks, especially while OpenAI's older Evals platform is winding down toward read-only status on October 31, 2026.

openai, evals, quality, datasets, regression testing

OpenAI Working with evals

Guide · OpenAI · Intermediate

You need API-level guidance for testing outputs, comparing models, and catching regressions during upgrades.

openai, evals, quality, regression testing, reliability

OpenAI Evaluate agent workflows

Guide · OpenAI · Intermediate

You need the current OpenAI path for tracing, grading, and regression-testing agent workflows instead of only single-prompt evals.

openai, agents, evals, traces, graders

Claude models overview video thumbnail

Claude models overview

Model docs · Anthropic · Beginner to advanced

You need the official comparison of the current Claude 5 family, context windows, aliases, and release families before choosing cost, speed, and reliability tradeoffs.

claude, anthropic, claude 5, model selection, frontier models

OpenRouter provider routing

Guide · OpenRouter · Intermediate

You need to learn how OpenRouter routes across providers, handles fallbacks, and exposes preference controls before relying on it in production.

openrouter, provider routing, fallbacks, model selection, reliability

AI Evals for Engineers & PMs video thumbnail

AI Evals for Engineers & PMs

Cohort course · Hamel Husain and Shreya Shankar · Intermediate

You are shipping AI features and need a serious evaluation workflow.

evals, product, llm reliability

AI Evals for Engineers and PMs

Course · Shreya Shankar · Intermediate

Use this when you want Shreya Shankar's material for evals and related AI skills.

Evals, LLM reliability, Product quality