AI directory search

Search across educators, skills, and resources.

Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.

54 matches for "Evals"

Video matches

Watch first when you want a fast feel for the topic before opening courses, docs, or profiles.

LLM evaluation with W&B video thumbnail ►

LLM evaluation with W&B

Weights & Biases · evals, llm apps, observability, mlops

AI evals with Phoenix video thumbnail ►

AI evals with Phoenix

Arize AI · evals, observability, tracing, rag debugging

Promptfoo red teaming video thumbnail ►

Promptfoo red teaming

Promptfoo · evals, prompt testing, red teaming, security

AI Evals for Engineers & PMs video thumbnail ►

AI Evals for Engineers & PMs

Hamel Husain and Shreya Shankar · evals, product, llm reliability

Educators

Hamel Husain profile photo

Hamel Husain

Hamel's AI evals guides · Intermediate to advanced

Very practical material on evaluating LLM apps before they disappoint users.

Skills

Evals, RAG, LLM product quality

Matt Pocock profile photo

Matt Pocock

AI Hero · Beginner to advanced

Practical developer-focused AI education across LLM fundamentals, AI SDK app development, MCP, Claude Code workflows, agent-ready codebases, evals, TDD, handoffs, and reusable skills such as /teach, /grill-me, /to-prd, /to-issues, /tdd, /triage, and /handoff.

Skills

AI coding, Claude Skills, Agentic workflows, AI SDK, MCP, LLM fundamentals, Personalized learning

Providers and platforms

Weights & Biases profile photo

Weights & Biases

W&B Courses · Intermediate

Good for builders who need to measure, debug, and improve LLM apps rather than just demo them.

Topics

LLM apps, Evals, Experiment tracking, MLOps

OpenAI profile photo

OpenAI

OpenAI docs, Academy, and Cookbook · Beginner to advanced

Official model and implementation material for learning GPT-6 Astra and cost-sensitive GPT-5.6 choices, Codex workflows, subagents, memories, agent evals, MCP and connector patterns, retrieval, background jobs, prompt engineering, production best practices, model optimization, structured outputs, and OpenAI's Academy learning path.

Topics

GPT-6 Astra, GPT models, Reasoning models, Model selection, Agents, Subagents, RAG, Structured outputs, MCP, Evals, Memories

Arize AI profile photo

Arize AI

Phoenix · Intermediate

Useful for debugging and evaluating LLM applications once you move beyond prototypes.

Topics

Observability, Evals, Tracing, RAG debugging

Langfuse profile photo

Langfuse

Langfuse Docs · Intermediate

Good operational material for tracing, scoring, and improving production LLM apps.

Topics

Observability, Prompt management, Evals, Tracing

Vellum profile photo

Vellum

Vellum Guides · Beginner to intermediate

Useful for product and ops teams that need practical LLM product concepts without getting lost in research.

Topics

Prompt management, Evals, Workflow design

Useful for teams building repeatable AI product processes around prompts, datasets, and evaluations.

Topics

Prompt management, Evals, LLM workflows

Promptfoo profile photo

Promptfoo

Promptfoo Docs · Intermediate

Very practical for regression testing prompts, model changes, and LLM outputs.

Topics

Prompt testing, Evals, Red teaming

Maven AI courses profile photo

Maven AI courses

Maven AI courses · Beginner to advanced

Useful discovery surface for live courses taught by practitioners across AI product, work, and engineering.

Topics

AI product, AI leadership, AI workflows, Evals

OpenRouter profile photo

OpenRouter

OpenRouter docs · Beginner to intermediate

Useful for learning model comparison, latest-family aliases, routing, fallback behavior, agent construction, project-specific evals, and API-compatible experimentation across proprietary and open model families.

Topics

Model routing, Model comparison, Auto Router, Agent SDK, Coding-agent harnesses, GPT models, Claude models, Gemini, Llama, Mistral, DeepSeek, Qwen, API examples, Evaluation

Learning paths

Resources

AI SDK v6 Crash Course

Workshop · Matt Pocock · Intermediate

You want a structured AI SDK v6 course that covers model choice, text and object generation, UI streams, agents, persistence, context engineering, evals, and advanced app patterns.

ai sdk, llm apps, agents, streaming, evals

The AI Engineer Roadmap

Free tutorial · Matt Pocock · Beginner to intermediate

You want a guided path through core AI concepts, model selection, the AI engineering mindset, evals, and techniques for improving LLM-powered apps.

ai engineering, model selection, evals, llm apps

LLM Evals

Guide · Hamel Husain · Intermediate

Your AI app needs quality checks before users see it.

evals, quality, llm apps

Evaluating AI Agents

Short course · DeepLearning.AI · Intermediate

You need to test, trace, and improve agent workflows instead of judging only single LLM responses.

agent evals, evals, agents, reliability, tracing

Building and Evaluating Advanced RAG Applications

Short course · DeepLearning.AI · Intermediate

You already know basic RAG and need better retrieval, evaluation, and production-quality patterns.

rag, evals, retrieval, llm apps, ai engineering

Promptfoo

Open-source eval and red-team framework · OpenAI / Promptfoo · Intermediate to advanced

Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required.

evals, llm evaluation, ai quality, open-source frameworks

DeepEval

Pytest-style LLM evaluation framework · Confident AI · Intermediate to advanced

Engineering teams that want LLM and agent evaluations to behave like software unit tests, with thresholds, assertions, and CI-friendly failures.

evals, llm evaluation, ai quality, open-source frameworks

Ragas

RAG and AI-application evaluation library · Vibrant Labs · Intermediate to advanced

Evaluating retrieval quality, grounded generation, agent tool use, and production-aligned test data for RAG applications.

evals, llm evaluation, ai quality, open-source frameworks

Inspect AI

Model and agent evaluation harness · UK AI Security Institute · Intermediate to advanced

Rigorous, reproducible model and agent capability or safety evaluations involving tools, multi-turn interaction, coding, or sandboxed environments.

evals, llm evaluation, ai quality, open-source frameworks

Arize Phoenix

AI observability and evaluation platform · Arize AI · Intermediate to advanced

Teams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting.

evals, llm evaluation, ai quality, evaluation platforms

OpenEvals

Reusable evaluator library · LangChain · Intermediate to advanced

Python or TypeScript developers who want composable evaluator functions without adopting a full evaluation platform.

evals, llm evaluation, ai quality, open-source frameworks

Giskard

Agent testing and red-team framework · Giskard AI · Intermediate to advanced

Behavioral tests and adversarial scans of multi-turn agents, chatbots, and RAG systems from a pytest-compatible workflow.

evals, llm evaluation, ai quality, open-source frameworks

Braintrust

AI evaluation and observability platform · Braintrust Data · Intermediate to advanced

Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.

evals, llm evaluation, ai quality, evaluation platforms

LangSmith

Agent evaluation and observability platform · LangChain · Intermediate to advanced

Agent teams, especially LangChain and LangGraph users, that need tracing, datasets, experiments, human review, and production feedback in one system.

evals, llm evaluation, ai quality, evaluation platforms

Langfuse

Open-source evals and observability platform · ClickHouse · Intermediate to advanced

Teams prioritizing open-source data control and one workflow across tracing, prompts, datasets, experiments, annotation, and evaluation.

evals, llm evaluation, ai quality, evaluation platforms

W&B Weave

LLM and agent evaluation platform · Weights & Biases · Intermediate to advanced

Teams already using Weights & Biases or wanting versioned datasets, scorers, traces, and evaluation comparisons in an experiment-oriented workflow.

evals, llm evaluation, ai quality, evaluation platforms

Opik

Open-source agent evaluation platform · Comet · Intermediate to advanced

Teams wanting hosted or self-hosted observability with failure-driven regression suites, experiments, metrics, and human review.

evals, llm evaluation, ai quality, evaluation platforms

MLflow GenAI evaluation

Open-source GenAI evaluation and monitoring · MLflow Project · Intermediate to advanced

Teams extending an existing MLflow or MLOps stack to LLM and agent tracing, evaluation-driven development, human feedback, and monitoring.

evals, llm evaluation, ai quality, evaluation platforms

Patronus AI

AI reliability and evaluation platform · Patronus AI · Intermediate to advanced

Safety- and reliability-sensitive RAG or agent systems needing specialized hallucination, retrieval, PII, bias, policy, and custom-criteria evaluators.

evals, llm evaluation, ai quality, evaluation platforms

OpenAI Evals

Hosted LLM and agent evaluation API · OpenAI · Intermediate to advanced

Existing OpenAI API teams maintaining dataset-based prompt or model regression tests during the remaining service window.

evals, llm evaluation, ai quality, cloud and provider evals

Vertex AI Gen AI evaluation service

Managed model and agent evaluation service · Google Cloud · Intermediate to advanced

Vertex AI teams evaluating model outputs, prompts, RAG responses, function calls, and agent trajectories with managed or custom metrics.

evals, llm evaluation, ai quality, cloud and provider evals

Microsoft Foundry evaluations

Cloud evaluation portal and SDK · Microsoft · Intermediate to advanced

Microsoft Foundry teams evaluating models, RAG systems, and single- or multi-turn agents for quality, safety, task completion, and tool use.

evals, llm evaluation, ai quality, cloud and provider evals

Amazon Bedrock Evaluations

Managed model and RAG evaluation jobs · Amazon Web Services · Intermediate to advanced

AWS teams comparing Bedrock models or assessing knowledge bases and external RAG sources with automated metrics, LLM judges, or human reviewers.

evals, llm evaluation, ai quality, cloud and provider evals

Anthropic evaluation guide

Evaluation methodology guide · Anthropic · Intermediate to advanced

Teams designing task-specific evaluation suites for Claude applications that need criteria, datasets, edge cases, grading patterns, and code examples.

evals, llm evaluation, ai quality, cloud and provider evals

Harbor

Agent evaluation and optimization harness · Harbor Framework Team · Intermediate to advanced

Running coding and computer-use agents against reproducible, sandboxed task suites at local or cloud scale.

evals, llm evaluation, ai quality, benchmarks and learning resources

Language Model Evaluation Harness

Language-model benchmark runner · EleutherAI · Intermediate to advanced

Reproducible few-shot and zero-shot evaluation of base or instruction-tuned language models on established academic benchmarks.

evals, llm evaluation, ai quality, benchmarks and learning resources

Lighteval

Multi-backend LLM evaluation toolkit · Hugging Face · Intermediate to advanced

Teams evaluating local, distributed, or API-served models across a large multilingual task catalog while retaining sample-level results.

evals, llm evaluation, ai quality, benchmarks and learning resources

SWE-bench

Software-engineering agent benchmark · SWE-bench team · Intermediate to advanced

Measuring whether coding agents can resolve real GitHub issues by producing repository patches that pass executable tests.

evals, llm evaluation, ai quality, benchmarks and learning resources

Terminal-Bench

Terminal-agent benchmark · Harbor Framework and Laude Institute · Intermediate to advanced

Comparing agents on difficult, verifiable coding, systems, security, and scientific work inside sandboxed terminals.

evals, llm evaluation, ai quality, benchmarks and learning resources

Demystifying evals for AI agents

Agent-evaluation learning resource · Anthropic · Intermediate to advanced

Product and engineering teams designing practical evaluations for multi-turn, tool-using agents.

evals, llm evaluation, ai quality, benchmarks and learning resources

LLM Evals: Everything You Need to Know

Application-evaluation learning resource · Hamel Husain and Shreya Shankar · Intermediate to advanced

Practitioners building product-specific evals from traces, domain-expert judgments, and observed failures instead of relying on generic benchmarks.

evals, llm evaluation, ai quality, benchmarks and learning resources

Claude and the ART enzyme system explained

AI-assisted science explainer · Anthropic · Intermediate

Anthropic reported ART on September 23, 2026. Use this explainer to separate what its agents found, what the lab confirmed, and what remains a hypothesis.

claude, ai agents, ai for science, biology, genome mining

Grok 4.7 explained

Frontier coding model launch explainer · SpaceXAI · Intermediate to advanced

SpaceXAI released Grok 4.7 on September 21, 2026 for coding, agentic tasks, and knowledge work. Use this guide to compare its task reliability and total cost with your current model.

grok 4.7, coding agents, long-running agents, model selection, evals

Union Alpha revealed as Pareto 26.9

Multi-model routing launch explainer · Unbiased · Intermediate to advanced

Union Alpha was revealed as Unbiased's Pareto 26.9 on September 17, 2026: a hosted system that routes work across several models and returns one checked answer.

pareto 26.9, union alpha, model routing, ensembles, coding agents

The Anatomy of Harness Engineering

Coding agent evaluation guide · Google Developers Blog · Intermediate to advanced

You want Google's September 9, 2026 guide to evaluating coding agents with small behavioral checks, outcome-based assertions, and batch runs that catch regressions without treating a single benchmark score as the whole story.

coding agents, evals, behavioral evaluations, regression testing, harness engineering

How Meta built safety into Muse

Agent security architecture guide · Meta AI · Intermediate to advanced

You want Meta's September 8, 2026 technical account of defense-in-depth for a long-running personal agent, including isolated runtime cells, credential surrogates, a separate permission authority, tainted-egress tracking, scoped approvals, browser controls, red teaming, and prompt-injection evals.

meta, muse, agent security, prompt injection, least privilege