AI directory search

Search across educators, skills, and resources.

Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.

46 matches for "evaluation"

Video matches

Watch first when you want a fast feel for the topic before opening courses, docs, or profiles.

12GB Model, 8 Hours, One 3D Game: OrcaSAQ2 27B Tested video thumbnail ►

12GB Model, 8 Hours, One 3D Game: OrcaSAQ2 27B Tested

Fahd Mirza · 2026, local models, qwen3.8, hermes agent

LLM evaluation with W&B video thumbnail ►

LLM evaluation with W&B

Weights & Biases · evals, llm apps, observability, mlops

AI Evals for Engineers & PMs video thumbnail ►

AI Evals for Engineers & PMs

Hamel Husain and Shreya Shankar · evals, product, llm reliability

Educators

Ben Lorica profile photo

Ben Lorica

The Data Exchange · Intermediate

Good practitioner interviews across data, ML, and AI engineering.

Skills

Data systems, ML engineering, AI trends

Greg Kamradt profile photo

Greg Kamradt

Data Independent AI tutorials · Beginner to intermediate

Practical walkthroughs for retrieval, LLM application patterns, and common developer questions.

Skills

RAG, LLM apps, Prompting, Evaluation

Providers and platforms

Weights & Biases profile photo

Weights & Biases

W&B Courses · Intermediate

Good for builders who need to measure, debug, and improve LLM apps rather than just demo them.

Topics

LLM apps, Evals, Experiment tracking, MLOps

Useful for teams building repeatable AI product processes around prompts, datasets, and evaluations.

Topics

Prompt management, Evals, LLM workflows

A strong foundation for people who need the math and modeling basics under applied AI.

Topics

ML foundations, Supervised learning, Unsupervised learning, Model evaluation

OpenRouter profile photo

OpenRouter

OpenRouter docs · Beginner to intermediate

Useful for learning model comparison, latest-family aliases, routing, fallback behavior, agent construction, project-specific evals, and API-compatible experimentation across proprietary and open model families.

Topics

Model routing, Model comparison, Auto Router, Agent SDK, Coding-agent harnesses, GPT models, Claude models, Gemini, Llama, Mistral, DeepSeek, Qwen, API examples, Evaluation

Fireworks AI profile photo

Fireworks AI

Fireworks AI model catalog · Intermediate to advanced

Official material for comparing serverless and dedicated model serving, training specialized models, and measuring quality, token use, cost, and task duration on production-shaped evaluations.

Topics

Ember-1, Kimi K3, Serverless inference, Dedicated deployment, Model training, Model evaluation, Reasoning efficiency

Resources

Building and Evaluating Advanced RAG Applications

Short course · DeepLearning.AI · Intermediate

You already know basic RAG and need better retrieval, evaluation, and production-quality patterns.

rag, evals, retrieval, llm apps, ai engineering

Promptfoo

Open-source eval and red-team framework · OpenAI / Promptfoo · Intermediate to advanced

Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required.

evals, llm evaluation, ai quality, open-source frameworks

DeepEval

Pytest-style LLM evaluation framework · Confident AI · Intermediate to advanced

Engineering teams that want LLM and agent evaluations to behave like software unit tests, with thresholds, assertions, and CI-friendly failures.

evals, llm evaluation, ai quality, open-source frameworks

Ragas

RAG and AI-application evaluation library · Vibrant Labs · Intermediate to advanced

Evaluating retrieval quality, grounded generation, agent tool use, and production-aligned test data for RAG applications.

evals, llm evaluation, ai quality, open-source frameworks

Inspect AI

Model and agent evaluation harness · UK AI Security Institute · Intermediate to advanced

Rigorous, reproducible model and agent capability or safety evaluations involving tools, multi-turn interaction, coding, or sandboxed environments.

evals, llm evaluation, ai quality, open-source frameworks

Arize Phoenix

AI observability and evaluation platform · Arize AI · Intermediate to advanced

Teams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting.

evals, llm evaluation, ai quality, evaluation platforms

OpenEvals

Reusable evaluator library · LangChain · Intermediate to advanced

Python or TypeScript developers who want composable evaluator functions without adopting a full evaluation platform.

evals, llm evaluation, ai quality, open-source frameworks

Giskard

Agent testing and red-team framework · Giskard AI · Intermediate to advanced

Behavioral tests and adversarial scans of multi-turn agents, chatbots, and RAG systems from a pytest-compatible workflow.

evals, llm evaluation, ai quality, open-source frameworks

Braintrust

AI evaluation and observability platform · Braintrust Data · Intermediate to advanced

Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.

evals, llm evaluation, ai quality, evaluation platforms

LangSmith

Agent evaluation and observability platform · LangChain · Intermediate to advanced

Agent teams, especially LangChain and LangGraph users, that need tracing, datasets, experiments, human review, and production feedback in one system.

evals, llm evaluation, ai quality, evaluation platforms

Langfuse

Open-source evals and observability platform · ClickHouse · Intermediate to advanced

Teams prioritizing open-source data control and one workflow across tracing, prompts, datasets, experiments, annotation, and evaluation.

evals, llm evaluation, ai quality, evaluation platforms

W&B Weave

LLM and agent evaluation platform · Weights & Biases · Intermediate to advanced

Teams already using Weights & Biases or wanting versioned datasets, scorers, traces, and evaluation comparisons in an experiment-oriented workflow.

evals, llm evaluation, ai quality, evaluation platforms

Opik

Open-source agent evaluation platform · Comet · Intermediate to advanced

Teams wanting hosted or self-hosted observability with failure-driven regression suites, experiments, metrics, and human review.

evals, llm evaluation, ai quality, evaluation platforms

MLflow GenAI evaluation

Open-source GenAI evaluation and monitoring · MLflow Project · Intermediate to advanced

Teams extending an existing MLflow or MLOps stack to LLM and agent tracing, evaluation-driven development, human feedback, and monitoring.

evals, llm evaluation, ai quality, evaluation platforms

Patronus AI

AI reliability and evaluation platform · Patronus AI · Intermediate to advanced

Safety- and reliability-sensitive RAG or agent systems needing specialized hallucination, retrieval, PII, bias, policy, and custom-criteria evaluators.

evals, llm evaluation, ai quality, evaluation platforms

OpenAI Evals

Hosted LLM and agent evaluation API · OpenAI · Intermediate to advanced

Existing OpenAI API teams maintaining dataset-based prompt or model regression tests during the remaining service window.

evals, llm evaluation, ai quality, cloud and provider evals

Vertex AI Gen AI evaluation service

Managed model and agent evaluation service · Google Cloud · Intermediate to advanced

Vertex AI teams evaluating model outputs, prompts, RAG responses, function calls, and agent trajectories with managed or custom metrics.

evals, llm evaluation, ai quality, cloud and provider evals

Microsoft Foundry evaluations

Cloud evaluation portal and SDK · Microsoft · Intermediate to advanced

Microsoft Foundry teams evaluating models, RAG systems, and single- or multi-turn agents for quality, safety, task completion, and tool use.

evals, llm evaluation, ai quality, cloud and provider evals

Amazon Bedrock Evaluations

Managed model and RAG evaluation jobs · Amazon Web Services · Intermediate to advanced

AWS teams comparing Bedrock models or assessing knowledge bases and external RAG sources with automated metrics, LLM judges, or human reviewers.

evals, llm evaluation, ai quality, cloud and provider evals

Anthropic evaluation guide

Evaluation methodology guide · Anthropic · Intermediate to advanced

Teams designing task-specific evaluation suites for Claude applications that need criteria, datasets, edge cases, grading patterns, and code examples.

evals, llm evaluation, ai quality, cloud and provider evals

Harbor

Agent evaluation and optimization harness · Harbor Framework Team · Intermediate to advanced

Running coding and computer-use agents against reproducible, sandboxed task suites at local or cloud scale.

evals, llm evaluation, ai quality, benchmarks and learning resources

Language Model Evaluation Harness

Language-model benchmark runner · EleutherAI · Intermediate to advanced

Reproducible few-shot and zero-shot evaluation of base or instruction-tuned language models on established academic benchmarks.

evals, llm evaluation, ai quality, benchmarks and learning resources

Lighteval

Multi-backend LLM evaluation toolkit · Hugging Face · Intermediate to advanced

Teams evaluating local, distributed, or API-served models across a large multilingual task catalog while retaining sample-level results.

evals, llm evaluation, ai quality, benchmarks and learning resources

SWE-bench

Software-engineering agent benchmark · SWE-bench team · Intermediate to advanced

Measuring whether coding agents can resolve real GitHub issues by producing repository patches that pass executable tests.

evals, llm evaluation, ai quality, benchmarks and learning resources

Terminal-Bench

Terminal-agent benchmark · Harbor Framework and Laude Institute · Intermediate to advanced

Comparing agents on difficult, verifiable coding, systems, security, and scientific work inside sandboxed terminals.

evals, llm evaluation, ai quality, benchmarks and learning resources

Demystifying evals for AI agents

Agent-evaluation learning resource · Anthropic · Intermediate to advanced

Product and engineering teams designing practical evaluations for multi-turn, tool-using agents.

evals, llm evaluation, ai quality, benchmarks and learning resources

LLM Evals: Everything You Need to Know

Application-evaluation learning resource · Hamel Husain and Shreya Shankar · Intermediate to advanced

Practitioners building product-specific evals from traces, domain-expert judgments, and observed failures instead of relying on generic benchmarks.

evals, llm evaluation, ai quality, benchmarks and learning resources

NVIDIA Kumo Tabular explained

Tabular foundation model explainer · NVIDIA · Intermediate to advanced

NVIDIA released Kumo Tabular on September 29, 2026. Use this guide to understand its in-context prediction workflow, reproduce a baseline, and test its vendor-reported results on your own tables.

kumo tabular, tabular foundation models, classification, regression, in-context learning

The Anatomy of Harness Engineering

Coding agent evaluation guide · Google Developers Blog · Intermediate to advanced

You want Google's September 9, 2026 guide to evaluating coding agents with small behavioral checks, outcome-based assertions, and batch runs that catch regressions without treating a single benchmark score as the whole story.

coding agents, evals, behavioral evaluations, regression testing, harness engineering

An alignment assessment of recent cybersecurity incidents

AI safety incident assessment · Anthropic · Advanced

You want Anthropic's September 9, 2026 assessment of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations, including the model behaviors and evaluation-design failures involved.

anthropic, ai safety, cybersecurity, agent evaluation, alignment

The AI policy window is open. We need to act.

Frontier AI safety policy guide · OpenAI · Advanced

You want OpenAI's September 9, 2026 policy proposal on frontier AI standards, independent assessments, incident reporting, and preserving human control as capabilities advance.

openai, ai safety, frontier models, evaluations, governance

Agentic models, measured on the injections that move money

Open security benchmark · Hugging Face Community · Intermediate to advanced

You want a reproducible September 5, 2026 benchmark for comparing how agentic models handle indirect prompt injection, with public data, a public harness, control runs, tool-call traces, and outcome metrics tied to unauthorized payment actions.

agents, prompt injection, agent security, evals, tool use

How GitHub makes AI coding more cost efficient

Coding agent evaluation guide · GitHub · Intermediate to advanced

You want GitHub's September 2, 2026 evidence for measuring coding-agent efficiency across the whole task, including selective output compression, preserving useful context, benchmark regressions, and controlled production experiments.

github copilot, coding agents, context engineering, evals, cost optimization

Research acceleration: The view inside OpenAI

Research report · OpenAI · Intermediate

You want OpenAI's September 6, 2026 evidence and measurement framework for how coding agents are changing research workflows, including task delegation, parallel agent use, capability tracking, and the limits of interpreting productivity signals.

openai, codex, research agents, automated research, agent adoption

Project HydraFusion

Model orchestration research preview · GitHub · Intermediate to advanced

You want GitHub's September 4, 2026 technical explanation of runtime model orchestration for coding tasks, including plan decomposition, draft-critique-revise patterns, model cascading, evaluation design, and quality-versus-cost tradeoffs.

github copilot, model orchestration, model selection, coding agents, evals

Designing the hf CLI for agents

Agent tooling and evaluation guide · Hugging Face · Intermediate

You want a measured guide to agent-friendly CLI design, including structured output, retry-safe commands, independent grading, and benchmarks showing where higher-level tools reduce calls and token use.

hugging face, coding agents, hf cli, evals, codex