AI directory search

Search across educators, skills, and resources.

Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.

21 matches for "testing"

Video matches

Watch first when you want a fast feel for the topic before opening courses, docs, or profiles.

Promptfoo red teaming video thumbnail ►

Promptfoo red teaming

Promptfoo · evals, prompt testing, red teaming, security

Made With ML video thumbnail ►

Made With ML

Made With ML · mlops, testing, deployment

Providers and platforms

Promptfoo profile photo

Promptfoo

Promptfoo Docs · Intermediate

Very practical for regression testing prompts, model changes, and LLM outputs.

Topics

Prompt testing, Evals, Red teaming

Made With ML profile photo

Made With ML

Made With ML · Intermediate

Useful path for production ML fundamentals that transfer to AI engineering.

Topics

MLOps, Testing, Deployment, ML systems

Ollama profile photo

Ollama

Ollama docs · Beginner to intermediate

Practical route into running and testing local models on your own machine.

Topics

Local models, LLM tools, AI engineering, Privacy

Agent tools and skill directories

Workflow skill catalog

AI Skills for Real Engineers

Matt Pocock / AI Hero · Skill catalog

Use this when you want opinionated coding-agent workflows instead of generic prompt snippets.

Start

Start with /teach, /grill-me, /to-prd, /to-issues, /tdd, /triage, or /handoff depending on the job.

Resources

My Skill Makes Claude Code GREAT At TDD

Guide / Claude skill · Matt Pocock · Intermediate

You want an agent workflow that implements behavior with a red, green, refactor loop instead of jumping straight to broad code changes.

claude skills, tdd, ai coding, testing

Giskard

Agent testing and red-team framework · Giskard AI · Intermediate to advanced

Behavioral tests and adversarial scans of multi-turn agents, chatbots, and RAG systems from a pytest-compatible workflow.

evals, llm evaluation, ai quality, open-source frameworks

The Anatomy of Harness Engineering

Coding agent evaluation guide · Google Developers Blog · Intermediate to advanced

You want Google's September 9, 2026 guide to evaluating coding agents with small behavioral checks, outcome-based assertions, and batch runs that catch regressions without treating a single benchmark score as the whole story.

coding agents, evals, behavioral evaluations, regression testing, harness engineering

Modernizing complex legacy code with AI agents

Agent migration case study · Mistral AI · Intermediate to advanced

You want Mistral's September 9, 2026 field report on migrating 40,000 lines of Fortran 77 to C++, including a parity harness built before migration, agent-generated documentation, planner-coder-tester-reviewer workflows, and the human checkpoints needed for maintainable results.

mistral, coding agents, legacy modernization, parity testing, agent workflows

Prompting Claude Fable 5.1

Prompting guide · Anthropic · Intermediate to advanced

You are testing or migrating to Claude Fable 5.1 and want model-specific guidance for effort, progress updates, tool-call batching, long tasks, compaction, subagents, file edits, and verification.

anthropic, claude fable 5.1, prompting, coding agents, tool use

How to Choose the Best AI Model (Live, in Your Editor)

Model selection guide · OpenRouter · Intermediate

You want a practical six-step workflow for shortlisting models from live data, testing them on your own prompts, measuring cost per completed task, and choosing or routing from inside a coding assistant.

openrouter, model selection, evals, benchmarks, cost

OpenAI eval design guide

Guide · OpenAI · Intermediate

You need practical guidance for designing representative eval datasets, choosing graders, and turning model testing into an engineering loop instead of ad hoc spot checks, especially while OpenAI's older Evals platform is winding down toward read-only status on October 31, 2026.

openai, evals, quality, datasets, regression testing

OpenAI Working with evals

Guide · OpenAI · Intermediate

You need API-level guidance for testing outputs, comparing models, and catching regressions during upgrades.

openai, evals, quality, regression testing, reliability

OpenAI Evaluate agent workflows

Guide · OpenAI · Intermediate

You need the current OpenAI path for tracing, grading, and regression-testing agent workflows instead of only single-prompt evals.

openai, agents, evals, traces, graders

OpenRouter models guide

Models guide · OpenRouter · Beginner to advanced

You need to compare many model families through one catalog before testing prompts across providers.

openrouter, model comparison, model routing, opus, claude

OpenAI model selection cookbook

Cookbook guide · OpenAI · Intermediate

You want a practical OpenAI walkthrough for model selection tradeoffs, eval design, and rollout testing instead of treating model choice as a static table lookup.

openai, model selection, evals, latency, cost

xAI Grok Build 0.1

Model docs · xAI · Intermediate

You want the specific Grok Build model details, pricing, and capabilities before testing xAI for agentic coding work.

xai, grok build, coding agents, model selection, agentic coding

Promptfoo Intro video thumbnail ►

Promptfoo Intro

Open source docs · Promptfoo · Intermediate

You need regression tests for prompts, models, and LLM outputs.

evals, prompt testing, red teaming

Made With ML video thumbnail ►

Made With ML

Free course · Made With ML · Intermediate

You need production ML habits that transfer to AI systems.

mlops, testing, deployment