►
Promptfoo red teaming
Promptfoo · evals, prompt testing, red teaming, security
AI directory search
Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.
21 matches for "testing"
Promptfoo Docs · Intermediate
Very practical for regression testing prompts, model changes, and LLM outputs.
Topics
Prompt testing, Evals, Red teaming
Made With ML · Intermediate
Useful path for production ML fundamentals that transfer to AI engineering.
Topics
MLOps, Testing, Deployment, ML systems
Ollama docs · Beginner to intermediate
Practical route into running and testing local models on your own machine.
Topics
Local models, LLM tools, AI engineering, Privacy
Workflow skill catalog
Matt Pocock / AI Hero · Skill catalog
Use this when you want opinionated coding-agent workflows instead of generic prompt snippets.
Start
Start with /teach, /grill-me, /to-prd, /to-issues, /tdd, /triage, or /handoff depending on the job.
Guide · Matt Pocock · Intermediate
You want TypeScript checks, tests, linters, and review loops that help agents produce better code and catch regressions quickly.
ai coding, typescript, testing, feedback loops
Guide / Claude skill · Matt Pocock · Intermediate
You want an agent workflow that implements behavior with a red, green, refactor loop instead of jumping straight to broad code changes.
claude skills, tdd, ai coding, testing
Agent testing and red-team framework · Giskard AI · Intermediate to advanced
Behavioral tests and adversarial scans of multi-turn agents, chatbots, and RAG systems from a pytest-compatible workflow.
evals, llm evaluation, ai quality, open-source frameworks
Coding agent evaluation guide · Google Developers Blog · Intermediate to advanced
You want Google's September 9, 2026 guide to evaluating coding agents with small behavioral checks, outcome-based assertions, and batch runs that catch regressions without treating a single benchmark score as the whole story.
coding agents, evals, behavioral evaluations, regression testing, harness engineering
Agent migration case study · Mistral AI · Intermediate to advanced
You want Mistral's September 9, 2026 field report on migrating 40,000 lines of Fortran 77 to C++, including a parity harness built before migration, agent-generated documentation, planner-coder-tester-reviewer workflows, and the human checkpoints needed for maintainable results.
mistral, coding agents, legacy modernization, parity testing, agent workflows
Prompting guide · Anthropic · Intermediate to advanced
You are testing or migrating to Claude Fable 5.1 and want model-specific guidance for effort, progress updates, tool-call batching, long tasks, compaction, subagents, file edits, and verification.
anthropic, claude fable 5.1, prompting, coding agents, tool use
Model selection guide · OpenRouter · Intermediate
You want a practical six-step workflow for shortlisting models from live data, testing them on your own prompts, measuring cost per completed task, and choosing or routing from inside a coding assistant.
openrouter, model selection, evals, benchmarks, cost
Guide · OpenAI · Intermediate
You need practical guidance for designing representative eval datasets, choosing graders, and turning model testing into an engineering loop instead of ad hoc spot checks, especially while OpenAI's older Evals platform is winding down toward read-only status on October 31, 2026.
openai, evals, quality, datasets, regression testing
Guide · OpenAI · Intermediate
You need API-level guidance for testing outputs, comparing models, and catching regressions during upgrades.
openai, evals, quality, regression testing, reliability
Guide · OpenAI · Intermediate
You need the current OpenAI path for tracing, grading, and regression-testing agent workflows instead of only single-prompt evals.
openai, agents, evals, traces, graders
Models guide · OpenRouter · Beginner to advanced
You need to compare many model families through one catalog before testing prompts across providers.
openrouter, model comparison, model routing, opus, claude
Cookbook guide · OpenAI · Intermediate
You want a practical OpenAI walkthrough for model selection tradeoffs, eval design, and rollout testing instead of treating model choice as a static table lookup.
openai, model selection, evals, latency, cost
Model docs · xAI · Intermediate
You want the specific Grok Build model details, pricing, and capabilities before testing xAI for agentic coding work.
xai, grok build, coding agents, model selection, agentic coding
►
Open source docs · Promptfoo · Intermediate
You need regression tests for prompts, models, and LLM outputs.
evals, prompt testing, red teaming
►
Free course · Made With ML · Intermediate
You need production ML habits that transfer to AI systems.
mlops, testing, deployment