AI directory search

Search across educators, skills, and resources.

Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.

15 matches for "regression"

Providers and platforms

Promptfoo profile photo

Promptfoo

Promptfoo Docs · Intermediate

Very practical for regression testing prompts, model changes, and LLM outputs.

Topics

Prompt testing, Evals, Red teaming

Learning paths

Resources

Promptfoo

Open-source eval and red-team framework · OpenAI / Promptfoo · Intermediate to advanced

Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required.

evals, llm evaluation, ai quality, open-source frameworks

Braintrust

AI evaluation and observability platform · Braintrust Data · Intermediate to advanced

Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.

evals, llm evaluation, ai quality, evaluation platforms

Opik

Open-source agent evaluation platform · Comet · Intermediate to advanced

Teams wanting hosted or self-hosted observability with failure-driven regression suites, experiments, metrics, and human review.

evals, llm evaluation, ai quality, evaluation platforms

OpenAI Evals

Hosted LLM and agent evaluation API · OpenAI · Intermediate to advanced

Existing OpenAI API teams maintaining dataset-based prompt or model regression tests during the remaining service window.

evals, llm evaluation, ai quality, cloud and provider evals

NVIDIA Kumo Tabular explained

Tabular foundation model explainer · NVIDIA · Intermediate to advanced

NVIDIA released Kumo Tabular on September 29, 2026. Use this guide to understand its in-context prediction workflow, reproduce a baseline, and test its vendor-reported results on your own tables.

kumo tabular, tabular foundation models, classification, regression, in-context learning

The Anatomy of Harness Engineering

Coding agent evaluation guide · Google Developers Blog · Intermediate to advanced

You want Google's September 9, 2026 guide to evaluating coding agents with small behavioral checks, outcome-based assertions, and batch runs that catch regressions without treating a single benchmark score as the whole story.

coding agents, evals, behavioral evaluations, regression testing, harness engineering

How GitHub makes AI coding more cost efficient

Coding agent evaluation guide · GitHub · Intermediate to advanced

You want GitHub's September 2, 2026 evidence for measuring coding-agent efficiency across the whole task, including selective output compression, preserving useful context, benchmark regressions, and controlled production experiments.

github copilot, coding agents, context engineering, evals, cost optimization

How to Evaluate AI Agents

Evaluation guide · Hugging Face Community · Intermediate

You need a practical evaluation plan built around representative tasks, controlled environments, observable traces, outcome and constraint metrics, repeated trials, and production failures turned into regression tests.

hugging face, agents, evals, evaluation harnesses, trajectory metrics

OpenAI eval design guide

Guide · OpenAI · Intermediate

You need practical guidance for designing representative eval datasets, choosing graders, and turning model testing into an engineering loop instead of ad hoc spot checks, especially while OpenAI's older Evals platform is winding down toward read-only status on October 31, 2026.

openai, evals, quality, datasets, regression testing

OpenAI Working with evals

Guide · OpenAI · Intermediate

You need API-level guidance for testing outputs, comparing models, and catching regressions during upgrades.

openai, evals, quality, regression testing, reliability

OpenAI Evaluate agent workflows

Guide · OpenAI · Intermediate

You need the current OpenAI path for tracing, grading, and regression-testing agent workflows instead of only single-prompt evals.

openai, agents, evals, traces, graders

Promptfoo Intro video thumbnail ►

Promptfoo Intro

Open source docs · Promptfoo · Intermediate

You need regression tests for prompts, models, and LLM outputs.

evals, prompt testing, red teaming