Promptfoo
Promptfoo Docs · Intermediate
Very practical for regression testing prompts, model changes, and LLM outputs.
Topics
Prompt testing, Evals, Red teaming
AI directory search
Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.
15 matches for "regression"
Promptfoo Docs · Intermediate
Very practical for regression testing prompts, model changes, and LLM outputs.
Topics
Prompt testing, Evals, Red teaming
AI product teams
Learn first
Good matches
Open next
Guide · Matt Pocock · Intermediate
You want TypeScript checks, tests, linters, and review loops that help agents produce better code and catch regressions quickly.
ai coding, typescript, testing, feedback loops
Open-source eval and red-team framework · OpenAI / Promptfoo · Intermediate to advanced
Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required.
evals, llm evaluation, ai quality, open-source frameworks
AI evaluation and observability platform · Braintrust Data · Intermediate to advanced
Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.
evals, llm evaluation, ai quality, evaluation platforms
Open-source agent evaluation platform · Comet · Intermediate to advanced
Teams wanting hosted or self-hosted observability with failure-driven regression suites, experiments, metrics, and human review.
evals, llm evaluation, ai quality, evaluation platforms
Hosted LLM and agent evaluation API · OpenAI · Intermediate to advanced
Existing OpenAI API teams maintaining dataset-based prompt or model regression tests during the remaining service window.
evals, llm evaluation, ai quality, cloud and provider evals
Tabular foundation model explainer · NVIDIA · Intermediate to advanced
NVIDIA released Kumo Tabular on September 29, 2026. Use this guide to understand its in-context prediction workflow, reproduce a baseline, and test its vendor-reported results on your own tables.
kumo tabular, tabular foundation models, classification, regression, in-context learning
Coding agent evaluation guide · Google Developers Blog · Intermediate to advanced
You want Google's September 9, 2026 guide to evaluating coding agents with small behavioral checks, outcome-based assertions, and batch runs that catch regressions without treating a single benchmark score as the whole story.
coding agents, evals, behavioral evaluations, regression testing, harness engineering
Coding agent evaluation guide · GitHub · Intermediate to advanced
You want GitHub's September 2, 2026 evidence for measuring coding-agent efficiency across the whole task, including selective output compression, preserving useful context, benchmark regressions, and controlled production experiments.
github copilot, coding agents, context engineering, evals, cost optimization
Evaluation guide · Hugging Face Community · Intermediate
You need a practical evaluation plan built around representative tasks, controlled environments, observable traces, outcome and constraint metrics, repeated trials, and production failures turned into regression tests.
hugging face, agents, evals, evaluation harnesses, trajectory metrics
Guide · OpenAI · Intermediate
You need practical guidance for designing representative eval datasets, choosing graders, and turning model testing into an engineering loop instead of ad hoc spot checks, especially while OpenAI's older Evals platform is winding down toward read-only status on October 31, 2026.
openai, evals, quality, datasets, regression testing
Guide · OpenAI · Intermediate
You need API-level guidance for testing outputs, comparing models, and catching regressions during upgrades.
openai, evals, quality, regression testing, reliability
Guide · OpenAI · Intermediate
You need the current OpenAI path for tracing, grading, and regression-testing agent workflows instead of only single-prompt evals.
openai, agents, evals, traces, graders
►
Open source docs · Promptfoo · Intermediate
You need regression tests for prompts, models, and LLM outputs.
evals, prompt testing, red teaming