AI directory search

Search across educators, skills, and resources.

Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.

3 matches for "agent evaluation"

Resources

The Anatomy of Harness Engineering

Coding agent evaluation guide · Google Developers Blog · Intermediate to advanced

You want Google's September 9, 2026 guide to evaluating coding agents with small behavioral checks, outcome-based assertions, and batch runs that catch regressions without treating a single benchmark score as the whole story.

coding agents, evals, behavioral evaluations, regression testing, harness engineering

An alignment assessment of recent cybersecurity incidents

AI safety incident assessment · Anthropic · Advanced

You want Anthropic's September 9, 2026 assessment of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations, including the model behaviors and evaluation-design failures involved.

anthropic, ai safety, cybersecurity, agent evaluation, alignment

How GitHub makes AI coding more cost efficient

Coding agent evaluation guide · GitHub · Intermediate to advanced

You want GitHub's September 2, 2026 evidence for measuring coding-agent efficiency across the whole task, including selective output compression, preserving useful context, benchmark regressions, and controlled production experiments.

github copilot, coding agents, context engineering, evals, cost optimization