AI directory search

Search across educators, skills, and resources.

Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.

5 matches for "evaluations"

Providers and platforms

Useful for teams building repeatable AI product processes around prompts, datasets, and evaluations.

Topics

Prompt management, Evals, LLM workflows

Resources

The Anatomy of Harness Engineering

Coding agent evaluation guide · Google Developers Blog · Intermediate to advanced

You want Google's September 9, 2026 guide to evaluating coding agents with small behavioral checks, outcome-based assertions, and batch runs that catch regressions without treating a single benchmark score as the whole story.

coding agents, evals, behavioral evaluations, regression testing, harness engineering

An alignment assessment of recent cybersecurity incidents

AI safety incident assessment · Anthropic · Advanced

You want Anthropic's September 9, 2026 assessment of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations, including the model behaviors and evaluation-design failures involved.

anthropic, ai safety, cybersecurity, agent evaluation, alignment

The AI policy window is open. We need to act.

Frontier AI safety policy guide · OpenAI · Advanced

You want OpenAI's September 9, 2026 policy proposal on frontier AI standards, independent assessments, incident reporting, and preserving human control as capabilities advance.

openai, ai safety, frontier models, evaluations, governance

Agent and Model Evaluations in Gemini Enterprise Agent Platform

Evaluation guide · Google AI for Developers · Intermediate to advanced

You want Google's practical July 31, 2026 guide to running the same agent and model evaluations during development and on production traffic, with metrics for quality, safety, grounding, tool use, and trajectories.

google, gemini, agents, evals, model selection