Humanloop
Humanloop Blog and Docs · Intermediate
Useful for teams building repeatable AI product processes around prompts, datasets, and evaluations.
Topics
Prompt management, Evals, LLM workflows
AI directory search
Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.
5 matches for "evaluations"
Humanloop Blog and Docs · Intermediate
Useful for teams building repeatable AI product processes around prompts, datasets, and evaluations.
Topics
Prompt management, Evals, LLM workflows
Coding agent evaluation guide · Google Developers Blog · Intermediate to advanced
You want Google's September 9, 2026 guide to evaluating coding agents with small behavioral checks, outcome-based assertions, and batch runs that catch regressions without treating a single benchmark score as the whole story.
coding agents, evals, behavioral evaluations, regression testing, harness engineering
AI safety incident assessment · Anthropic · Advanced
You want Anthropic's September 9, 2026 assessment of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations, including the model behaviors and evaluation-design failures involved.
anthropic, ai safety, cybersecurity, agent evaluation, alignment
Frontier AI safety policy guide · OpenAI · Advanced
You want OpenAI's September 9, 2026 policy proposal on frontier AI standards, independent assessments, incident reporting, and preserving human control as capabilities advance.
openai, ai safety, frontier models, evaluations, governance
Evaluation guide · Google AI for Developers · Intermediate to advanced
You want Google's practical July 31, 2026 guide to running the same agent and model evaluations during development and on production traffic, with metrics for quality, safety, grounding, tool use, and trajectories.
google, gemini, agents, evals, model selection