►
LLM evaluation with W&B
Weights & Biases · evals, llm apps, observability, mlops
AI directory search
Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.
54 matches for "Evals"
Watch first when you want a fast feel for the topic before opening courses, docs, or profiles.
►
Weights & Biases · evals, llm apps, observability, mlops
►
Arize AI · evals, observability, tracing, rag debugging
►
Promptfoo · evals, prompt testing, red teaming, security
►
Hamel Husain and Shreya Shankar · evals, product, llm reliability
Hamel's AI evals guides · Intermediate to advanced
Very practical material on evaluating LLM apps before they disappoint users.
Skills
Evals, RAG, LLM product quality
AI Evals for Engineers and PMs · Intermediate
Useful if you need to judge whether an AI feature is actually improving.
Skills
Evals, LLM reliability, Product quality
AI Hero · Beginner to advanced
Practical developer-focused AI education across LLM fundamentals, AI SDK app development, MCP, Claude Code workflows, agent-ready codebases, evals, TDD, handoffs, and reusable skills such as /teach, /grill-me, /to-prd, /to-issues, /tdd, /triage, and /handoff.
Skills
AI coding, Claude Skills, Agentic workflows, AI SDK, MCP, LLM fundamentals, Personalized learning
Hamza Farooq on Maven · Beginner to intermediate
Useful for PMs who need to design, evaluate, and ship reliable AI systems beyond impressive demos.
Skills
Agentic AI, AI product strategy, Evals, Production AI
W&B Courses · Intermediate
Good for builders who need to measure, debug, and improve LLM apps rather than just demo them.
Topics
LLM apps, Evals, Experiment tracking, MLOps
OpenAI docs, Academy, and Cookbook · Beginner to advanced
Official model and implementation material for learning GPT-6 Astra and cost-sensitive GPT-5.6 choices, Codex workflows, subagents, memories, agent evals, MCP and connector patterns, retrieval, background jobs, prompt engineering, production best practices, model optimization, structured outputs, and OpenAI's Academy learning path.
Topics
GPT-6 Astra, GPT models, Reasoning models, Model selection, Agents, Subagents, RAG, Structured outputs, MCP, Evals, Memories
Useful for debugging and evaluating LLM applications once you move beyond prototypes.
Topics
Observability, Evals, Tracing, RAG debugging
Langfuse Docs · Intermediate
Good operational material for tracing, scoring, and improving production LLM apps.
Topics
Observability, Prompt management, Evals, Tracing
Vellum Guides · Beginner to intermediate
Useful for product and ops teams that need practical LLM product concepts without getting lost in research.
Topics
Prompt management, Evals, Workflow design
Humanloop Blog and Docs · Intermediate
Useful for teams building repeatable AI product processes around prompts, datasets, and evaluations.
Topics
Prompt management, Evals, LLM workflows
Promptfoo Docs · Intermediate
Very practical for regression testing prompts, model changes, and LLM outputs.
Topics
Prompt testing, Evals, Red teaming
Maven AI courses · Beginner to advanced
Useful discovery surface for live courses taught by practitioners across AI product, work, and engineering.
Topics
AI product, AI leadership, AI workflows, Evals
OpenRouter docs · Beginner to intermediate
Useful for learning model comparison, latest-family aliases, routing, fallback behavior, agent construction, project-specific evals, and API-compatible experimentation across proprietary and open model families.
Topics
Model routing, Model comparison, Auto Router, Agent SDK, Coding-agent harnesses, GPT models, Claude models, Gemini, Llama, Mistral, DeepSeek, Qwen, API examples, Evaluation
AI product teams
Learn first
Good matches
Open next
Workshop · Matt Pocock · Intermediate
You want a structured AI SDK v6 course that covers model choice, text and object generation, UI streams, agents, persistence, context engineering, evals, and advanced app patterns.
ai sdk, llm apps, agents, streaming, evals
Free tutorial · Matt Pocock · Beginner to intermediate
You want a guided path through core AI concepts, model selection, the AI engineering mindset, evals, and techniques for improving LLM-powered apps.
ai engineering, model selection, evals, llm apps
Guide · Hamel Husain · Intermediate
Your AI app needs quality checks before users see it.
evals, quality, llm apps
Short course · DeepLearning.AI · Intermediate
You need to test, trace, and improve agent workflows instead of judging only single LLM responses.
agent evals, evals, agents, reliability, tracing
Short course · DeepLearning.AI · Intermediate
You already know basic RAG and need better retrieval, evaluation, and production-quality patterns.
rag, evals, retrieval, llm apps, ai engineering
Open-source eval and red-team framework · OpenAI / Promptfoo · Intermediate to advanced
Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required.
evals, llm evaluation, ai quality, open-source frameworks
Pytest-style LLM evaluation framework · Confident AI · Intermediate to advanced
Engineering teams that want LLM and agent evaluations to behave like software unit tests, with thresholds, assertions, and CI-friendly failures.
evals, llm evaluation, ai quality, open-source frameworks
RAG and AI-application evaluation library · Vibrant Labs · Intermediate to advanced
Evaluating retrieval quality, grounded generation, agent tool use, and production-aligned test data for RAG applications.
evals, llm evaluation, ai quality, open-source frameworks
Model and agent evaluation harness · UK AI Security Institute · Intermediate to advanced
Rigorous, reproducible model and agent capability or safety evaluations involving tools, multi-turn interaction, coding, or sandboxed environments.
evals, llm evaluation, ai quality, open-source frameworks
AI observability and evaluation platform · Arize AI · Intermediate to advanced
Teams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting.
evals, llm evaluation, ai quality, evaluation platforms
Reusable evaluator library · LangChain · Intermediate to advanced
Python or TypeScript developers who want composable evaluator functions without adopting a full evaluation platform.
evals, llm evaluation, ai quality, open-source frameworks
Agent testing and red-team framework · Giskard AI · Intermediate to advanced
Behavioral tests and adversarial scans of multi-turn agents, chatbots, and RAG systems from a pytest-compatible workflow.
evals, llm evaluation, ai quality, open-source frameworks
AI evaluation and observability platform · Braintrust Data · Intermediate to advanced
Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.
evals, llm evaluation, ai quality, evaluation platforms
Agent evaluation and observability platform · LangChain · Intermediate to advanced
Agent teams, especially LangChain and LangGraph users, that need tracing, datasets, experiments, human review, and production feedback in one system.
evals, llm evaluation, ai quality, evaluation platforms
Open-source evals and observability platform · ClickHouse · Intermediate to advanced
Teams prioritizing open-source data control and one workflow across tracing, prompts, datasets, experiments, annotation, and evaluation.
evals, llm evaluation, ai quality, evaluation platforms
LLM and agent evaluation platform · Weights & Biases · Intermediate to advanced
Teams already using Weights & Biases or wanting versioned datasets, scorers, traces, and evaluation comparisons in an experiment-oriented workflow.
evals, llm evaluation, ai quality, evaluation platforms
Open-source agent evaluation platform · Comet · Intermediate to advanced
Teams wanting hosted or self-hosted observability with failure-driven regression suites, experiments, metrics, and human review.
evals, llm evaluation, ai quality, evaluation platforms
Open-source GenAI evaluation and monitoring · MLflow Project · Intermediate to advanced
Teams extending an existing MLflow or MLOps stack to LLM and agent tracing, evaluation-driven development, human feedback, and monitoring.
evals, llm evaluation, ai quality, evaluation platforms
AI reliability and evaluation platform · Patronus AI · Intermediate to advanced
Safety- and reliability-sensitive RAG or agent systems needing specialized hallucination, retrieval, PII, bias, policy, and custom-criteria evaluators.
evals, llm evaluation, ai quality, evaluation platforms
Hosted LLM and agent evaluation API · OpenAI · Intermediate to advanced
Existing OpenAI API teams maintaining dataset-based prompt or model regression tests during the remaining service window.
evals, llm evaluation, ai quality, cloud and provider evals
Managed model and agent evaluation service · Google Cloud · Intermediate to advanced
Vertex AI teams evaluating model outputs, prompts, RAG responses, function calls, and agent trajectories with managed or custom metrics.
evals, llm evaluation, ai quality, cloud and provider evals
Cloud evaluation portal and SDK · Microsoft · Intermediate to advanced
Microsoft Foundry teams evaluating models, RAG systems, and single- or multi-turn agents for quality, safety, task completion, and tool use.
evals, llm evaluation, ai quality, cloud and provider evals
Managed model and RAG evaluation jobs · Amazon Web Services · Intermediate to advanced
AWS teams comparing Bedrock models or assessing knowledge bases and external RAG sources with automated metrics, LLM judges, or human reviewers.
evals, llm evaluation, ai quality, cloud and provider evals
Evaluation methodology guide · Anthropic · Intermediate to advanced
Teams designing task-specific evaluation suites for Claude applications that need criteria, datasets, edge cases, grading patterns, and code examples.
evals, llm evaluation, ai quality, cloud and provider evals
Agent evaluation and optimization harness · Harbor Framework Team · Intermediate to advanced
Running coding and computer-use agents against reproducible, sandboxed task suites at local or cloud scale.
evals, llm evaluation, ai quality, benchmarks and learning resources
Language-model benchmark runner · EleutherAI · Intermediate to advanced
Reproducible few-shot and zero-shot evaluation of base or instruction-tuned language models on established academic benchmarks.
evals, llm evaluation, ai quality, benchmarks and learning resources
Multi-backend LLM evaluation toolkit · Hugging Face · Intermediate to advanced
Teams evaluating local, distributed, or API-served models across a large multilingual task catalog while retaining sample-level results.
evals, llm evaluation, ai quality, benchmarks and learning resources
Software-engineering agent benchmark · SWE-bench team · Intermediate to advanced
Measuring whether coding agents can resolve real GitHub issues by producing repository patches that pass executable tests.
evals, llm evaluation, ai quality, benchmarks and learning resources
Terminal-agent benchmark · Harbor Framework and Laude Institute · Intermediate to advanced
Comparing agents on difficult, verifiable coding, systems, security, and scientific work inside sandboxed terminals.
evals, llm evaluation, ai quality, benchmarks and learning resources
Agent-evaluation learning resource · Anthropic · Intermediate to advanced
Product and engineering teams designing practical evaluations for multi-turn, tool-using agents.
evals, llm evaluation, ai quality, benchmarks and learning resources
Application-evaluation learning resource · Hamel Husain and Shreya Shankar · Intermediate to advanced
Practitioners building product-specific evals from traces, domain-expert judgments, and observed failures instead of relying on generic benchmarks.
evals, llm evaluation, ai quality, benchmarks and learning resources
AI-assisted science explainer · Anthropic · Intermediate
Anthropic reported ART on September 23, 2026. Use this explainer to separate what its agents found, what the lab confirmed, and what remains a hypothesis.
claude, ai agents, ai for science, biology, genome mining
Frontier coding model launch explainer · SpaceXAI · Intermediate to advanced
SpaceXAI released Grok 4.7 on September 21, 2026 for coding, agentic tasks, and knowledge work. Use this guide to compare its task reliability and total cost with your current model.
grok 4.7, coding agents, long-running agents, model selection, evals
Multi-model routing launch explainer · Unbiased · Intermediate to advanced
Union Alpha was revealed as Unbiased's Pareto 26.9 on September 17, 2026: a hosted system that routes work across several models and returns one checked answer.
pareto 26.9, union alpha, model routing, ensembles, coding agents
Coding agent evaluation guide · Google Developers Blog · Intermediate to advanced
You want Google's September 9, 2026 guide to evaluating coding agents with small behavioral checks, outcome-based assertions, and batch runs that catch regressions without treating a single benchmark score as the whole story.
coding agents, evals, behavioral evaluations, regression testing, harness engineering
Agent security architecture guide · Meta AI · Intermediate to advanced
You want Meta's September 8, 2026 technical account of defense-in-depth for a long-running personal agent, including isolated runtime cells, credential surrogates, a separate permission authority, tainted-egress tracking, scoped approvals, browser controls, red teaming, and prompt-injection evals.
meta, muse, agent security, prompt injection, least privilege