►
LLM evaluation with W&B
Weights & Biases · evals, llm apps, observability, mlops
AI directory search
Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.
16 matches for "observability"
Watch first when you want a fast feel for the topic before opening courses, docs, or profiles.
Useful for debugging and evaluating LLM applications once you move beyond prototypes.
Topics
Observability, Evals, Tracing, RAG debugging
Langfuse Docs · Intermediate
Good operational material for tracing, scoring, and improving production LLM apps.
Topics
Observability, Prompt management, Evals, Tracing
Engineering teams
Learn first
Good matches
Open next
AI observability and evaluation platform · Arize AI · Intermediate to advanced
Teams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting.
evals, llm evaluation, ai quality, evaluation platforms
AI evaluation and observability platform · Braintrust Data · Intermediate to advanced
Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.
evals, llm evaluation, ai quality, evaluation platforms
Agent evaluation and observability platform · LangChain · Intermediate to advanced
Agent teams, especially LangChain and LangGraph users, that need tracing, datasets, experiments, human review, and production feedback in one system.
evals, llm evaluation, ai quality, evaluation platforms
Open-source evals and observability platform · ClickHouse · Intermediate to advanced
Teams prioritizing open-source data control and one workflow across tracing, prompts, datasets, experiments, annotation, and evaluation.
evals, llm evaluation, ai quality, evaluation platforms
Open-source agent evaluation platform · Comet · Intermediate to advanced
Teams wanting hosted or self-hosted observability with failure-driven regression suites, experiments, metrics, and human review.
evals, llm evaluation, ai quality, evaluation platforms
Coding agent workflow release notes · Qwen · Intermediate to advanced
You want Qwen Code's September 10, 2026 update on workflow run history, agent and token observability, context-usage inspection, combining Plan with YOLO mode, and named parallel channel tasks in isolated workspaces.
qwen code, coding agents, workflow observability, context management, planning
Coding agent guide · OpenAI · Intermediate to advanced
You need lifecycle hooks that run scripts or MCP tools around Codex sessions, prompts, tool calls, compaction, subagents, permissions, validation, logging, or persistent-memory workflows.
openai, codex, hooks, mcp, automation
►
Free course · Weights & Biases · Intermediate
You need to debug and measure LLM app quality.
evals, llm apps, observability
►
Open source tool and docs · Arize AI · Intermediate
You need to trace, inspect, and evaluate LLM app behavior.
evals, observability, tracing
►
Docs and cookbooks · Langfuse · Intermediate
You need production LLM tracing, scoring, and prompt operations.
observability, tracing, prompt management