►
12GB Model, 8 Hours, One 3D Game: OrcaSAQ2 27B Tested
Fahd Mirza · 2026, local models, qwen3.8, hermes agent
AI directory search
Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.
46 matches for "evaluation"
Watch first when you want a fast feel for the topic before opening courses, docs, or profiles.
The Data Exchange · Intermediate
Good practitioner interviews across data, ML, and AI engineering.
Skills
Data systems, ML engineering, AI trends
Data Independent AI tutorials · Beginner to intermediate
Practical walkthroughs for retrieval, LLM application patterns, and common developer questions.
Skills
RAG, LLM apps, Prompting, Evaluation
W&B Courses · Intermediate
Good for builders who need to measure, debug, and improve LLM apps rather than just demo them.
Topics
LLM apps, Evals, Experiment tracking, MLOps
Humanloop Blog and Docs · Intermediate
Useful for teams building repeatable AI product processes around prompts, datasets, and evaluations.
Topics
Prompt management, Evals, LLM workflows
Stanford CS229 Machine Learning · Intermediate
A strong foundation for people who need the math and modeling basics under applied AI.
Topics
ML foundations, Supervised learning, Unsupervised learning, Model evaluation
OpenRouter docs · Beginner to intermediate
Useful for learning model comparison, latest-family aliases, routing, fallback behavior, agent construction, project-specific evals, and API-compatible experimentation across proprietary and open model families.
Topics
Model routing, Model comparison, Auto Router, Agent SDK, Coding-agent harnesses, GPT models, Claude models, Gemini, Llama, Mistral, DeepSeek, Qwen, API examples, Evaluation
Fireworks AI model catalog · Intermediate to advanced
Official material for comparing serverless and dedicated model serving, training specialized models, and measuring quality, token use, cost, and task duration on production-shaped evaluations.
Topics
Ember-1, Kimi K3, Serverless inference, Dedicated deployment, Model training, Model evaluation, Reasoning efficiency
Short course · DeepLearning.AI · Intermediate
You already know basic RAG and need better retrieval, evaluation, and production-quality patterns.
rag, evals, retrieval, llm apps, ai engineering
Open-source eval and red-team framework · OpenAI / Promptfoo · Intermediate to advanced
Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required.
evals, llm evaluation, ai quality, open-source frameworks
Pytest-style LLM evaluation framework · Confident AI · Intermediate to advanced
Engineering teams that want LLM and agent evaluations to behave like software unit tests, with thresholds, assertions, and CI-friendly failures.
evals, llm evaluation, ai quality, open-source frameworks
RAG and AI-application evaluation library · Vibrant Labs · Intermediate to advanced
Evaluating retrieval quality, grounded generation, agent tool use, and production-aligned test data for RAG applications.
evals, llm evaluation, ai quality, open-source frameworks
Model and agent evaluation harness · UK AI Security Institute · Intermediate to advanced
Rigorous, reproducible model and agent capability or safety evaluations involving tools, multi-turn interaction, coding, or sandboxed environments.
evals, llm evaluation, ai quality, open-source frameworks
AI observability and evaluation platform · Arize AI · Intermediate to advanced
Teams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting.
evals, llm evaluation, ai quality, evaluation platforms
Reusable evaluator library · LangChain · Intermediate to advanced
Python or TypeScript developers who want composable evaluator functions without adopting a full evaluation platform.
evals, llm evaluation, ai quality, open-source frameworks
Agent testing and red-team framework · Giskard AI · Intermediate to advanced
Behavioral tests and adversarial scans of multi-turn agents, chatbots, and RAG systems from a pytest-compatible workflow.
evals, llm evaluation, ai quality, open-source frameworks
AI evaluation and observability platform · Braintrust Data · Intermediate to advanced
Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.
evals, llm evaluation, ai quality, evaluation platforms
Agent evaluation and observability platform · LangChain · Intermediate to advanced
Agent teams, especially LangChain and LangGraph users, that need tracing, datasets, experiments, human review, and production feedback in one system.
evals, llm evaluation, ai quality, evaluation platforms
Open-source evals and observability platform · ClickHouse · Intermediate to advanced
Teams prioritizing open-source data control and one workflow across tracing, prompts, datasets, experiments, annotation, and evaluation.
evals, llm evaluation, ai quality, evaluation platforms
LLM and agent evaluation platform · Weights & Biases · Intermediate to advanced
Teams already using Weights & Biases or wanting versioned datasets, scorers, traces, and evaluation comparisons in an experiment-oriented workflow.
evals, llm evaluation, ai quality, evaluation platforms
Open-source agent evaluation platform · Comet · Intermediate to advanced
Teams wanting hosted or self-hosted observability with failure-driven regression suites, experiments, metrics, and human review.
evals, llm evaluation, ai quality, evaluation platforms
Open-source GenAI evaluation and monitoring · MLflow Project · Intermediate to advanced
Teams extending an existing MLflow or MLOps stack to LLM and agent tracing, evaluation-driven development, human feedback, and monitoring.
evals, llm evaluation, ai quality, evaluation platforms
AI reliability and evaluation platform · Patronus AI · Intermediate to advanced
Safety- and reliability-sensitive RAG or agent systems needing specialized hallucination, retrieval, PII, bias, policy, and custom-criteria evaluators.
evals, llm evaluation, ai quality, evaluation platforms
Hosted LLM and agent evaluation API · OpenAI · Intermediate to advanced
Existing OpenAI API teams maintaining dataset-based prompt or model regression tests during the remaining service window.
evals, llm evaluation, ai quality, cloud and provider evals
Managed model and agent evaluation service · Google Cloud · Intermediate to advanced
Vertex AI teams evaluating model outputs, prompts, RAG responses, function calls, and agent trajectories with managed or custom metrics.
evals, llm evaluation, ai quality, cloud and provider evals
Cloud evaluation portal and SDK · Microsoft · Intermediate to advanced
Microsoft Foundry teams evaluating models, RAG systems, and single- or multi-turn agents for quality, safety, task completion, and tool use.
evals, llm evaluation, ai quality, cloud and provider evals
Managed model and RAG evaluation jobs · Amazon Web Services · Intermediate to advanced
AWS teams comparing Bedrock models or assessing knowledge bases and external RAG sources with automated metrics, LLM judges, or human reviewers.
evals, llm evaluation, ai quality, cloud and provider evals
Evaluation methodology guide · Anthropic · Intermediate to advanced
Teams designing task-specific evaluation suites for Claude applications that need criteria, datasets, edge cases, grading patterns, and code examples.
evals, llm evaluation, ai quality, cloud and provider evals
Agent evaluation and optimization harness · Harbor Framework Team · Intermediate to advanced
Running coding and computer-use agents against reproducible, sandboxed task suites at local or cloud scale.
evals, llm evaluation, ai quality, benchmarks and learning resources
Language-model benchmark runner · EleutherAI · Intermediate to advanced
Reproducible few-shot and zero-shot evaluation of base or instruction-tuned language models on established academic benchmarks.
evals, llm evaluation, ai quality, benchmarks and learning resources
Multi-backend LLM evaluation toolkit · Hugging Face · Intermediate to advanced
Teams evaluating local, distributed, or API-served models across a large multilingual task catalog while retaining sample-level results.
evals, llm evaluation, ai quality, benchmarks and learning resources
Software-engineering agent benchmark · SWE-bench team · Intermediate to advanced
Measuring whether coding agents can resolve real GitHub issues by producing repository patches that pass executable tests.
evals, llm evaluation, ai quality, benchmarks and learning resources
Terminal-agent benchmark · Harbor Framework and Laude Institute · Intermediate to advanced
Comparing agents on difficult, verifiable coding, systems, security, and scientific work inside sandboxed terminals.
evals, llm evaluation, ai quality, benchmarks and learning resources
Agent-evaluation learning resource · Anthropic · Intermediate to advanced
Product and engineering teams designing practical evaluations for multi-turn, tool-using agents.
evals, llm evaluation, ai quality, benchmarks and learning resources
Application-evaluation learning resource · Hamel Husain and Shreya Shankar · Intermediate to advanced
Practitioners building product-specific evals from traces, domain-expert judgments, and observed failures instead of relying on generic benchmarks.
evals, llm evaluation, ai quality, benchmarks and learning resources
Tabular foundation model explainer · NVIDIA · Intermediate to advanced
NVIDIA released Kumo Tabular on September 29, 2026. Use this guide to understand its in-context prediction workflow, reproduce a baseline, and test its vendor-reported results on your own tables.
kumo tabular, tabular foundation models, classification, regression, in-context learning
Coding agent evaluation guide · Google Developers Blog · Intermediate to advanced
You want Google's September 9, 2026 guide to evaluating coding agents with small behavioral checks, outcome-based assertions, and batch runs that catch regressions without treating a single benchmark score as the whole story.
coding agents, evals, behavioral evaluations, regression testing, harness engineering
AI safety incident assessment · Anthropic · Advanced
You want Anthropic's September 9, 2026 assessment of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations, including the model behaviors and evaluation-design failures involved.
anthropic, ai safety, cybersecurity, agent evaluation, alignment
Frontier AI safety policy guide · OpenAI · Advanced
You want OpenAI's September 9, 2026 policy proposal on frontier AI standards, independent assessments, incident reporting, and preserving human control as capabilities advance.
openai, ai safety, frontier models, evaluations, governance
Open security benchmark · Hugging Face Community · Intermediate to advanced
You want a reproducible September 5, 2026 benchmark for comparing how agentic models handle indirect prompt injection, with public data, a public harness, control runs, tool-call traces, and outcome metrics tied to unauthorized payment actions.
agents, prompt injection, agent security, evals, tool use
Coding agent evaluation guide · GitHub · Intermediate to advanced
You want GitHub's September 2, 2026 evidence for measuring coding-agent efficiency across the whole task, including selective output compression, preserving useful context, benchmark regressions, and controlled production experiments.
github copilot, coding agents, context engineering, evals, cost optimization
Research report · OpenAI · Intermediate
You want OpenAI's September 6, 2026 evidence and measurement framework for how coding agents are changing research workflows, including task delegation, parallel agent use, capability tracking, and the limits of interpreting productivity signals.
openai, codex, research agents, automated research, agent adoption
Model orchestration research preview · GitHub · Intermediate to advanced
You want GitHub's September 4, 2026 technical explanation of runtime model orchestration for coding tasks, including plan decomposition, draft-critique-revise patterns, model cascading, evaluation design, and quality-versus-cost tradeoffs.
github copilot, model orchestration, model selection, coding agents, evals
Agent tooling and evaluation guide · Hugging Face · Intermediate
You want a measured guide to agent-friendly CLI design, including structured output, retry-safe commands, independent grading, and benchmarks showing where higher-level tools reduce calls and token use.
hugging face, coding agents, hf cli, evals, codex