Hamel Husain
Hamel's AI evals guides · Intermediate to advanced
Very practical material on evaluating LLM apps before they disappoint users.
Skills
Evals, RAG, LLM product quality
AI directory search
Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.
24 matches for "quality"
Hamel's AI evals guides · Intermediate to advanced
Very practical material on evaluating LLM apps before they disappoint users.
Skills
Evals, RAG, LLM product quality
AI Evals for Engineers and PMs · Intermediate
Useful if you need to judge whether an AI feature is actually improving.
Skills
Evals, LLM reliability, Product quality
Foundation Marketing · Beginner to intermediate
Good for marketers learning how AI fits content strategy, distribution, and research without ignoring quality.
Skills
AI content marketing, Distribution, SEO, Content systems
SparkToro · Beginner to intermediate
Clear strategic context for how AI changes discovery, search, and audience behavior.
Skills
AI search, Audience research, Marketing strategy, Content quality
Neural Networks series · Beginner to intermediate
High-quality visual intuition for neural networks and transformers before or alongside coding-heavy courses.
Topics
Neural networks, Transformers, Mathematical intuition
PMs, designers, founders
Learn first
Good matches
Open next
AI product teams
Learn first
Good matches
Open next
Guide · Hamel Husain · Intermediate
Your AI app needs quality checks before users see it.
evals, quality, llm apps
Short course · DeepLearning.AI · Intermediate
You already know basic RAG and need better retrieval, evaluation, and production-quality patterns.
rag, evals, retrieval, llm apps, ai engineering
Specialization · Duke University · Beginner to intermediate
You want a structured product-management route for scoping, evaluating, and shipping AI products.
ai product, product management, ai strategy, product quality
Model orchestration research preview · GitHub · Intermediate to advanced
You want GitHub's September 4, 2026 technical explanation of runtime model orchestration for coding tasks, including plan decomposition, draft-critique-revise patterns, model cascading, evaluation design, and quality-versus-cost tradeoffs.
github copilot, model orchestration, model selection, coding agents, evals
Evaluation guide · Google AI for Developers · Intermediate to advanced
You want Google's practical July 31, 2026 guide to running the same agent and model evaluations during development and on production traffic, with metrics for quality, safety, grounding, tool use, and trajectories.
google, gemini, agents, evals, model selection
Practical guide · Hugging Face · Intermediate to advanced
You want a runnable guide to ColBERT-style late-interaction retrieval, including MaxSim scoring, retrieve-and-rerank, indexing, visual document retrieval, evaluation, and the storage-quality tradeoff versus dense embeddings.
hugging face, sentence transformers, rag, retrieval, embeddings
Guide · OpenAI · Intermediate
You need practical guidance for designing representative eval datasets, choosing graders, and turning model testing into an engineering loop instead of ad hoc spot checks, especially while OpenAI's older Evals platform is winding down toward read-only status on October 31, 2026.
openai, evals, quality, datasets, regression testing
Guide · OpenAI · Intermediate
You want OpenAI's current quickstart for turning examples into dataset-backed evals and improvement loops instead of relying on a deprecated docs path.
openai, datasets, evals, fine-tuning, quality
Guide · OpenAI · Intermediate
You need API-level guidance for testing outputs, comparing models, and catching regressions during upgrades.
openai, evals, quality, regression testing, reliability
Guide · OpenAI · Intermediate
You need the official pattern for compressing long agent conversations and preserving the right context instead of letting transcripts grow until quality or cost breaks down.
openai, compaction, context management, long-running agents, reasoning
GitHub repo · Perplexity · Intermediate
You want Perplexity's open evaluation suites for benchmarking grounded search quality and comparing search API behavior against other retrieval stacks.
perplexity, search evals, benchmarks, grounded answers, evaluation
Guide · OpenRouter · Intermediate to advanced
You want the official OpenRouter guide for routing across the quality, price, and latency frontier instead of hard-coding one tradeoff for every workload.
openrouter, pareto router, routing, latency, cost
Guide · Cohere · Intermediate
You are learning RAG quality work and want a focused official guide on reranking before debugging retrieval only through prompt changes.
cohere, reranking, rag, retrieval, search relevance
►
Free course · Weights & Biases · Intermediate
You need to debug and measure LLM app quality.
evals, llm apps, observability
Guides · Hamel Husain · Intermediate to advanced
Use this when you want Hamel Husain's material for evals and related AI skills.
Evals, RAG, LLM product quality
Course · Shreya Shankar · Intermediate
Use this when you want Shreya Shankar's material for evals and related AI skills.
Evals, LLM reliability, Product quality
Blog · Rand Fishkin · Beginner to intermediate
Use this when you want Rand Fishkin's material for ai search and related AI skills.
AI search, Audience research, Marketing strategy, Content quality