AI learning guide

Best AI resources for production AI engineering

Learn deployment, observability, latency, cost, MLOps, and product quality.

Best production systems book: AI Engineering. Chip Huyen's resource for building reliable AI applications. Start here when a demo needs to become a system.

Best lifecycle course: Full Stack Deep Learning Lectures. Full Stack Deep Learning course videos on the ML and AI product lifecycle. Use it for deployment, iteration, product quality, and operating concerns.

Best tracing and eval tooling: Phoenix by Arize. Open-source observability and evaluation tooling. Use it when you need to see why an AI workflow failed.

Production AI is where demos meet constraints

A demo can ignore latency, cost, monitoring, user feedback, model upgrades, privacy, retries, and failure handling. Production AI cannot. The best resources teach the operating system around the model.

Chip Huyen is the strongest production AI starting point. Full Stack Deep Learning gives lifecycle context. Phoenix and Langfuse help when you need observability and evals in a running system. Pair those with current provider docs on batch processing, prompt caching, flex modes, and retrieval so your production choices match how the APIs actually behave now.

Build one measured feature

A good learning project should include logs, eval examples, model comparison, cost measurement, latency measurement, and a written list of failure modes. Without those pieces, it is still mostly a demo.

Avoid resources that imply production is just deployment. The hard part is knowing whether the system is good, whether it is improving, and what happens when it is wrong.

Recommended courses and resources

  1. AI Engineering

    Book · Chip Huyen · Intermediate to advanced

    You are moving from demos to production systems.

  2. Building and Evaluating Advanced RAG Applications

    Short course · DeepLearning.AI · Intermediate

    You already know basic RAG and need better retrieval, evaluation, and production-quality patterns.

  3. Ragas

    RAG and AI-application evaluation library · Vibrant Labs · Intermediate to advanced

    Evaluating retrieval quality, grounded generation, agent tool use, and production-aligned test data for RAG applications.

  4. Arize Phoenix

    AI observability and evaluation platform · Arize AI · Intermediate to advanced

    Teams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting.

  5. Braintrust

    AI evaluation and observability platform · Braintrust Data · Intermediate to advanced

    Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.

Roll a learning mission

Pick one small move from this guide instead of opening ten tabs.

About this guide

Author: Learnetto Editorial Team. Learnetto maintains this AI learning directory by organizing public course pages, official documentation, educator material, and practical learning resources.

How it is made: Learnetto uses public course pages, official documentation, educator material, and directory data to compile these recommendations. AI may help draft and organize the page, but recommendations are checked against the listed sources, page topic, and learner intent.

Review policy: We only add a named personal reviewer when that person has substantially reviewed the page. Until then, the page is attributed to Learnetto rather than a founder, editor, or individual expert.

Last updated: October 7, 2026. Suggest a correction if a course, doc, or recommendation is outdated.

Videos to watch

From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI video thumbnail ►

From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI

AI Engineer

LLM evaluation with W&B video thumbnail ►

LLM evaluation with W&B

Weights & Biases

AI evals with Phoenix video thumbnail ►

AI evals with Phoenix

Arize AI

AI Engineering with Chip Huyen video thumbnail ►

AI Engineering with Chip Huyen

Chip Huyen

Full Stack Deep Learning lecture video thumbnail ►

Full Stack Deep Learning lecture

Full Stack Deep Learning

MLOps community production AI video thumbnail ►

MLOps community production AI

MLOps Community

ML Zoomcamp supervised learning video thumbnail ►

ML Zoomcamp supervised learning

DataTalks.Club

Educators and sources

Educator / source Best for Skills Start with
Developers, AI engineers AI engineering, Agents, Developer tools Watch AI Engineer talks for production patterns and tool choices.
Engineers, ML practitioners AI engineering, Systems, Production ML Use the book page and related essays as a production engineering path.
Digital writers, founders, creators AI-assisted writing, Content systems, Personal brand, Idea development Use a writing template with AI as a first-pass collaborator, then rewrite in your own voice.
Product managers, AI product leaders, founders Agentic AI, AI product strategy, Evals, Production AI Use the course to evaluate one AI product opportunity and define what reliability would mean before implementation.
Data and AI practitioners Data systems, ML engineering, AI trends Search episodes by topic: RAG, evaluation, agents, MLOps.
Developers fine-tuning and deploying models Open models, Fine-tuning, Deployment, Transformers Pick one fine-tuning or inference guide and reproduce it end to end.
Developers and data science learners Machine learning, Deep learning, LLM apps, MLOps Pick a playlist that matches your current level and follow the code.

Resources

AI Engineering

Book · Chip Huyen · Intermediate to advanced

You are moving from demos to production systems.

Ragas

RAG and AI-application evaluation library · Vibrant Labs · Intermediate to advanced

Evaluating retrieval quality, grounded generation, agent tool use, and production-aligned test data for RAG applications.

Arize Phoenix

AI observability and evaluation platform · Arize AI · Intermediate to advanced

Teams that want evaluation connected to OpenTelemetry traces, datasets, experiments, prompt iterations, and production troubleshooting.

Braintrust

AI evaluation and observability platform · Braintrust Data · Intermediate to advanced

Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.

LangSmith

Agent evaluation and observability platform · LangChain · Intermediate to advanced

Agent teams, especially LangChain and LangGraph users, that need tracing, datasets, experiments, human review, and production feedback in one system.

Langfuse

Open-source evals and observability platform · ClickHouse · Intermediate to advanced

Teams prioritizing open-source data control and one workflow across tracing, prompts, datasets, experiments, annotation, and evaluation.

Opik

Open-source agent evaluation platform · Comet · Intermediate to advanced

Teams wanting hosted or self-hosted observability with failure-driven regression suites, experiments, metrics, and human review.

MLflow GenAI evaluation

Open-source GenAI evaluation and monitoring · MLflow Project · Intermediate to advanced

Teams extending an existing MLflow or MLOps stack to LLM and agent tracing, evaluation-driven development, human feedback, and monitoring.

Google's AI Principles

Responsible AI principles · Google · Beginner

Use this as Google's primary statement of the principles it applies to AI development and deployment, alongside the playbook's recommendation to build a framework for trust.

Qwen Code September 10 weekly update

Coding agent workflow release notes · Qwen · Intermediate to advanced

You want Qwen Code's September 10, 2026 update on workflow run history, agent and token observability, context-usage inspection, combining Plan with YOLO mode, and named parallel channel tasks in isolated workspaces.

How GitHub makes AI coding more cost efficient

Coding agent evaluation guide · GitHub · Intermediate to advanced

You want GitHub's September 2, 2026 evidence for measuring coding-agent efficiency across the whole task, including selective output compression, preserving useful context, benchmark regressions, and controlled production experiments.

Anthropic ant apply

Agent infrastructure guide · Anthropic · Intermediate to advanced

You want Anthropic's September 3, 2026 workflow for declaring agents, environments, skills, memory stores, and scheduled deployments in a repository, previewing changes, and applying them reproducibly in CI.

OpenAI Codex hooks

Coding agent guide · OpenAI · Intermediate to advanced

You need lifecycle hooks that run scripts or MCP tools around Codex sessions, prompts, tool calls, compaction, subagents, permissions, validation, logging, or persistent-memory workflows.

How to Evaluate AI Agents

Evaluation guide · Hugging Face Community · Intermediate

You need a practical evaluation plan built around representative tasks, controlled environments, observable traces, outcome and constraint metrics, repeated trials, and production failures turned into regression tests.

Gemini 3.8 Flash

Model docs · Google AI for Developers · Intermediate

You need Google's official September 2, 2026 GA reference for its newest Flash model, including the stable model ID, 1M-token input window, thinking levels, multimodal inputs, tools, and production inference options.