AI learning guide

Best AI resources for local models

Run open models locally and understand inference, privacy, and tradeoffs.

Best discovery hub: Hugging Face model hub. Hugging Face catalog for open model checkpoints, datasets, and demos. Start here when comparing open models and practical availability.

Best Llama primary source: Llama API models. Official Meta Llama documentation. Use it for access, deployment, integrations, and model-family details.

Best fast Qwen start: Qwen quickstart. Official Qwen deployment quickstart. Use it when you want a practical route into running Qwen models locally or hosted.

Local models are about tradeoffs

Running models locally can help with privacy, control, latency, offline work, and experimentation. It can also create hardware limits, weaker quality, operational burden, and confusing setup.

Start with Hugging Face to discover models, then use official Llama or Qwen docs when you want to run or deploy a specific family. Treat model cards and licenses as part of the learning material.

Test the actual workload

A local model that looks good in a chat demo may fail at extraction, coding, multilingual work, or long-context retrieval. Choose a resource that teaches you how to test your own tasks, not only install a model.

Pay attention to quantization, context length, memory, inference speed, tool support, and hosting route. Those practical constraints decide whether local AI is useful for your case.

Recommended courses and resources

  1. How to Cost Your AI-Powered Filters

    Technical article · FS Data Lab · Advanced

    FS Data Lab published this guide on October 1, 2026. Use it when you want a rigorous way to estimate GPU latency and inference cost for AI-powered SQL filters, compare filter orderings, and understand how compute, memory bandwidth, selectivity, batching, and KV-cache reuse affect a query plan.

  2. NVIDIA Kumo Tabular explained

    Tabular foundation model explainer · NVIDIA · Intermediate to advanced

    NVIDIA released Kumo Tabular on September 29, 2026. Use this guide to understand its in-context prediction workflow, reproduce a baseline, and test its vendor-reported results on your own tables.

  3. Gemini 3.8 Flash

    Model docs · Google AI for Developers · Intermediate

    You need Google's official September 2, 2026 GA reference for its newest Flash model, including the stable model ID, 1M-token input window, thinking levels, multimodal inputs, tools, and production inference options.

  4. OpenAI external models

    Guide · OpenAI · Intermediate

    You want to learn how OpenAI handles access to non-OpenAI model families before designing a mixed-provider or routed workflow.

  5. Llama 4 prompt formats

    Model docs · Meta Llama · Beginner to advanced

    You need the current official Llama 4 prompt formats and model-card details before choosing Maverick or Scout for hosted or local deployment.

Roll a learning mission

Pick one small move from this guide instead of opening ten tabs.

About this guide

Author: Learnetto Editorial Team. Learnetto maintains this AI learning directory by organizing public course pages, official documentation, educator material, and practical learning resources.

How it is made: Learnetto uses public course pages, official documentation, educator material, and directory data to compile these recommendations. AI may help draft and organize the page, but recommendations are checked against the listed sources, page topic, and learner intent.

Review policy: We only add a named personal reviewer when that person has substantially reviewed the page. Until then, the page is attributed to Learnetto rather than a founder, editor, or individual expert.

Last updated: October 7, 2026. Suggest a correction if a course, doc, or recommendation is outdated.

Videos to watch

Don't Use Claude, Use This 320B Local Coding AI Instead video thumbnail ►

Don't Use Claude, Use This 320B Local Coding AI Instead

The Stack

Build Your Own Agentic Harness in Python - Full Tutorial video thumbnail ►

Build Your Own Agentic Harness in Python - Full Tutorial

Tech With Tim

Recursive's $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push EP 299 video thumbnail ►

Recursive's $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push EP 299

Peter H. Diamandis

12GB Model, 8 Hours, One 3D Game: OrcaSAQ2 27B Tested video thumbnail ►

12GB Model, 8 Hours, One 3D Game: OrcaSAQ2 27B Tested

Fahd Mirza

Hugging Face agents course video thumbnail ►

Hugging Face agents course

Hugging Face

The last six months in LLMs in five minutes video thumbnail ►

The last six months in LLMs in five minutes

Simon Willison

Educators and sources

Educator / source Best for Skills Start with
Developers, technical generalists LLM tools, Prompting, AI safety, Local models, Model selection Read the recent model-roundup posts, then try the llm command-line tool with two or three different providers.
Researchers, advanced builders NLP research, Open models, Multilingual AI Browse research posts and community programs.
Developers fine-tuning and deploying models Open models, Fine-tuning, Deployment, Transformers Pick one fine-tuning or inference guide and reproduce it end to end.
Developers learning Transformer applications Transformers, NLP, Open models, Fine-tuning Use the book notebooks alongside the Hugging Face course.

Resources

How to Cost Your AI-Powered Filters

Technical article · FS Data Lab · Advanced

FS Data Lab published this guide on October 1, 2026. Use it when you want a rigorous way to estimate GPU latency and inference cost for AI-powered SQL filters, compare filter orderings, and understand how compute, memory bandwidth, selectivity, batching, and KV-cache reuse affect a query plan.

NVIDIA Kumo Tabular explained

Tabular foundation model explainer · NVIDIA · Intermediate to advanced

NVIDIA released Kumo Tabular on September 29, 2026. Use this guide to understand its in-context prediction workflow, reproduce a baseline, and test its vendor-reported results on your own tables.

Gemini 3.8 Flash

Model docs · Google AI for Developers · Intermediate

You need Google's official September 2, 2026 GA reference for its newest Flash model, including the stable model ID, 1M-token input window, thinking levels, multimodal inputs, tools, and production inference options.

OpenAI external models

Guide · OpenAI · Intermediate

You want to learn how OpenAI handles access to non-OpenAI model families before designing a mixed-provider or routed workflow.

Llama 4 prompt formats

Model docs · Meta Llama · Beginner to advanced

You need the current official Llama 4 prompt formats and model-card details before choosing Maverick or Scout for hosted or local deployment.

Mistral models overview

Model docs · Mistral AI · Beginner to advanced

You need to compare current Mistral families such as Devstral, Magistral, Voxtral, OCR, and general-purpose models.

Qwen API platform

API docs · Qwen · Beginner to advanced

You need official Qwen model-family context, deployment docs, and quickstarts before choosing a hosted or local workflow.

Together AI model catalog

Model catalog · Together AI · Beginner to advanced

You need to browse hosted open and proprietary models by provider and capability.

Together AI serverless models

Model docs · Together AI · Intermediate

You need to learn how serverless hosted model inference works before deploying an app.

Llama Cookbook

GitHub repo · Meta Llama · Beginner to advanced

You want Meta's practical recipes for inference, fine-tuning, RAG, and end-to-end Llama applications.

Llama API models

Model docs · Meta Llama · Beginner to advanced

You need Meta's current model-card and prompt-format index before choosing which Llama family, size, or integration path to test.

Mistral model selection guide

Guide · Mistral AI · Beginner to advanced

You want Mistral's official comparison of model families, pricing, context, and licensing before implementation.

Qwen quickstart

Quickstart · Qwen · Beginner to intermediate

You want the fastest official route into running Qwen3 with Hugging Face, vLLM, or SGLang.

Qwen3-Coder

Model launch guide · Qwen · Intermediate to advanced

You want the official Qwen explanation of where Qwen3-Coder fits for repo work, agentic coding, and open-model software engineering.

A Local-First Agent for Private and Cost-Effective Knowledge Work

Engineering study · Perplexity · Intermediate to advanced

You want a practical primary-source study of co-designing a compact local model and agent harness, including on-demand skills, context compaction, sandboxing, verification, and optional frontier-model escalation.

Hugging Face model hub

Model catalog · Hugging Face · Beginner to advanced

You need to discover, compare, and run open model checkpoints, datasets, and demos.