►
Trends in AI Engineering: Careers, Coding Agents & Evals | Hugo Bowne-Anderson & Alexey Grigorev
Alexey Grigorev · 2026, AI engineering careers, coding agents, evals
AI directory search
Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.
98 matches for "CI"
Watch first when you want a fast feel for the topic before opening courses, docs, or profiles.
►
Alexey Grigorev · 2026, AI engineering careers, coding agents, evals
►
AI Engineer · 2026, GitHub Copilot, custom agents, GitHub Actions
►
Graham Neubig · 2026, deep research, search agents, retrieval
►
AI Engineer · 2026, reinforcement learning, small models, financial agents
►
OpenAI · 2026, Codex, production monitoring, Grafana
►
AI Engineer · 2026, agent evaluation, tracing, LLM as judge
►
OpenAI · 2026, code review, coding agents, software governance
►
Kurzgesagt – In a Nutshell · 2026, ai safety, agent security, reward hacking
►
The Stack · 2026, iquest-q1, local models, coding agents
Initial Commit · Beginner to intermediate
Practical AI use for small businesses, product decisions, and repeatable operations.
Skills
Founder workflows, Product ops, Automation
DeepLearning.AI Short Courses · Beginner to advanced
Structured, practical courses from prompt engineering through agentic workflows.
Skills
Prompting, Agents, RAG, ML foundations
fast.ai · Beginner to intermediate
Good practical teaching on making ML more understandable and useful.
Skills
Practical ML, Ethics, Education
Cohere For AI · Advanced
Useful for people who want research context and open science around language models.
Skills
NLP research, Open models, Multilingual AI
AI Automation Society · Beginner to intermediate
Large Skool community focused on taking people from watching AI videos to building practical automations and repeatable operating systems.
Skills
AI automation, n8n, Business workflows, AI operating systems
AI Automation Agency Hub · Beginner to intermediate
Very large Skool community for learning how to package and sell AI automation services to businesses.
Skills
AI agencies, Client acquisition, Automation offers, Business systems
ChatGPT Users · Beginner to intermediate
Beginner-friendly Skool community around practical ChatGPT use for business, content, sales, and everyday client work.
Skills
ChatGPT, Business prompts, Decision support, Client workflows
The AI Advantage · Beginner to intermediate
Brings AI education to a broad entrepreneurial audience that may not identify as technical.
Skills
AI productivity, Business growth, Mindset, Prompting
AI Automation Mastermind · Beginner to intermediate
A small-business oriented community for learning AI and automation as operational leverage.
Skills
AI automation, Operations, Efficiency, Small business workflows
AI Automation Circle · Beginner to intermediate
Structured Skool community built around live feedback and support for people learning to build automations.
Skills
AI automation, Live hotseats, Client systems, Workflow support
AI Coaching Advantage · Beginner to intermediate
Focused on helping coaches and creators move beyond prompt tips into AI-assisted client and business operations.
Skills
AI coaching workflows, Client operations, Files, Email drafts, Business organization
AI Creator Academy by DCT · Beginner to intermediate
Explicitly non-tech AI training for everyday people who want creator, digital product, and income workflows.
Skills
AI creator skills, Digital products, AI art, Canva, Side hustles
AI Systemizers · Beginner to intermediate
Business-owner community for using AI, automation, and systems to increase revenue without adding operational drag.
Skills
Business systems, AI automation, Sales systems, Founder freedom
The AI Founder's Vault+ · Beginner to intermediate
Explicitly built for non-technical founders and creators who want tools, templates, and systems for building with AI.
Skills
Plug-and-play AI, Founder tools, Templates, CRM workflows
Stock Image AI Prompt Lab · Beginner to intermediate
Niche Skool group for learning commercial-ready AI stock image prompting and iteration.
Skills
AI image prompts, Stock images, Commercial creative workflows
The AI-Driven Business Summit · Beginner to intermediate
Summit community focused on expert sessions, tools, and AI revenue systems for women business owners.
Skills
AI business, Skool growth, Community monetization, AI tools
The Stoa of AI · Beginner to intermediate
Skool community explicitly for small business owners with no technical background who want to reduce costs and move faster with AI.
Skills
AI strategy, AI assistants, Business tools, Cost reduction
AI Content Creators · Beginner to intermediate
Creator community teaching prompts, tools, and tutorials for producing and monetizing AI content.
Skills
AI content, Images, Videos, Voice cloning, Monetization
The Active Income Network · Beginner to intermediate
Community around real entrepreneurs building businesses using AI and modern social tools.
Skills
AI business, Offer validation, Lead generation, Sales outreach
Nick Saraev · Beginner to intermediate
Business-oriented AI automation education for people looking to create offers and workflows without deep technical prerequisites.
Skills
AI agencies, Automation, No-code tools, Business workflows
Connor Cahill · Beginner to intermediate
Teaches AI automation and agency-building concepts for people turning AI workflows into client services.
Skills
AI agencies, Automation, Client offers, No-code systems
Michele Torti · Beginner to intermediate
Practical creator in the AI business and automation space for people learning income-oriented AI workflows.
Skills
AI business, Automation, Freelancing, Client workflows
Practical AI content and personal-brand workflows for non-technical creators and operators.
Skills
AI content, Personal brand workflows, Writing with AI, Social growth
Typeshare and digital writing · Beginner to intermediate
Good for learning how AI can support structured writing and content production without replacing judgment.
Skills
AI-assisted writing, Content systems, Personal brand, Idea development
Anthropic Academy and Claude docs · Beginner to advanced
Official material for learning Anthropic's current Claude family, model tradeoffs, Claude Code, MCP, computer use, practical prompt workflows, and team model-governance workflows without relying on third-party summaries.
Topics
Claude models, Claude Code, MCP, Computer use, AI fluency, Frontier model selection, Model governance, Release notes
A fast, practical way to build vocabulary and intuition before going deeper into LLMs or AI engineering.
Topics
ML foundations, Classification, Embeddings, Neural networks
TypeSafe AI and Jev documentation · Intermediate to advanced
Official material for using Jev where software needs a bounded judgment, probability, or score rather than generated prose.
Topics
Jev, System One models, Typed decisions, Calibrated probabilities, Choice, Score, Noul, RLCD
OpenAI docs, Academy, and Cookbook · Beginner to advanced
Official model and implementation material for learning GPT-6 Astra and cost-sensitive GPT-5.6 choices, Codex workflows, subagents, memories, agent evals, MCP and connector patterns, retrieval, background jobs, prompt engineering, production best practices, model optimization, structured outputs, and OpenAI's Academy learning path.
Topics
GPT-6 Astra, GPT models, Reasoning models, Model selection, Agents, Subagents, RAG, Structured outputs, MCP, Evals, Memories
Kaggle Learn · Beginner
Useful for people who need the data and ML basics before working seriously with AI tools.
Topics
Python, Machine learning, Data preparation, Computer vision
Weaviate Academy · Beginner to intermediate
Good structured learning around vector databases, retrieval, and search relevance.
Topics
Vector search, RAG, Hybrid search, Embeddings
Useful for debugging and evaluating LLM applications once you move beyond prototypes.
Topics
Observability, Evals, Tracing, RAG debugging
Langfuse Docs · Intermediate
Good operational material for tracing, scoring, and improving production LLM apps.
Topics
Observability, Prompt management, Evals, Tracing
MLOps Community · Intermediate to advanced
Good for learning how practitioners actually ship and maintain ML/AI systems.
Topics
MLOps, Production ML, AI systems, Community learning
DeepLearning.AI · Beginner to advanced
A broad catalog of structured AI courses, including many short practical courses from tool creators.
Topics
Generative AI, Deep learning, Prompting, Agents
MarkTechPost · Intermediate
Useful as a discovery feed, but verify against papers and official repos before relying on it.
Topics
Research discovery, Tool discovery, AI news
CS50's Introduction to Artificial Intelligence with Python · Beginner to intermediate
A rigorous entry point into classical AI concepts, search, optimization, machine learning, neural networks, and language processing.
Topics
Search, Knowledge representation, ML foundations, AI foundations
Introduction to Artificial Intelligence · Beginner to intermediate
Good grounding in search, games, probability, and reinforcement learning before LLM-specific work.
Topics
Search, Planning, Reinforcement learning, AI foundations
IBM SkillsBuild AI · Beginner
Accessible AI literacy and business-facing AI learning for non-specialists.
Topics
AI foundations, Generative AI basics, Business use cases
Elements of AI · Beginner
A clear non-technical entry point for understanding AI concepts and social implications.
Topics
AI foundations, Ethics, Society, ML basics
Codecademy AI courses · Beginner
Good for beginners who want guided coding exercises while learning AI concepts.
Topics
AI foundations, Python, Prompting, LLM apps
Udacity School of AI · Beginner to advanced
Project-heavy AI learning for people who want structured programs and portfolio work.
Topics
Machine learning, Deep learning, AI product, Computer vision
Coursera AI courses · Beginner to advanced
Broad catalog for comparing university, company, and practitioner-led AI programs.
Topics
AI foundations, Machine learning, Generative AI, Business AI
edX artificial intelligence courses · Beginner to advanced
Useful for more academic AI courses and professional certificate programs.
Topics
AI foundations, Machine learning, Robotics, Ethics
ComfyUI examples · Intermediate
Useful for understanding node-based generative image workflows and reproducible pipelines.
Topics
Image generation, Workflow graphs, Stable Diffusion
OpenRouter docs · Beginner to intermediate
Useful for learning model comparison, latest-family aliases, routing, fallback behavior, agent construction, project-specific evals, and API-compatible experimentation across proprietary and open model families.
Topics
Model routing, Model comparison, Auto Router, Agent SDK, Coding-agent harnesses, GPT models, Claude models, Gemini, Llama, Mistral, DeepSeek, Qwen, API examples, Evaluation
Fireworks AI model catalog · Intermediate to advanced
Official material for comparing serverless and dedicated model serving, training specialized models, and measuring quality, token use, cost, and task duration on production-shaped evaluations.
Topics
Ember-1, Kimi K3, Serverless inference, Dedicated deployment, Model training, Model evaluation, Reasoning efficiency
Gemini API model docs · Beginner to advanced
Official Gemini material for learning Gemini 3.8 Flash and the current stable, preview, latest, and experimental model lineup, plus the now-GA Interactions API, Lyria music generation, prompt design, function calling, background execution, Deep Research, Computer Use, Hooks, Live API, File Search, coding-agent setup, multimodal tradeoffs, and AI Studio workflows.
Topics
Gemini 3.8 Flash, Gemini models, Multimodal AI, Long context, Model selection, AI Studio, Interactions API, Background execution, Deep Research, Computer Use, Hooks, Live API, Coding agents, Music generation, File Search, API examples
Meta Model API and Llama docs · Beginner to advanced
Official path into current Llama families, prompt formats, open-weight deployment, Meta's newer hosted Model API surface, and integration decisions for both local and managed workflows.
Topics
Llama, Open models, Meta Model API, Muse Spark, Muse Code, Local models, Fine-tuning, Model deployment, Safety models
Skill system
OpenClaw · Official docs
Use this to understand the OpenClaw skill format: markdown instruction files in directories with SKILL.md frontmatter and tool-use guidance.
Start
Read the official skills docs before installing community skills, then test one low-risk local skill in a disposable workspace.
AI worker platform
Mission Control AI · Preconfigured AI workers
Use this when you want role-specific AI workers with SOPs, integrations, and governance policies already built in.
Start
Map one operational process, identify data and approval boundaries, and evaluate whether a prebuilt worker fits before building custom agents.
AI coworker
Viktor · Slack and Teams AI employee
Use this when the adoption surface should be chat-native and team-facing rather than a separate automation builder.
Start
Pick one team workflow, define who approves outputs, and measure whether Viktor saves time without creating review burden.
Everyone
Learn first
Open next
Builders choosing between Claude, GPT, Gemini, Llama, Mistral, Cohere, DeepSeek, Qwen, Grok, Perplexity, and hosted open models
Learn first
Good matches
Open next
Short course · DeepLearning.AI · Intermediate
You need to test, trace, and improve agent workflows instead of judging only single LLM responses.
agent evals, evals, agents, reliability, tracing
Course · DeepLearning.AI · Beginner
You want a non-technical foundation for deciding where generative AI fits in a team or business.
ai leadership, ai strategy, business ai, ai adoption
Specialization · Duke University · Beginner to intermediate
You want a structured product-management route for scoping, evaluating, and shipping AI products.
ai product, product management, ai strategy, product quality
API docs · Perplexity · Intermediate
You want to compare search, Sonar, Agent API, and cited research workflows from the primary source.
ai search, deep research, grounded answers, sonar, research workflows
Open-source eval and red-team framework · OpenAI / Promptfoo · Intermediate to advanced
Declarative regression tests and side-by-side comparisons of prompts, models, RAG systems, and agent configurations, especially when red teaming is also required.
evals, llm evaluation, ai quality, open-source frameworks
Pytest-style LLM evaluation framework · Confident AI · Intermediate to advanced
Engineering teams that want LLM and agent evaluations to behave like software unit tests, with thresholds, assertions, and CI-friendly failures.
evals, llm evaluation, ai quality, open-source frameworks
Model and agent evaluation harness · UK AI Security Institute · Intermediate to advanced
Rigorous, reproducible model and agent capability or safety evaluations involving tools, multi-turn interaction, coding, or sandboxed environments.
evals, llm evaluation, ai quality, open-source frameworks
AI evaluation and observability platform · Braintrust Data · Intermediate to advanced
Teams that want a feedback loop from production traces and failures into datasets, regression experiments, CI gates, and continuous scoring.
evals, llm evaluation, ai quality, evaluation platforms
Agent evaluation and observability platform · LangChain · Intermediate to advanced
Agent teams, especially LangChain and LangGraph users, that need tracing, datasets, experiments, human review, and production feedback in one system.
evals, llm evaluation, ai quality, evaluation platforms
Open-source evals and observability platform · ClickHouse · Intermediate to advanced
Teams prioritizing open-source data control and one workflow across tracing, prompts, datasets, experiments, annotation, and evaluation.
evals, llm evaluation, ai quality, evaluation platforms
Open-source GenAI evaluation and monitoring · MLflow Project · Intermediate to advanced
Teams extending an existing MLflow or MLOps stack to LLM and agent tracing, evaluation-driven development, human feedback, and monitoring.
evals, llm evaluation, ai quality, evaluation platforms
AI reliability and evaluation platform · Patronus AI · Intermediate to advanced
Safety- and reliability-sensitive RAG or agent systems needing specialized hallucination, retrieval, PII, bias, policy, and custom-criteria evaluators.
evals, llm evaluation, ai quality, evaluation platforms
Evaluation methodology guide · Anthropic · Intermediate to advanced
Teams designing task-specific evaluation suites for Claude applications that need criteria, datasets, edge cases, grading patterns, and code examples.
evals, llm evaluation, ai quality, cloud and provider evals
Agent evaluation and optimization harness · Harbor Framework Team · Intermediate to advanced
Running coding and computer-use agents against reproducible, sandboxed task suites at local or cloud scale.
evals, llm evaluation, ai quality, benchmarks and learning resources
Language-model benchmark runner · EleutherAI · Intermediate to advanced
Reproducible few-shot and zero-shot evaluation of base or instruction-tuned language models on established academic benchmarks.
evals, llm evaluation, ai quality, benchmarks and learning resources
Software-engineering agent benchmark · SWE-bench team · Intermediate to advanced
Measuring whether coding agents can resolve real GitHub issues by producing repository patches that pass executable tests.
evals, llm evaluation, ai quality, benchmarks and learning resources
Terminal-agent benchmark · Harbor Framework and Laude Institute · Intermediate to advanced
Comparing agents on difficult, verifiable coding, systems, security, and scientific work inside sandboxed terminals.
evals, llm evaluation, ai quality, benchmarks and learning resources
Application-evaluation learning resource · Hamel Husain and Shreya Shankar · Intermediate to advanced
Practitioners building product-specific evals from traces, domain-expert judgments, and observed failures instead of relying on generic benchmarks.
evals, llm evaluation, ai quality, benchmarks and learning resources
Learning path · Google Cloud · Beginner
Use this four-activity learning path for an introductory overview of generative AI, large language models, and responsible AI principles.
generative ai, large language models, responsible ai, foundations, learning path
Responsible AI principles · Google · Beginner
Use this as Google's primary statement of the principles it applies to AI development and deployment, alongside the playbook's recommendation to build a framework for trust.
responsible ai, ai principles, ai governance, ai safety, trust
Course · Google Cloud · Beginner
Use this one-hour course when security and data-protection leaders need a framework for identifying AI-specific risks, protecting sensitive data, and applying Google's Secure AI Framework.
ai security, secure ai framework, data protection, compliance, risk management
Connected-finance safety guide · OpenAI · Beginner
OpenAI announced on October 2, 2026 that ChatGPT Finances is rolling out to Free and Go users in the United States. Use this guide to understand the connected-account workflow, check the source data behind answers, and review privacy controls before linking an account.
chatgpt, personal finance, connected accounts, plaid, privacy
Desktop agent safety guide · GitHub · Intermediate
GitHub released Copilot computer use in public preview on October 1, 2026. Use this guide to decide when visual desktop control is appropriate, enable it deliberately, and test its permission and stop controls before real work.
github copilot, computer use, desktop agents, copilot cli, gui automation
Model migration explainer · Anthropic · Intermediate to advanced
Anthropic released Claude Sonnet 5.5 on September 28, 2026. Use this guide to understand the efficiency claims, breaking API changes, and a safe migration test from Sonnet 5.
claude sonnet 5.5, claude code, coding agents, model migration, adaptive thinking
Reasoning-efficient coding model explainer · Fireworks AI · Intermediate to advanced
Fireworks released Ember-1 on September 23, 2026 as a Kimi K3-based model trained to use fewer reasoning tokens. Use this explainer to evaluate the quality, cost, latency, and context-growth claims on your own agent workload.
ember-1, kimi k3, reasoning models, reasoning tokens, cost optimization
AI-assisted science explainer · Anthropic · Intermediate
Anthropic reported ART on September 23, 2026. Use this explainer to separate what its agents found, what the lab confirmed, and what remains a hypothesis.
claude, ai agents, ai for science, biology, genome mining
System One Model launch · TypeSafe AI · Intermediate to advanced
Jev launched on September 15, 2026 as TypeSafe AI's first System One model: a non-chat model that turns application state into typed decisions and probabilities for software workflows.
jev, system one models, rlcd, structured decisions, calibrated probabilities
Managed agent API launch and guide · OpenAI · Intermediate to advanced
You want OpenAI's September 10, 2026 launch guide for building long-running cloud agents with the managed Codex harness, context compaction, tool search, multi-agent delegation, and your choice of hosted or self-managed sandbox.
openai, agents, codex, context compaction, multi-agent
Voice agent model launch · OpenAI · Intermediate
You want OpenAI's September 10, 2026 launch of GPT-Live-1 in the API for full-duplex voice agents, including its $0.05-per-minute front-end pricing and guidance on pairing it with a backend model and agent harness.
openai, voice agents, realtime, api, model selection
AI safety incident assessment · Anthropic · Advanced
You want Anthropic's September 9, 2026 assessment of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations, including the model behaviors and evaluation-design failures involved.
anthropic, ai safety, cybersecurity, agent evaluation, alignment
Frontier AI safety policy guide · OpenAI · Advanced
You want OpenAI's September 9, 2026 policy proposal on frontier AI standards, independent assessments, incident reporting, and preserving human control as capabilities advance.
openai, ai safety, frontier models, evaluations, governance
Open security benchmark · Hugging Face Community · Intermediate to advanced
You want a reproducible September 5, 2026 benchmark for comparing how agentic models handle indirect prompt injection, with public data, a public harness, control runs, tool-call traces, and outcome metrics tied to unauthorized payment actions.
agents, prompt injection, agent security, evals, tool use
Coding agent evaluation guide · GitHub · Intermediate to advanced
You want GitHub's September 2, 2026 evidence for measuring coding-agent efficiency across the whole task, including selective output compression, preserving useful context, benchmark regressions, and controlled production experiments.
github copilot, coding agents, context engineering, evals, cost optimization
Short course · DeepLearning.AI · Beginner
You want a concise, hands-on course for replacing vague vibe-coding prompts with project constitutions, feature specs, plan-implement-verify loops, legacy-code workflows, and a portable custom agent skill.
coding agents, spec-driven development, agent skills, planning, verification
Agent infrastructure guide · Anthropic · Intermediate to advanced
You want Anthropic's September 3, 2026 workflow for declaring agents, environments, skills, memory stores, and scheduled deployments in a repository, previewing changes, and applying them reproducibly in CI.
anthropic, ant cli, managed agents, infrastructure as code, skills
Open-source guide and repo · Hugging Face · Intermediate
You want a September 3, 2026 walkthrough of funes, an open-source local memory layer that indexes agent traces, preserves provenance, and lets Claude Code, Codex, pi, and Hermes recall decisions across sessions and machines.
hugging face, coding agents, agent memory, codex, claude code