All AI videos

AI video

I Built a Team of Self-Improving AI Agents (opus 5.5)

A practical walkthrough of scheduled Claude Code agents that create an artifact, grade it against a rubric, record feedback, and use that feedback in the next run.

Duncan Rogoff | Learn Claude Code · 2026 featured video

Watch on YouTube

What this video shows

Duncan Rogoff demonstrates scheduled Claude Code sessions for web pages, short videos, long-form video, and newsletter drafts. Each session creates or revises an artifact, scores it against written criteria, and records feedback for a later run. For visual work, his agent captures screenshots and compares them with examples before suggesting the next change.

The useful pattern comes from controlled iteration rather than autonomous model training. Karpathy's autoresearch repository changes one training file, runs a fixed five-minute experiment, and keeps a change only when a numeric validation metric improves. Rogoff applies a looser version to subjective work, where an LLM grader and feedback document replace the numeric metric. That makes a stable rubric, a fixed test set, versioned outputs, and human review necessary.

Compare the workflow with Karpathy's autoresearch repository, which uses a fixed time budget and validation metric. Use Anthropic's guide to define success criteria and build evaluations before trusting an LLM grader.

What you will learn

  • The agent changes files and instructions between runs, but the underlying model does not retrain itself in this workflow.
  • Changing one variable at a time makes it easier to connect a score change to the edit that produced it.
  • A fixed test set prevents the agent from choosing only examples that suit its latest rule.
  • An LLM grader can handle visual or editorial criteria, but you should test its agreement with human reviewers before using its scores as a gate.

How to apply this safely

  1. Choose one low-risk artifact, define a baseline, and write acceptance criteria that a reviewer can apply consistently.
  2. Keep the input examples and scoring rubric fixed while the agent changes one instruction or implementation detail per run.
  3. Save each artifact, score, change, and grader explanation so you can reproduce improvements and reject regressions.
  4. Require human approval before publishing, merging code, sending messages, or changing the rubric that decides whether a run succeeded.

Important limitations

  • The creator shows his workflow and selected outputs rather than a blinded comparison across repeated runs, so the video does not establish that the system improves reliably over time.
  • Subjective LLM scores can drift or reward superficial traits. A second grader does not replace a representative test set or periodic human calibration.
  • Scheduled agents need bounded permissions, time limits, budget limits, audit logs, and a safe failure path before they work unattended.

Sources to check

Continue learning on Learnetto

AI evals guide

Compare evaluation methods and tools for production AI systems.