All AI videos

AI video

CMU AI Agents 2026: 10. Deep Research

CMU assistant professor Akari Asai surveys deep-research agent architecture, evaluation, retrieval, training, and citation quality in a Fall 2026 AI Agents lecture.

Graham Neubig · 2026 featured video

Watch on YouTube

What this video shows

Asai contrasts single-hop retrieval with research tasks that require decomposition, repeated search, source reading, revision, synthesis, and citations. She surveys benchmarks for short factual answers, long reports, expert domains, and citation support, then connects their design to search systems and agent training.

The lecture gives a research map rather than one production recipe. It also warns that an agent can retrieve benchmark answers or original datasets from the web, which invalidates a score unless evaluators block those sources or use a controlled corpus. Learnetto recommends scoring claim support, source quality, completeness, and reproducibility separately.

Review the official CMU AI Agents course page for the course scope and instructors. Compare the lecture with the Deep Research Bench paper and its frozen RetroSearch evaluation environment.

What you will learn

  • Complex research questions require an agent to revise queries and gather evidence across several sources.
  • Answer quality and citation support are separate dimensions because fluent reports can misrepresent evidence.
  • Expert-domain benchmarks need qualified annotators and defensible answer criteria.
  • Evaluators must prevent agents from retrieving benchmark labels or reference reports during the test.

How to apply this safely

  1. Define a research task with explicit scope, recency, source, and citation requirements.
  2. Log every query, visited source, extracted claim, and final citation so a reviewer can trace the report.
  3. Score factual correctness, evidence support, source quality, completeness, cost, and latency separately.
  4. Run leakage checks against benchmark repositories and repeat tests in a frozen or recorded search environment when comparing versions.

Important limitations

  • The lecture surveys active research, so benchmark coverage and model results can become dated as systems and datasets change.
  • Automated citation judges can disagree with experts, while a frozen corpus improves repeatability at the cost of current-web realism.

Sources to check

Continue learning on Learnetto