What this video shows
Asai contrasts single-hop retrieval with research tasks that require decomposition, repeated search, source reading, revision, synthesis, and citations. She surveys benchmarks for short factual answers, long reports, expert domains, and citation support, then connects their design to search systems and agent training.
The lecture gives a research map rather than one production recipe. It also warns that an agent can retrieve benchmark answers or original datasets from the web, which invalidates a score unless evaluators block those sources or use a controlled corpus. Learnetto recommends scoring claim support, source quality, completeness, and reproducibility separately.
Review the official CMU AI Agents course page for the course scope and instructors. Compare the lecture with the Deep Research Bench paper and its frozen RetroSearch evaluation environment.
What you will learn
- Complex research questions require an agent to revise queries and gather evidence across several sources.
- Answer quality and citation support are separate dimensions because fluent reports can misrepresent evidence.
- Expert-domain benchmarks need qualified annotators and defensible answer criteria.
- Evaluators must prevent agents from retrieving benchmark labels or reference reports during the test.
How to apply this safely
- Define a research task with explicit scope, recency, source, and citation requirements.
- Log every query, visited source, extracted claim, and final citation so a reviewer can trace the report.
- Score factual correctness, evidence support, source quality, completeness, cost, and latency separately.
- Run leakage checks against benchmark repositories and repeat tests in a frozen or recorded search environment when comparing versions.
Important limitations
- The lecture surveys active research, so benchmark coverage and model results can become dated as systems and datasets change.
- Automated citation judges can disagree with experts, while a frozen corpus improves repeatability at the cost of current-web realism.
Sources to check
- CMU 11-768 AI Agents Official Fall 2026 course page for the lecture.
- Deep Research Bench Primary paper for a controlled evaluation of open-web research agents.
Continue learning on Learnetto
Best deep-research courses
Compare research workflows, source evaluation, and synthesis lessons.
AI evals guide
Build datasets and graders for research quality.
AI research delegation lab
Model dependencies, parallel work, cost, and clean-run probability.