What this video shows
Blanes describes an evaluation loop used while Nova Act moved from research preview to an AWS service. Teams define success with the customer, collect telemetry and interview evidence, classify failures into model, engineering, or product gaps, then feed the highest-value cases into decisions and future evaluations.
The trust-cliff section says customers in this program treated roughly 80% reliability as added monitoring work and reacted differently near 90%. Learnetto treats those percentages as speaker-reported observations, not general thresholds. The durable method is to measure the real workflow, publish known limits, and keep a reviewed regression set as customer behavior changes.
Read the official Nova Act user guide for the service scope behind the examples. Use OpenAI's evaluation best-practices guide as an adjacent method for task-specific datasets and continuous evaluation.
What you will learn
- Production traces and customer interviews reveal tasks and failure costs that synthetic test authors may miss.
- Classifying each gap as model, harness, or product work helps the right team address it.
- Known limits should be visible to users because an aggregate score does not describe unsupported workflows.
- Regression sets need new reviewed cases as customers adopt more complex tasks and the product changes.
How to apply this safely
- Ask users to define a successful outcome, acceptable intervention rate, and failure cost for one workflow.
- Collect consented traces, explicit feedback, support cases, and interview notes, then remove sensitive data before reuse.
- Label each failure by task, severity, root cause, and owning layer, and add representative cases to a versioned test set.
- Publish the supported scope and rerun the test set for model, prompt, tool, and interface changes.
Important limitations
- The reliability percentages and customer examples come from an Amazon speaker and are not presented with sample sizes, distributions, or an independent study.
- Production-derived evaluations can overrepresent frequent users and observed failures. Teams still need safety cases, edge cases, privacy review, and held-out tests.
Sources to check
- What is Amazon Nova Act? Official service documentation for the browser-agent product discussed in the talk.
- Evaluation best practices An adjacent primary-source guide to task-specific and continuous evaluation.
Continue learning on Learnetto
Best AI agent evaluation courses
Build practical datasets, graders, traces, and review loops.
AI evals guide
Choose deterministic, model-based, and production checks.
Best AI agent courses
Connect evaluation work with tool use and workflow design.