All AI videos

AI video

Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab

Amazon AGI Lab program lead Felipe Blanes explains how the Nova Act team updates evaluations from customer use instead of relying only on static pre-release benchmarks.

AI Engineer · 2026 featured video

Watch on YouTube

What this video shows

Blanes describes an evaluation loop used while Nova Act moved from research preview to an AWS service. Teams define success with the customer, collect telemetry and interview evidence, classify failures into model, engineering, or product gaps, then feed the highest-value cases into decisions and future evaluations.

The trust-cliff section says customers in this program treated roughly 80% reliability as added monitoring work and reacted differently near 90%. Learnetto treats those percentages as speaker-reported observations, not general thresholds. The durable method is to measure the real workflow, publish known limits, and keep a reviewed regression set as customer behavior changes.

Read the official Nova Act user guide for the service scope behind the examples. Use OpenAI's evaluation best-practices guide as an adjacent method for task-specific datasets and continuous evaluation.

What you will learn

  • Production traces and customer interviews reveal tasks and failure costs that synthetic test authors may miss.
  • Classifying each gap as model, harness, or product work helps the right team address it.
  • Known limits should be visible to users because an aggregate score does not describe unsupported workflows.
  • Regression sets need new reviewed cases as customers adopt more complex tasks and the product changes.

How to apply this safely

  1. Ask users to define a successful outcome, acceptable intervention rate, and failure cost for one workflow.
  2. Collect consented traces, explicit feedback, support cases, and interview notes, then remove sensitive data before reuse.
  3. Label each failure by task, severity, root cause, and owning layer, and add representative cases to a versioned test set.
  4. Publish the supported scope and rerun the test set for model, prompt, tool, and interface changes.

Important limitations

  • The reliability percentages and customer examples come from an Amazon speaker and are not presented with sample sizes, distributions, or an independent study.
  • Production-derived evaluations can overrepresent frequent users and observed failures. Teams still need safety cases, edge cases, privacy review, and held-out tests.

Sources to check

Continue learning on Learnetto

AI evals guide

Choose deterministic, model-based, and production checks.