All AI videos

AI video

From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model

Amazon AGI engineer Tarun Sunkaraneni profiles a multimodal training input pipeline and reports how staged data-system changes raised GPU utilization from 15% to 90% in his workload.

AI Engineer · 2026 featured video

Watch on YouTube

What this video shows

Sunkaraneni starts with a Qwen3-VL-style training workload whose images come from S3. He reports that the baseline spends about 85% of its loop waiting for loading, decoding, resizing, and transfer, leaving GPU utilization near 15%. He then adds concurrent workers and prefetching so the trainer receives prepared batches instead of requesting each batch synchronously.

The later changes reduce copies by passing Ray object references and spread workers when one machine network interface becomes the next constraint. The reported 15% to 90% utilization change belongs to this workload and hardware setup. Learnetto recommends repeating the timing breakdown on your own pipeline before adopting Ray or changing cluster size.

Read the Ray object documentation before using object references to reduce repeated transfers. Compare the worker design with Python asyncio task documentation so you can separate asynchronous I/O from CPU-bound work.

What you will learn

  • Measure data wait time separately from model compute before tuning kernels or buying more accelerators.
  • Concurrency helps independent loads and transforms overlap, while prefetching prepares future batches before the trainer requests them.
  • Shared object references can reduce serialization and copying, but object-store pressure and spilling still need monitoring.
  • A faster local pipeline can expose a network bottleneck at cluster scale, so repeat profiling after each material change.

How to apply this safely

  1. Record loading, decode, transform, transfer, forward, and backward time for a representative run.
  2. Change one stage first, then compare throughput, utilization, memory, and failed batches against the same sample.
  3. Add bounded prefetching and backpressure so workers cannot fill memory when training slows.
  4. Run the chosen design at production node count and inspect network saturation, object-store use, data correctness, and cost per completed sample.

Important limitations

  • The performance numbers come from the speaker demonstration and have not been independently reproduced by Learnetto. Storage layout, image sizes, transforms, hardware, and network topology can change the result.
  • Ray adds scheduling and operational overhead. A smaller workload may perform well with framework data loaders, local caching, or simpler multiprocessing.

Sources to check

  • Ray objects Official documentation for object references, argument passing, and object storage.
  • Python asyncio tasks Python documentation for scheduling and coordinating asynchronous work.

Continue learning on Learnetto

AI evals guide

Keep data and output checks fixed while you tune system performance.