Building and Scaling GenAI Workloads with Amazon EKS

This event has limited capacity. Register today to secure your spot.

Hands-on workshop Level 400

Upcoming Sessions

No upcoming sessions — check back soon

Share this workshop

About the event

Two hours. Real NVIDIA GPUs. A working GenAI stack you'll actually want to redeploy.

You'll deploy a large language model on Amazon EKS, push it until it breaks, then optimize it — using the same patterns AWS customers run in production today.

This is not a lecture. You'll be hands-on-keys the entire time, in a live AWS environment we provision for you.

What you'll build:

  • Deploy and serve an LLM on GPU-accelerated EKS nodes using vLLM
  • Scale inference across GPUs with Ray for bursty, multi-user traffic
  • Wire up real-time monitoring with Prometheus and Grafana — tokens/sec, latency, GPU utilization
  • Benchmark under load, find where it breaks, then fix it with data
  • Apply KV cache offloading to cut response time up to 3.6×

What you leave with:

  • A working end-to-end inference stack
  • The Terraform to redeploy it in your own account
  • The optimization playbook to defend your GPU spend

Led by AWS experts who've helped teams ship GenAI at scale.