Validating Distributed LLM Serving Benchmarks with NVIDIA srt-slurm, SLURM Recipes, Parameter Sweeps, and Pareto Analysis
In this tutorial, we discover NVIDIA’s srt-slurm framework and find out how we use srtctl to transform declarative YAML configurations into reproducible SLURM benchmark workflows for distributed LLM serving. We arrange the undertaking in Google Colab, examine its inside structure, outline a cluster configuration, dry-run built-in and customized recipes, and mannequin a disaggregated prefill-and-decode deployment…
