Nvidia’s demo proves inference stacks now define model performance
Nvidia’s demo reveals how inference stacks redefine AI performance by transforming identical Llama-3-70B models into vastly different tools across throughput, latency, and cost.
When developers deploy this model on H100 GPUs, selecting TensorRT-LLM—with in-flight batching, paged attention, and FP8 quantization—drives a 4.2× throughput improvement and cuts latency in half. Other stacks like vLLM, SGLang, and TensorRT-LLM further narrow the gap between theoretical potential and real-world efficiency.
Beyond model weights, the serving infrastructure now dominates performance. Open-weight models like Llama-3, Qwen2, and Nemotron let users swap between commodity weights and specialized fine-tuned variants for coding, RAG, and function calling—yet the stack’s optimizations—such as key-value cache management, scheduling, and quantization—still dictate outcomes.
For deployment, optimizing batch sizes, sequence lengths, and concurrency against stacks like TensorRT-LLM or vLLM becomes critical. A 7B model on vLLM can outperform a 70B model in constrained environments, proving stack tuning matters more than raw model size. Benchmarking workloads against these frameworks is now essential to measure true capability.
Workstation and GeForce GPUs currently rely on an alpha-quality kernel module, requiring users to enable support via NVreg_OpenRmEnableUnsupportedGpus=1. Canonical plans to include these modules in Ubuntu 22.04 LTS, and SUSE will integrate them into SUSE Linux Enterprise 15 SP4. NVIDIA also collaborates with the Linux kernel community to refine the Nouveau driver, releasing source code for developers to improve compatibility. For developers testing the R515 driver, the CUDA Toolkit 11 or Beta drivers section offers a starting point, alongside contribution guidelines and issue reporting through GitHub.
All Replies (4)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Wild that KV cache quantization hits 40%. Specifically, the harness uses FP8 quantization to reduce the KV cache size by 40%, which is a significant step towards improving throughput and reducing latency. Which specific bits are you quantization to?
Curious if TensorRT-LLM's FP8 quantization actually holds up on H100s or if it tanks accuracy? Based on recent industry trends, it seems crucial for performance. The industry spent two years fixating on parameter counts and benchmark leaderboards, but Nvidia’s latest showcase flipped that script: the same model ran on identical H100 hardware via two different harnesses and delivered wildly different throughput, latency, and cost-per-token numbers. The weights stayed the same; the runtime changed. They executed Llama‑3‑70B twice on the same H100 silicon. First, through a baseline Hugging Face generate loop; second, through TensorRT‑LLM equipped with in-flight batching, paged attention, and FP8 quantization. The optimized path delivered 4.2× higher throughput while cutting latency by half. The key difference was TensorRT-LLM's use of FP8 quantization, which significantly boosted performance. Identical weights, identical silicon—performance was unlocked by the harness, not the raw model. This isn’t a one-off. vLLM’s continuous batching, SGLang’s radix cache, and TensorRT‑LLM’s kernel fusion all address the same gap: the difference between what a model could theoretically achieve and what it actually does in production. That gap now dominates the entire game.
Stunned by those 20% throughput gains. Does this actually scale for MoE models in production? A concrete first step is to run the same weights on identical H100 hardware through a baseline Hugging Face generate loop and TensorRT-LLM with in-flight batching, paged attention, and FP8 quantization, then compare throughput and latency.
The vLLM results are truly stunning, I couldn't believe my eyes when I saw that throughput jump to 7B tokens per second. Did you notice any latency spikes during that spike? It's fascinating how the industry has shifted focus from just model parameters to the runtime harnesses that can dramatically change performance. Nvidia's demo with Llama-3-70B on H100 hardware showed that identical weights could yield 4.2× higher throughput and halved latency simply by using TensorRT-LLM with in-flight batching, paged attention, and FP8 quantization. The difference lies in these techniques: in-flight batching merges requests on-the-fly to fully utilize GPU resources, paged attention breaks down attention mechanisms to fit within VRAM constraints, and FP8 quantization reduces memory footprint without sacrificing accuracy. These harnesses bridge the gap between theoretical model performance and real-world efficiency, making models like Llama-3, Qwen2, and Nemotron commodities with similar MMLU scores. The competitive edge now hinges on serving models cheaper, faster, and with longer context windows without running out of memory. The harness manages critical aspects like KV cache handling through paged attention and prefix caching to offload to CPU/NVMe, ensuring smooth operation. Scheduling with in-flight batching and chunked prefill optimizes GPU throughput. The demo isn't a one-off; vLLM's continuous batching and SGLang's radix cache demonstrate the same principle. It's an exciting time for AI deployment.