跳到正文
原文
NVIDIA Technical Blog:Agentic AI / Generative AI·· 1 天前AI 评分43

NVIDIA Dynamo-Triton 部署 HSTU 生成式推荐模型,延迟最高降低 5.93 倍

Deploying an HSTU Generative Recommender with NVIDIA Dynamo-Triton

AI 导读

NVIDIA Dynamo-Triton(原 Triton Inference Server)现已通过 recsys-examples 仓库支持端到端 HSTU 生成式推荐推理工作流,结合 PyTorch AOTI、FlexKV KV 缓存与 NV embedding cache。

正文

Generative recommender (GR) systems are emerging as a powerful new direction for large-scale personalization. Instead of treating recommendation as a set of isolated retrieval, ranking, and prediction stages, GRs reformulate recommendation as sequence modeling over user behavior. A user’s interactions, context, candidate items, and actions become tokens in a high-cardinality event stream, and the model learns to generate or score the next relevant items from that sequence.

This approach is especially attractive for modern recommendation workloads, where user histories can be long, item catalogs are constantly changing, and personalization quality depends on modeling rich sequential behavior. But it also introduces a serving challenge: GR models need low-latency inference despite long histories, large embedding tables, and sequence-heavy model architectures.

NVIDIA Dynamo-Triton (formerly NVIDIA Triton Inference Server) now supports an end-to-end Hierarchical Sequential Transduction Unit (HSTU) GR inference workflow through the NVIDIA recsys-examples repository. The workflow combines HTSUs, PyTorch Ahead-of-Time Inductor compilation, FlexKV-backed KV caching, native C++ validation, NV embedding cache, and Dynamo-Triton deployment.

The result is a practical path for serving HSTU ranking models with strong latency performance. At dynamic batch size 8 on an NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPU, Dynamo-Triton with PyTorch AOTI achieved best-case speedups of up to 4.47x for the three-layer HSTU model and 5.93x for the eight-layer model under a 100% GPU KV-cache hit rate, relative to the same AOTI configuration without KV caching.

This post shows you how to move an HSTU generative recommender from PyTorch development to production inference with NVIDIA Dynamo-Triton, PyTorch AOTI, and FlexKV. You will learn how to export and ahead-of-time compile the model, validate the resulting deployment artifact in Python and native C++, and serve it through Dynamo-Triton without rewriting the model for a separate runtime.

It also examines how GPU-backed KV caching reduces repeated computation and presents benchmarks demonstrating up to 5.93x lower latency, highlighting the practical performance benefits of this deployment workflow.

Why use HSTUs for generative recommendation?

HSTUs were introduced for GR workloads that operate over high-cardinality, nonstationary event streams. In a traditional recommender system, retrieval and ranking are often built from a collection of specialized models and feature pipelines. GRs instead model recommendation as a sequential prediction problem, allowing the model to reason over user context, item history, action history, and candidate items in one sequence-aware architecture.

In the NVIDIA HSTU ranking example, the model input is built from categorical tokens. Contextual tokens represent user-side information, item tokens represent items, and optional action tokens represent user interactions with those items.

Comparison diagram of Traditional Deep Learning Recommendation Model versus Generative RecSys HSTU workflow.
Figure 1. Traditional deep learning recommendation model (left) versus generative HSTU workflow (right)

The HSTU preprocessing path retrieves embeddings, interleaves item and action embeddings when action tokens are present, appends contextual information, and applies positional encoding. HSTU blocks then process the sequence, and a prediction head produces multitask ranking outputs.

This structure is well-suited for recommendation systems where recency, order, and repeated interaction patterns matter. However, it also means inference can become expensive when each request repeatedly processes long historical sequences. Production systems need to preserve HSTU modeling benefits while reducing redundant computation during serving.

Why is serving large sequential recommenders challenging?

Serving large sequential recommenders is different from serving a small dense ranking model. The serving stack must handle jagged sequence inputs, large categorical embedding state, long histories, and request patterns where the same user may return repeatedly with only a small amount of new information. Recomputing the full key-value state for a user history on every request wastes work and increases latency.

This is where KV caching becomes important. A KV cache stores reusable key-value data from prior sequence computation, allowing the model to avoid recomputing cached portions of the user history. For recommender inference, this is particularly useful when a user’s long-term history remains mostly stable while new candidate items or recent actions arrive.

Diagram illustrating HSTU serving with sequential input history and candidates, and masked attention mechanism.
Figure 2. HSTU serving

The NVIDIA HSTU inference workflow includes a KVCacheManager that uses GPU memory and host storage for KV data caches. The GPU cache is organized as a paged KV data table and supports lookup, allocation, append, and eviction. When GPU cache space is constrained, older users can be evicted according to an LRU-style policy. Host-side storage provides another tier for cached KV data, and the workflow includes a FlexKV-backed backend for the KV-cache runtime.

The HSTU attention kernel can consume KV data from the paged cache, and the exported inference path includes cache-aware custom operations for lookup, allocation, onboarding, appending, and offloading. This allows the serving path to preserve the model’s sequence semantics while reducing redundant computation.

PyTorch AOTI for native inference

The Pytorch AOTI (Ahead-of-Time Inductor) workflow starts from a PyTorch model and exports it using torch.export and PyTorch AOTI. AOTI compiles the model ahead of time into a package that can be loaded by a native C++ runtime. This reduces Python runtime overhead and provides a deployment-friendly artifact for the Dynamo-Triton PyTorch AOTI backend.

Flowchart of the PyTorch AOTI workflow from model export to native C++ runtime loading.
Figure 3. PyTorch AOTI workflow

The exported model package contains the AOTI model archive plus metadata and embedding table files. In the NVIDIA example, the embedding implementation combines DynamicEmb inference embedding tables and NV Embedding Cache that reduces GPU memory usage by storing only the popular embeddings in GPU memory while keeping the entire table in CPU memory. The export path writes layer metadata and embedding table data alongside the compiled .pt2 archive so the model can be loaded without unnecessary duplicate embedding table copies.

The workflow validates the same exported artifacts in multiple ways. Python export scripts generate the package and replay tensors. Native C++ executables load and replay the exported model for correctness and performance validation. The Dynamo-Triton deployment then uses the same AOTI package and replay path, which helps keep development validation and production serving aligned.

Dynamo-Triton deployment path

Dynamo-Triton provides the production serving layer for the exported HSTU model. The AOTI deployment uses the Dynamo-Triton PyTorch backend with platform: "torch_aoti". This allows Dynamo-Triton to load and serve the ahead-of-time compiled PyTorch model package.

The full workflow includes the following five stages:

  • Build the required custom operators and runtime libraries
  • Export the HSTU ranking model with PyTorch AOTI
  • Start the FlexKV-backed KV-cache service
  • Validate the exported artifacts with native C++ replay
  • Serve the exported KV-cache AOTI model with Dynamo-Triton

This approach is important because recommender serving requires more than just a fast model kernel. Dynamo-Triton brings model repository management, request handling, backend integration, metrics, and deployment structure. AOTI brings a lower-overhead compiled model artifact. NV Embedding Cache lowers GPU memory requirements by keeping only the hot parts of embedding tables in GPU memory. FlexKV reduces recomputation for long user histories by caching attention blocks. Together these form a serving stack aimed at realistic generative recommender inference.

Architecture diagram of the HSTU GenRec inferencing stack including Dynamo-Triton, NV Embedding Cache, and KV Cache Manager.
Figure 4. HSTU GR inferencing stack with Dynamo-Triton

Benchmarking HSTU serving latency

This benchmark compares HSTU serving latency across Dynamo-Triton backends, model sizes, batch sizes, and KV-cache states.

The goal is to quantify the performance benefits of the production HSTU serving stack. It compares the Dynamo-Triton PyTorch AOTI backend with the Python backend, measures the additional latency reduction from GPU KV caching, and evaluates how those benefits scale across model depths and batch sizes. Ultimately, it shows developers the performance they can expect when moving from uncached Python-based inference to compiled, cache-aware HSTU deployment with Dynamo-Triton.

The benchmark results in recsys-examples use the KuaiRand-1K ranking configuration on a single GPU. The model structure includes three-layer and eight-layer HSTU variants, hidden size 512, four attention heads, BF16 model weights, BF16 KV cache, maximum history sequence length of 8,192 in total (4,096 item plus action pairs) history stream, maximum candidate sequence length of 100, and six contextual features. The effective sequence length before alignment is 8,298 tokens, exported with a maximum aligned sequence length of 8,320.

The benchmark protocol reports latency per logical request. For the Dynamo-Triton AOTI benchmark, each Dynamo-Triton call contains one logical batch, and latency per logical request is calculated by dividing end-to-end pass time by the number of Dynamo-Triton calls times the logical batch size. Dataset loading, validation, rebatching, user-ID generation, server startup, warmup, and post-warmup sleep are excluded from the measured time.

The hardware used for the Dynamo-Triton backend comparison and batch-size results is an NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPU.

Dynamo-Triton backend comparison

At Dynamo-Triton batch size 2, PyTorch AOTI improves latency over the Dynamo-Triton Python backend even without KV-cache hits. With hits from the GPU KV cache of 20 GB, the latency improvement is significantly larger.

Bar chart comparing latency for Dynamo-Triton Python backend, PyTorch AOTI, and PyTorch AOTI with KV-cache hit for 3-layer and 8-layer HSTU models.
Figure 5. Dynamo-Triton backend comparisons for three-layer and eight-layer HSTU

These results show two different gains. First, AOTI reduces serving overhead compared with the Python backend. Second, KV cache hits reduce model work by reusing cached sequence state. The cache benefit is especially visible on the deeper eight-layer model, where avoiding recomputation has more impact.

Batch-size scaling with AOTI and KV cache

The PyTorch AOTI backend results by batch size show that KV caching becomes increasingly effective as logical batch size grows.

Line chart showing three-layer HSTU batch-size scaling for No cache versus GPU KV-cache-hit latency.
Figure 6. Three-layer HSTU, with Dynamo-Triton batch-size scaling
Line chart showing eight-layer HSTU batch-size scaling for No cache versus GPU KV-cache-hit latency.
Figure 7. Eight-layer HSTU, with Dynamo-Triton batch-size scaling

At batch size 8, the three-layer HSTU model reaches 0.423 ms latency per logical request with GPU KV-cache hits. The eight-layer HSTU model reaches 0.678 ms. These are strong results for long-sequence ranking inference and demonstrate the value of combining compiled model execution with cache-aware serving.

What are the benefits of accelerating HSTU GR inference?

Recommendation systems operate under tight latency budgets. Additional ranking latency can affect page-load times, feed responsiveness, and ad-serving deadlines. Meanwhile, increasingly sequence-aware and personalized models can require more inference compute, particularly as user histories grow.

The HSTU serving workflow combines complementary technologies to address this challenge. HSTU provides the generative recommendation architecture, PyTorch AOTInductor produces an ahead-of-time compiled deployment artifact, FlexKV-backed KV caching enables reuse of previously computed attention state, and NVIDIA Dynamo-Triton provides the production serving environment.

This approach is especially valuable when successive requests share an unchanged prefix of a user’s interaction history. Rather than recomputing attention over that portion of the sequence, the model can reuse its cached key-value state and compute only what is required for newly appended tokens. The potential savings increase with longer sequences and deeper models, where repeated computation across multiple HSTU layers would otherwise add significant latency.

Get started accelerating HSTU GR inference

You can reproduce and extend this workflow from the NVIDIA/recsys-examples GitHub repo. The HSTU overview introduces the GR model structure, including contextual tokens, item tokens, action tokens, embedding tables, HSTU blocks, and prediction heads.

The AOTI inference guide walks through building the required images and libraries, preparing KuaiRand-1K data, training a checkpoint, exporting the KV-cache AOTI model, validating it with C++ replay, packaging the Dynamo-Triton runtime image, and replaying requests through the Dynamo-Triton server.

To learn more, check out these related resources:

Acknowledgments

This post is a cross-functional effort across several NVIDIA teams. We would like to thank J, Runchu Zhao, Yulu Liu, Lin Hu, Zhuofan Li, Jacob Subag, and Tomer Bar-On for their contributions.

来源:NVIDIA Technical Blog:Agentic AI / Generative AI · developer.nvidia.com