~/nvidia-ncp-genl · nvidia-ncp-genl_d1_8ae04dbf52d2 ▊
A team deploying an LLM inference service wants to scale the context (prefill) phase separately from the generation (decode) phase so each can use dedicated GPUs. Which NVIDIA Dynamo capability implements this approach?
Unlock this exam to join the per-question discussion and vote on questions.