NVIDIA · NCP-AAI

The NVIDIA stack thread

Which NVIDIA tool does which job. Opens in M4 with NIM as an agent's inference endpoint and the NIM/TensorRT-LLM/Triton/Kubernetes serving stack, and closes in M7 with the NeMo Agent Toolkit orchestrating on top of that same stack, GPU-throughput tuning, and NeMo Guardrails as a first-class platform component.

NCPA-T3 · 17 lessons across 3 modules

  1. M4M4-01NVIDIA NIM as an agent's inference endpoint: latency budgets and failure handling
  2. M4M4-02Scaling with containers, Kubernetes, and load balancing
  3. M4M4-03Profiling performance and reliability under distributed load
  4. M4M4-04MLOps and governance: CI/CD, monitoring, and audit
  5. M4M4-05Balancing deployment cost against high availability
  6. M4M4-06The serving stack in context: NIM, TensorRT-LLM, Triton, and Kubernetes
  7. M7M7-01The NeMo Agent Toolkit: framework-agnostic orchestration
  8. M7M7-02Tuning NIM for GPU throughput: batching, TensorRT-LLM, and vLLM backends
  9. M7M7-03TensorRT-LLM and Triton Inference Server for latency reduction
  10. M7M7-04NeMo Guardrails as a first-class platform component
  11. M7M7-05Multimodal input pipelines on NVIDIA hardware
  12. M7M7-06How the pieces fit: one NVIDIA agentic stack end to end
  13. M9M9-01NeMo Guardrails: the five rail stages
  14. M9M9-02Layered safety frameworks: filters and escalation
  15. M9M9-03PII, agentic security, and audit trails
  16. M9M9-04Mitigating bias and toxicity in agent outputs
  17. M9M9-05Licensing and regulatory compliance

Part of the throughlines running across the NCP-AAI prep course.