NVIDIA · NCP-AAI
The NVIDIA stack thread
Which NVIDIA tool does which job. Opens in M4 with NIM as an agent's inference endpoint and the NIM/TensorRT-LLM/Triton/Kubernetes serving stack, and closes in M7 with the NeMo Agent Toolkit orchestrating on top of that same stack, GPU-throughput tuning, and NeMo Guardrails as a first-class platform component.
NCPA-T3 · 17 lessons across 3 modules
- M4M4-01NVIDIA NIM as an agent's inference endpoint: latency budgets and failure handling
- M4M4-02Scaling with containers, Kubernetes, and load balancing
- M4M4-03Profiling performance and reliability under distributed load
- M4M4-04MLOps and governance: CI/CD, monitoring, and audit
- M4M4-05Balancing deployment cost against high availability
- M4M4-06The serving stack in context: NIM, TensorRT-LLM, Triton, and Kubernetes
- M7M7-01The NeMo Agent Toolkit: framework-agnostic orchestration
- M7M7-02Tuning NIM for GPU throughput: batching, TensorRT-LLM, and vLLM backends
- M7M7-03TensorRT-LLM and Triton Inference Server for latency reduction
- M7M7-04NeMo Guardrails as a first-class platform component
- M7M7-05Multimodal input pipelines on NVIDIA hardware
- M7M7-06How the pieces fit: one NVIDIA agentic stack end to end
- M9M9-01NeMo Guardrails: the five rail stages
- M9M9-02Layered safety frameworks: filters and escalation
- M9M9-03PII, agentic security, and audit trails
- M9M9-04Mitigating bias and toxicity in agent outputs
- M9M9-05Licensing and regulatory compliance
Part of the throughlines running across the NCP-AAI prep course.