NVIDIA · NCP-AAI
The model-efficiency thread
Making a trained model smaller and faster without wrecking it. Opens in M4 with quantization, distillation, pruning, and KV caching, runs through M7's parallelism families and Nsight profiling, and closes in M8 where dynamic batching, NIM, and Multi-Instance GPU turn those optimizations into a served deployment.
NCPG-T1 · 0 lessons across 0 modules
Part of the throughlines running across the NCP-AAI prep course.