NVIDIA · NCA-GENM
The efficiency thread
Representation and cost tradeoffs at the mechanical level: what tokenization throws away, why attention is quadratic, what an eval metric actually computes, how LoRA cuts fine-tuning memory, and how quantization and batching cut serving cost.
T4 · 0 lessons across 0 modules
Part of the throughlines running across the NCA-GENM prep course.