SASS and GPU Microarchitecture
PTX instructions to program Tensor Cores on the latest NVIDIA Blackwell GPUs
Performance hints from Jeff Dean and Sanjay Ghemawat
AI-ML-Roadmap-from-scratch?tab=readme-ov-file#module-4---machine-learning
Advanced Computer architecture and systems
scuttleblurb nvda1 CPU to GPU
Data gated conventional MAC shown in the video.
deeplearningbook MIT
Efficient Deep Learning Computing
6-172-performance-engineering-of-software-systems lecture-1-introduction-and-matrix-multiplication
deeplearning/performance/dl-performance-matrix-multiplication
why-gemm-is-at-the-heart-of-deep-learning
MIT introtodeeplearning - Jan 2026
arith24.arithsymposium.org/slides/s7-koenig.pdf