A100 numbers to know
Arithmetic is nearly free; moving bits is what you pay for. These are the constants for pricing an MNIST learner on an A100-SXM4-40GB before writing it, then a 57-card deck to memorize them.
Each level down costs 10× and holds 1,000× more
published (NVIDIA, Dally, Horowitz)measured on our A100estimate, within ~3×
| Level | Capacity N | Latency E | Energy / bit E | Bandwidth | What fits |
|---|---|---|---|---|---|
| Register file | 256 KB per SM; ≤ 255 per thread | within a cycle | ≲ 0.1 pJ | — | loop state |
| L1 + shared memory | 192 KB per SM; 163 KB per block | ~1 ns | ~0.1 pJ | ~19 TB/s FA | MNIST-small; LeNet-5 in FP16 |
| L2 cache | 40 MB, two partitions | ~10 ns | ~1 pJ | unpublished | MNIST-medium; 4-bit MNIST-train |
| HBM2 | 40 GB | ~100 ns | ~10 pJ (≤ 32) | 1,555 GB/s N | MNIST-train, 850× over |
| Host over PCIe 4 | — | ≥ 10 µs | 30+ pJ | 31.5 GB/s each way N | upload once |
This is Dally's ladder (local SRAM 5 pJ per word, on-chip SRAM 50, DRAM 640) laid over NVIDIA's capacities. Shared memory stops at 48 KB per block unless you opt in.
An add is worth 10 µm of wire
| Heuristic | What to remember | |
|---|---|---|
| Leaving the die | SerDes 10 · LPDDR4 4 · DDR4 20 · HBM ≈ 5 pJ per bit, plus 4.8 to cross the GPU die | D20 |
| Store or recompute | Recompute wins below n = 4,000: a 640 pJ off-chip round trip versus n × 160 fJ | D22 |
| Recipe for 1 pJ per op | 8–16-bit data; tens of ops per local fetch; about a thousand per DRAM fetch | H14 |
| Placement | Random versus local placement: 354× the communication energy for the same O(n) | D22 |
| Amortize overhead | A 30 pJ fetch/decode is 2000% of an FP16 multiply-add, 22% of a matrix instruction | D23 |
| Big SRAM | 256 MB ≈ 100 mm²: 58 pJ and 12.3 ns per word, 99% of it wire | D22 |
| Voltage | E = ½CV². 0.5 V is 2× better than 0.7 V and 4× better than 1.0 V | D23 |
Dally's prescription: do less, with smaller numbers, locally, combinationally, sparsely.
Tensor Cores make FP16 sixteen times faster than FP32
| Limit | Value N |
|---|---|
| Where 312 comes from | 108 SMs × 1,024 multiply-adds per clock × 1.41 GHz = 156 T per second |
| Per SM | 64 concurrent warps · 32 thread blocks · 64K registers · ≤ 255 registers per thread |
| Shared memory | Carve-out 0 / 8 / 16 / 32 / 64 / 100 / 132 / 164 KB per SM; 163 KB per block; 48 KB static unless opted in |
| L2 cache | Two partitions; residency controls pin data; compression gives up to 2× capacity, 4× bandwidth |
| Matrix instructions | FP16/BF16 16×8×16 · TF32 16×8×4 · int8 16×8×32 · 1-bit 16×8×256. Pad dimensions to 8; 64 is safest |
| On-die SRAM | 27 MB registers + 20 MB L1/shared (108 separate islands) + 40 MB L2 ≈ 87 MB |
| Roofline ridge | 312 TFLOPS ÷ 1,555 GB/s = 200 FLOPs per byte. A dense layer needs batch ≈ 200; an elementwise op is 1,000× below |
Ten decades of latency, seventeen of energy
published (NVIDIA, Dally, Horowitz)measured on our A100estimate, within ~3×
Below FP16, cheaper arithmetic barely moves the board
One BF16 numerical-policy flag made weight-gradient GEMMs 9× slower. Check the reduction policy before blaming backprop.
MNIST-train is 5 MB too big for the L2 cache
Nine rules for the MNIST kernel
- Price movement, not math. Count bits × level at 0.1 / 1 / 10 pJ per bit.
- Make the working set fit a level. Small → one block. Medium → L2. Original → tile through HBM, or shrink pixels into L2.
- One launch, no syncs. Fuse train + predict into one kernel.
- Time is energy until the chip is saturated. The fastest program is the cheapest.
- Earn 100 multiply-adds per HBM byte. Batch ≳ 200, or keep weights on-die. Fuse every elementwise op.
- Use Tensor-Core precisions. Below FP16, choose precision for footprint.
- Recompute rather than spill. Storing wins only on-die.
- Respect the square root and the grain. Keep hot state small; pad matrix dimensions to 8.
- Measure like the leaderboard. NVML energy counter, 4–5 s windows, loaded idle subtracted.
Flashcards: memorize the 25 core numbers first
Tap the card to flip it, then Got it or Again (the card returns later in the round).
With the card focused, Space flips, → / G = got it, ← / A = again. Progress stays in this browser.
All 57 cards as a list
Sources and method
Sources for every tag
- D23 W. Dally, “Energy Efficiency and AI Hardware,” Stanford AHA Retreat keynote, 2023: slides 8, 11–13, 28, 31, 35.
- D22 W. Dally, “On the Model of Computation: Point,” CACM 65(9).
- D20 W. Dally, Y. Turakhia, S. Han, “Domain-Specific Hardware Accelerators,” CACM 63(7): cost model, p. 56.
- H14 M. Horowitz, “Computing's Energy Problem (and what we can do about it),” ISSCC 2014, Fig. 1.1.9.
- N NVIDIA, “Ampere Architecture In-Depth” (Table 3, SM and L2 sections) and the Ampere Tuning Guide §1.4.
- M Our runs on one A100-SXM4-40GB (400 W limit, PyTorch 2.5.1, CUDA 12.4): NVML energy counter, loaded idle subtracted, host excluded. Eager PyTorch / cuBLAS figures, not hardware limits. Repo: cybertronai/sutro-problems.
- FA T. Dao et al., “FlashAttention,” Fig. 1. G general community figure. Dean's ladder: J. Dean, LADIS 2009 keynote.
How the estimates were made, and how far to trust them
- NVIDIA publishes capacity and bandwidth per level, but neither latency nor energy.
- The estimates apply Dally's 2020 cost model (50 fJ/bit at the sub-array + 100 fJ/bit-mm over 2√area; 0.3 ns + 0.4 ns/mm) to NVIDIA's capacities at ~0.06 µm² per SRAM bit.
- 192 KB → 0.31 mm a side → ~0.11 pJ/bit, ~0.5 ns. 40 MB → 4.5 mm a side → ~0.95 pJ/bit, ~4 ns, plus the unpublished trip from an SM to the L2 partition.
- HBM is Dally's own slide: 5 pJ/bit of access plus a 48 mm, 4.8 pJ/bit round trip. It sits under the hard bound TDP ÷ bandwidth = 32 pJ/bit.
- The model reproduces Dally's stated “100 MB → 0.7 pJ/bit” exactly. Treat every estimate as right within about 3×: enough to rank designs.
- “Word” on Dally's ladder slide is read as 64 bits, which matches his CACM figures. Tier sizes follow the current competition README (1,000 / 10,000 / 60,000).