A100 numbers to know

Arithmetic is nearly free; moving bits is what you pay for. These are the constants for pricing an MNIST learner on an A100-SXM4-40GB before writing it, then a 57-card deck to memorize them.

1 fJ / bit
One add. Worth 10 µm of wire: a byte fetched across a 16 mm die costs 1,600 adds. D23
×10 per level
0.1 → 1 → 10 pJ per bit from L1 to L2 to HBM. Latency climbs the same way. E
2.6 pJ
One FP16 multiply-add at the board. Only 6% of it is arithmetic. M
18 µs
One eager launch ≈ 0.6 mJ: the entire MNIST-small record. M

Each level down costs 10× and holds 1,000× more

published (NVIDIA, Dally, Horowitz)measured on our A100estimate, within ~3×

Three ladders, one shapeLog scales, one gridline per ×10. NVIDIA publishes capacity; latency and energy are estimates.
Capacity
NVIDIA-published
L1 + shared192 KB
L2 cache40 MB
HBM240 GB
Latency
estimate, per reference
L1 + shared~1 ns
L2 cache~10 ns
HBM2~100 ns
Energy
estimate, per bit
L1 + shared~0.1 pJ
L2 cache~1 pJ
HBM2~10 pJ
LevelCapacity NLatency EEnergy / bit EBandwidthWhat fits
Register file256 KB per SM; ≤ 255 per threadwithin a cycle≲ 0.1 pJ—loop state
L1 + shared memory192 KB per SM; 163 KB per block~1 ns~0.1 pJ~19 TB/s FAMNIST-small; LeNet-5 in FP16
L2 cache40 MB, two partitions~10 ns~1 pJunpublishedMNIST-medium; 4-bit MNIST-train
HBM240 GB~100 ns~10 pJ (≤ 32)1,555 GB/s NMNIST-train, 850× over
Host over PCIe 4—≥ 10 µs30+ pJ31.5 GB/s each way Nupload once

This is Dally's ladder (local SRAM 5 pJ per word, on-chip SRAM 50, DRAM 640) laid over NVIDIA's capacities. Shared memory stops at 48 KB per block unless you opt in.

An add is worth 10 µm of wire

100 fJ / bit-mm
On-chip wire, flat across process nodes. D23
0.4 ns / mm
Signal speed on a chip, about c/120. D22
50 fJ / bit
Reading a small 8 KB SRAM. Everything bigger is wire. D20
64,000×
A DRAM fetch of two operands versus adding them. D22
99.9%
Share of a CPU instruction's energy that is overhead. D20
4× → 2×
Four times the memory doubles energy and latency per access. D20
One DRAM read buys 40,000 8-bit addsEnergy per operation at 45 nm, log scale. Horowitz, ISSCC 2014.
int8 add0.03 pJH14
int32 add0.1 pJH14
int8 multiply0.2 pJH14
FP16 add0.4 pJH14
FP32 add0.9 pJH14
FP16 multiply1.1 pJH14
int32 multiply3.1 pJH14
FP32 multiply3.7 pJH14
8 KB cache read (64-bit)10 pJH14
32 KB cache read20 pJH14
A whole CPU instruction70 pJH14
1 MB cache read100 pJH14
DRAM read1.3–2.6 nJH14
HeuristicWhat to remember
Leaving the dieSerDes 10 · LPDDR4 4 · DDR4 20 · HBM ≈ 5 pJ per bit, plus 4.8 to cross the GPU dieD20
Store or recomputeRecompute wins below n = 4,000: a 640 pJ off-chip round trip versus n × 160 fJD22
Recipe for 1 pJ per op8–16-bit data; tens of ops per local fetch; about a thousand per DRAM fetchH14
PlacementRandom versus local placement: 354× the communication energy for the same O(n)D22
Amortize overheadA 30 pJ fetch/decode is 2000% of an FP16 multiply-add, 22% of a matrix instructionD23
Big SRAM256 MB ≈ 100 mm²: 58 pJ and 12.3 ns per word, 99% of it wireD22
VoltageE = ½CV². 0.5 V is 2× better than 0.7 V and 4× better than 1.0 VD23

Dally's prescription: do less, with smaller numbers, locally, combinationally, sparsely.

Tensor Cores make FP16 sixteen times faster than FP32

108 SMs
× 64 cores = 6,912; 4 Tensor Cores each = 432. N
1,555 GB/s
HBM2. PCIe 4 is 50× less; NVLink 3 is 600 GB/s. N
100 per byte
Multiply-adds you must earn per HBM byte (200 FLOPs/byte ridge). N
826 mm² · 400 W
TSMC N7, 54.2 B transistors, 1,410 MHz clock. N
Peak throughput by precisionTrillions of operations per second (TFLOPS for float, TOPS for integer). NVIDIA.
INT4 Tensor Core1,248 TOPS
INT8 Tensor Core624 TOPS
FP16 / BF16 Tensor Core312 TFLOPS
TF32 Tensor Core156 TFLOPS
FP32 cores19.5 TFLOPS
FP64 cores9.7 TFLOPS
LimitValue N
Where 312 comes from108 SMs × 1,024 multiply-adds per clock × 1.41 GHz = 156 T per second
Per SM64 concurrent warps · 32 thread blocks · 64K registers · ≤ 255 registers per thread
Shared memoryCarve-out 0 / 8 / 16 / 32 / 64 / 100 / 132 / 164 KB per SM; 163 KB per block; 48 KB static unless opted in
L2 cacheTwo partitions; residency controls pin data; compression gives up to 2× capacity, 4× bandwidth
Matrix instructionsFP16/BF16 16×8×16 · TF32 16×8×4 · int8 16×8×32 · 1-bit 16×8×256. Pad dimensions to 8; 64 is safest
On-die SRAM27 MB registers + 20 MB L1/shared (108 separate islands) + 40 MB L2 ≈ 87 MB
Roofline ridge312 TFLOPS ÷ 1,555 GB/s = 200 FLOPs per byte. A dense layer needs batch ≈ 200; an elementwise op is 1,000× below

Ten decades of latency, seventeen of energy

published (NVIDIA, Dally, Horowitz)measured on our A100estimate, within ~3×

Latency: the GPU's new rung is the launchLog scale, one gridline per ×10. Dean's CPU ladder for scale: L1 0.5 ns, L2 7 ns, DRAM 100 ns.
32-bit add0.15 nsD22
8 KB SRAM sub-array read0.3 nsD22
1 mm of on-chip wire0.4 nsD22
One A100 clock cycle0.71 nsN
L1 / shared-memory reference~1 nsE
L2 reference~10 nsE
Across the die and back (48 mm)~19 nsD23
HBM reference~100 nsE
1 MB from HBM0.64 µsN
Raw CUDA kernel launch~5–10 µsG
MNIST-small record, train + predict17 µsM
One PyTorch eager op~18 µsM
MNIST-train (47 MB) from HBM30 µsN
6-layer MLP forward, batch 10.22 msM
MNIST-train over PCIe 41.5 msN
MNIST-medium record (5% error)3.3 msM
Sweep all 40 GB of HBM26 msN
Epoch, 12 M-param MLP, FP32, batch 641.36 sM
Energy: one idle millisecond outweighs 25 billion multiply-addsLog scale, one gridline per ×10. Measured values are above loaded idle unless marked raw.
Add or flip-flop, per bit1 fJD23
32-bit add20 fJD22
One bit moved 1 mm0.1 pJD23
16-bit multiply-add, arithmetic alone0.16 pJD22
A byte from L1 / shared memory~1 pJE
1-bit multiply-add at the board1.2 pJM
INT8 multiply-add at the board1.6 pJM
FP16 multiply-add at the board2.6 pJM
A byte from L2~10 pJE
FP32 multiply-add at the board29 pJM
One CPU instruction70–250 pJH14
A byte from HBM~80 pJE
A byte from DDR4160 pJD20
Train one example, 12 M-param MLP, BF16163 µJM
MNIST-small record0.59 mJM
One eager op of pure overhead~0.6 mJM
One pass over MNIST-train from HBM~3.8 mJE
1 ms of loaded-idle A10065 mJM
MNIST-medium record (5% error)174 mJM
1 ms of saturated GEMM, raw~400 mJM
Epoch, 12 M-param MLP, FP32, batch 64162 JM

Below FP16, cheaper arithmetic barely moves the board

FP32 → FP16 buys 11×. The next two steps buy 1.6× and 1.4×.Energy per multiply-add above idle, saturated 8192³ GEMM on one A100-SXM4-40GB.
FP32, TF32 off29 pJ
FP16 Tensor Core2.6 pJ
INT8 Tensor Core1.6 pJ
1-bit AND + popcount1.2 pJ
65 W idle
With tensors loaded (52 W bare). Every millisecond costs 65 mJ raw. M
25–50 W
Above idle for every leaderboard entry. Energy ≈ 30–50 mJ per ms: time is energy. M
~330 W
Above idle for a saturated GEMM, 397 W raw. M
6%
Arithmetic's share of the FP16 figure: 0.16 of 2.6 pJ. Precision pays through footprint. M
Overhead, not arithmetic, sets the price: same MLP, 17× spreadTraining energy per multiply-add above idle, 12 M-parameter MLP, eager PyTorch.
FP32, batch 6480 pJ
FP32, full batch33 pJ
BF16, full batch4.8 pJ
Saturated FP16 GEMM, for scale2.6 pJ

One BF16 numerical-policy flag made weight-gradient GEMMs 9× slower. Check the reduction policy before blaming backprop.

MNIST-train is 5 MB too big for the L2 cache

What fits whereSize on a log scale against the three capacities. Shrink pixels to 4 bits and the training set moves into L2.
MNIST-small, train + test, FP3272 KBblock
LeNet-5, FP16119 KBblock
784×100 layer, FP16153 KBblock
784-512-10 MLP, FP16794 KBL2
MNIST-medium, train + test, bytes1.6 MBL2
MNIST-train at 1 bit per pixel5.9 MBL2
MNIST-train at 4 bits per pixel23.5 MBL2
12 M-param MLP, FP1624 MBL2
MNIST-train as bytes47 MBHBM
12 M-param MLP, FP3248 MBHBM
MNIST-train, FP32188 MBHBM
Closed-form learners beat MLPs by three orders of magnitudeA100 energy above idle for the whole task. Filled = current record, hollow = MLP baseline. Original (1% error) is still open.
Small · 67% accuracy0.59 mJ vs 3.3 JM
Medium · 5% error174 mJ vs 290 JM
0.6 ms · 0.2 J
One epoch of a 784-512-10 MLP if Tensor Cores stayed saturated. Eager takes ~1 s. E
n ≈ 2,000
Below this width, recomputing an activation beats a round trip to HBM. E
212 images
Fit in one 163 KB block (62 in the 48 KB default). One image is 784 bytes.

Nine rules for the MNIST kernel

  1. Price movement, not math. Count bits × level at 0.1 / 1 / 10 pJ per bit.
  2. Make the working set fit a level. Small → one block. Medium → L2. Original → tile through HBM, or shrink pixels into L2.
  3. One launch, no syncs. Fuse train + predict into one kernel.
  4. Time is energy until the chip is saturated. The fastest program is the cheapest.
  5. Earn 100 multiply-adds per HBM byte. Batch ≳ 200, or keep weights on-die. Fuse every elementwise op.
  6. Use Tensor-Core precisions. Below FP16, choose precision for footprint.
  7. Recompute rather than spill. Storing wins only on-die.
  8. Respect the square root and the grain. Keep hot state small; pad matrix dimensions to 8.
  9. Measure like the leaderboard. NVML energy counter, 4–5 s windows, loaded idle subtracted.

Flashcards: memorize the 25 core numbers first

Tap the card to flip it, then Got it or Again (the card returns later in the round). With the card focused, Space flips, → / G = got it, ← / A = again. Progress stays in this browser.

tap the card to see the answer
All 57 cards as a list

Sources and method

Sources for every tag
  • D23 W. Dally, “Energy Efficiency and AI Hardware,” Stanford AHA Retreat keynote, 2023: slides 8, 11–13, 28, 31, 35.
  • D22 W. Dally, “On the Model of Computation: Point,” CACM 65(9).
  • D20 W. Dally, Y. Turakhia, S. Han, “Domain-Specific Hardware Accelerators,” CACM 63(7): cost model, p. 56.
  • H14 M. Horowitz, “Computing's Energy Problem (and what we can do about it),” ISSCC 2014, Fig. 1.1.9.
  • N NVIDIA, “Ampere Architecture In-Depth” (Table 3, SM and L2 sections) and the Ampere Tuning Guide §1.4.
  • M Our runs on one A100-SXM4-40GB (400 W limit, PyTorch 2.5.1, CUDA 12.4): NVML energy counter, loaded idle subtracted, host excluded. Eager PyTorch / cuBLAS figures, not hardware limits. Repo: cybertronai/sutro-problems.
  • FA T. Dao et al., “FlashAttention,” Fig. 1. G general community figure. Dean's ladder: J. Dean, LADIS 2009 keynote.
How the estimates were made, and how far to trust them
  • NVIDIA publishes capacity and bandwidth per level, but neither latency nor energy.
  • The estimates apply Dally's 2020 cost model (50 fJ/bit at the sub-array + 100 fJ/bit-mm over 2√area; 0.3 ns + 0.4 ns/mm) to NVIDIA's capacities at ~0.06 µm² per SRAM bit.
  • 192 KB → 0.31 mm a side → ~0.11 pJ/bit, ~0.5 ns. 40 MB → 4.5 mm a side → ~0.95 pJ/bit, ~4 ns, plus the unpublished trip from an SM to the L2 partition.
  • HBM is Dally's own slide: 5 pJ/bit of access plus a 48 mm, 4.8 pJ/bit round trip. It sits under the hard bound TDP ÷ bandwidth = 32 pJ/bit.
  • The model reproduces Dally's stated “100 MB → 0.7 pJ/bit” exactly. Treat every estimate as right within about 3×: enough to rank designs.
  • “Word” on Dally's ladder slide is read as 64 bits, which matches his CACM figures. Tier sizes follow the current competition README (1,000 / 10,000 / 60,000).