GPU Memory Physics: DRAM vs SRAM, Kernel Fusion and FlashAttention
Here is a fact that surprises most engineers the first time they see it: a modern data-centre GPU can perform roughly 300 arithmetic operations in the time it takes to fetch a single byte from its own main memory. The arithmetic units are not the bottleneck in most AI workloads. The walk to memory is. Understanding the GPU memory hierarchy is therefore the single most useful skill for explaining why one kernel runs at 5 percent of peak and another, doing the same maths, runs at 70 percent.
This is Part 3, the finale, of our series “GPU Compute: From iPhone to H100”. In Part 2 on CUDA kernel execution and data starvation we watched a kernel launch and saw how idle compute units wait on data. Here we explain the physics behind that waiting, and the engineering trick, kernel fusion with tiling, that FlashAttention used to cut it.
You will leave with a mental model of every memory level, worked byte-counting arithmetic for attention, and a one-page cheat sheet.
What this covers: a school analogy, the H100 memory levels with verified numbers, the roofline idea, unfused versus fused attention, online softmax, serving and iPhone implications, failure modes, and a recap of the series.
Context and Background
The series so far
Part 1, the anatomy of GPU parallelism, introduced the execution hierarchy: a GPU is made of Streaming Multiprocessors (SMs), each SM runs thread blocks, each block is split into warps of 32 threads, and the thread is the smallest unit. Part 2 followed a kernel through launch, scheduling and stalls. This final part asks the question both of them kept raising: where does the data live while all of those threads are working on it?
Why memory, not arithmetic, is the story
A Graphics Processing Unit (GPU) was designed for throughput. It hides the delay of any single operation by keeping thousands of threads in flight. That trick works only if data arrives fast enough to feed them. Over the last decade, peak arithmetic throughput has grown much faster than memory bandwidth, a gap often called the memory wall. The H100 SXM5 offers about 989.5 teraFLOPS of dense FP16 Tensor Core throughput (per a vendor-derived specification table) against roughly 3.35 terabytes per second of High Bandwidth Memory (HBM3) bandwidth. Dividing the first by the second gives about 295 floating point operations per byte. A kernel that does fewer than that many operations on each byte it loads from HBM cannot keep the Tensor Cores busy, no matter how clever the code is.
Two kinds of memory, two kinds of physics
Dynamic Random Access Memory (DRAM) stores each bit as charge in a tiny capacitor. It is dense and cheap per bit, so it can be stacked in large capacities, but every read needs the charge to be sensed and refreshed, and the data has to travel off the compute die. HBM is DRAM stacked vertically and placed beside the GPU die on a silicon interposer, which gives a very wide bus and high bandwidth. It is still DRAM, and it is still off-chip.
Static Random Access Memory (SRAM) stores each bit in a small circuit of transistors (typically six) that holds its state while powered. It is far faster and sits on the same die as the compute units, but each bit costs much more silicon area, so capacities are small. This difference, big and slow versus small and fast, is the whole story of this post. On the H100 the SRAM-class structures are the register file, the combined L1 cache and shared memory in each SM, and the shared L2 cache. NVIDIA’s Hopper Tuning Guide states that the combined L1, texture and shared memory capacity is 256 KB per SM on compute capability 9.0, that shared memory is up to 228 KB per SM and 227 KB per thread block, and that the L2 cache grew from 40 MB on the A100 to 50 MB on the H100.
The Feynman Test: A School Full of Students
Richard Feynman’s rule was that if you cannot explain an idea to a first-year student, you do not understand it. So here is the GPU memory hierarchy as a school.
Imagine a school with 132 classrooms. Each classroom is an SM. Inside a classroom, rows of 32 students work in lockstep, which are warps from Part 1. A thread block is one class. Every student is a thread.
The students are doing homework that needs facts from books. There are three places books can live.
- The book open on your desk is a register. Only you can see it, and reading it takes no effort at all. Each student has only a few of these.
- The shelf at the back of your classroom is shared memory. Everyone in the class can use it, it is a few steps away, and the whole classroom shares one shelf. It is small.
- The town library across the road is HBM. It holds everything, tens of gigabytes of it, but every trip means putting on a coat, crossing the street, queueing, and walking back. Between the library and the classrooms sits a school-wide cupboard in the corridor, the L2 cache, which keeps copies of recently fetched books.
Now consider two ways of doing a four-step homework problem. In the naive way, a student walks to the library, copies a page, walks back, does step one, and writes the answer on a sheet. Then they walk the sheet back to the library, fetch it again for step two, and so on. Four steps, eight street crossings. In the fused way, the student fetches the pages once, does all four steps at the desk using the shelf as scratch space, and returns once with the final answer. Same homework, same answer, a quarter of the walking.
That is kernel fusion. And tiling is what you do when the homework is too big for the desk: you fetch a manageable chunk of pages, finish everything you can with that chunk, and then fetch the next chunk. FlashAttention is simply this idea applied to the attention step in a transformer. The rest of this post makes each analogy precise.
The iPhone fits the same picture, with a twist we return to later: the town library is shared with every other app on the phone, and the classrooms are much smaller.
The GPU Memory Hierarchy on an H100, Level by Level
The GPU memory hierarchy is a pyramid of storage in which each level up is smaller, faster and closer to the arithmetic units: HBM holds tens of gigabytes at about 3.35 TB/s, the L2 cache holds 50 MB, each SM holds up to 228 KB of shared memory plus a 256 KB register file, and registers are the fastest of all. Performance engineering means keeping data at the top for as long as possible.

Figure 1: The H100 memory pyramid. Capacity grows and speed falls as you move away from the Tensor Cores.
The figure shows the capacities and the direction of cost. The text diagram below adds the same information in a form you can paste into notes or a terminal.
+-----------------------------+
| Registers 64K x 32-bit | per SM, ~256 KB, private to a thread
+--------------+--------------+
|
+--------------v--------------+
| Shared memory / L1 SRAM | per SM, 256 KB combined, up to 228 KB shared
+--------------+--------------+
|
+--------------v--------------+
| L2 cache SRAM, shared | 50 MB, all 132 SMs share it
+--------------+--------------+
|
+--------------v--------------+
| HBM3 DRAM, off-die stacks | 80 GB, about 3.35 TB/s
+-----------------------------+
smaller, faster, costlier per byte ^ | bigger, slower, cheaper per byte
Figure 1 and the text diagram say the same thing: data you reuse should be promoted up the pyramid, and data you touch once should stream through it without being stored.
Registers: the book on your desk
Each SM on a Hopper GPU has a register file of 64K 32-bit registers, which is 256 KB, according to the architecture documentation. A thread can use up to 255 registers. Registers are the only place where the Tensor Cores and arithmetic units can read operands at full speed, and values in them are private to one thread. Why does this matter? Because a thread that needs more registers than are available spills them to local memory, which is a slow region that physically lives in HBM and is cached. Where do you notice this? In the compiler output, which reports spill loads and stores. When does it bite? Typically when you make a fused kernel more ambitious. How do you handle it? Reduce tile sizes or the number of live values, and accept that registers are the scarcest resource per thread.
Shared memory SRAM: the shelf in the classroom
Shared memory is a programmer-managed scratchpad that all threads in a block can read and write. It shares physical SRAM with the L1 cache. Its latency is a small multiple of register latency, and its bandwidth per SM is very high, because it is on-die and banked into 32 banks so that a warp can access 32 words in a single cycle if they fall in different banks. Why does the programmer manage it by hand? Because the algorithm usually knows better than any cache policy which tile will be reused next. Where is it used? Matrix multiplication tiles, reduction scratch, and every FlashAttention kernel. When should you use it? When data is reused by several threads in a block. How? Declare it in the kernel, load a tile cooperatively, call a barrier such as __syncthreads(), then compute from it. Hopper adds the Tensor Memory Accelerator (TMA), which the Hopper Tuning Guide describes as moving 1D to 5D tensors between global memory and shared memory asynchronously, and thread block clusters with distributed shared memory so that blocks in a cluster can access each other’s shared memory.
L2 cache and HBM: the corridor cupboard and the library
The L2 cache is hardware-managed SRAM shared by all SMs. At 50 MB it is large for on-chip memory, and it is why repeated reads of a small working set can be much faster than the HBM figure suggests. A caveat: independent microbenchmarks of the PCIe H100 variant, run by Chips and Cheese, reported that accessing the far partition of L2 takes nearly twice as long as the near partition, so L2 is not uniform. HBM is the main memory. The SXM5 H100 carries 80 GB of HBM3 with about 3.35 TB/s of bandwidth, whereas the earlier A100 40 GB had 1.55 TB/s (NVIDIA Hopper Tuning Guide). Unloaded access latency to HBM is on the order of hundreds of nanoseconds in published microbenchmarks of recent NVIDIA GPUs; I do not quote a single exact H100 figure here because the sources I could verify do not give one for the SXM5 part.
Putting numbers on “fast” and “slow”
Exact latencies vary by generation, clocks and access pattern, so treat the table below as orders of magnitude drawn from the capacities above and general microbenchmark literature, not as measured H100 values.
| Level | H100 capacity (verified) | Scope | Relative speed (orders of magnitude) | Managed by |
|---|---|---|---|---|
| Registers | about 256 KB per SM | One thread | Fastest, effectively zero extra latency | Compiler |
| Shared memory / L1 | 256 KB per SM combined, up to 228 KB shared | One thread block | Tens of cycles | Programmer / hardware |
| L2 cache | 50 MB | Whole GPU | Low hundreds of cycles | Hardware |
| HBM3 | 80 GB, about 3.35 TB/s | Whole GPU | Several hundred nanoseconds | Allocator |
The shape matters more than any one entry. Between registers and HBM there are roughly three orders of magnitude of latency and five orders of magnitude of capacity. Every optimisation in this post is a way of getting more arithmetic done per trip across that gap.
Roofline Thinking: Is Your Kernel Memory-Bound?
A kernel is limited either by how fast it can compute or by how fast it can move data. The roofline model, introduced by Williams, Waterman and Patterson in 2009, makes that visible by plotting achievable performance against arithmetic intensity, the number of floating point operations performed per byte moved from main memory. Below a threshold called the ridge point, bandwidth sets the ceiling. Above it, compute sets the ceiling.
For the H100 SXM5 with dense FP16 Tensor Core peak of about 989.5 TFLOPS and 3.35 TB/s, the ridge point is 989.5e12 divided by 3.35e12, or roughly 295 FLOP per byte. Many real operations sit far to the left of that. An element-wise add of two FP16 vectors does one operation per six bytes moved (two reads of two bytes, one write of two bytes), an intensity near 0.17. A layer norm, a softmax, a GELU activation and a dropout mask are all in the same regime. These are the memory-bound kernels: their runtime is set by bytes divided by bandwidth, and making the arithmetic cheaper does nothing.
The lesson for the rest of this post: for memory-bound kernels, the only lever is moving fewer bytes. There are three ways to do that. You can use smaller data types, which is quantisation (see our FP8 vs INT8 vs INT4 quantisation benchmark). You can reuse data on chip through tiling. And you can avoid writing intermediate results out and reading them back, which is kernel fusion.
The five W’s and one H of the key terms
Every term in this post is easier to hold if you can answer six questions about it, so here they are in compact form.
HBM. What: stacked DRAM on the GPU package. Why: capacity and bandwidth that a conventional memory bus cannot offer. Who: suppliers manufacture it, NVIDIA integrates it. Where: beside the die on an interposer. When: you pay for it on every first touch of data. How: very wide parallel bus, thousands of pins.
Shared memory SRAM. What: programmer-managed on-die scratchpad. Why: to reuse data without going to HBM. Who: threads of one block. Where: inside each SM. When: whenever two or more threads read the same value. How: declare, cooperatively load, synchronise, compute.
Tiling. What: splitting a big problem into pieces that fit in fast memory. Why: because the whole problem cannot fit. Where: matrix multiplication and attention. When: when data is reused. How: pick a tile size so that tile inputs, outputs and scratch fit in shared memory and registers.
Kernel fusion. What: combining several kernels into one so intermediates stay on chip. Why: each kernel boundary forces a round trip to HBM. Who: compilers such as XLA, Triton and TorchInductor, or expert authors by hand. Where: chains of memory-bound operations. When: when the intermediate is large and consumed immediately. How: the next section.
Why Attention Is Memory-Bound: Counting Bytes
Standard self-attention for one head takes queries Q, keys K and values V, each of shape N by d, where N is the sequence length and d is the head dimension. It computes the score matrix S = QK^T, which is N by N, applies a row-wise softmax to get P, and then computes the output O = PV. The matrix multiplications are compute-friendly. The trouble is the N by N matrix in the middle.
Worked arithmetic for one head
Take N = 4096 and d = 128 in FP16, which is two bytes per element. This is an illustrative configuration, chosen because the numbers are clean, not a benchmark of any particular model.
- Each of Q, K, V and O holds 4096 x 128 x 2 bytes, which is 1,048,576 bytes, or 1 MiB.
- The score matrix S holds 4096 x 4096 x 2 bytes, which is 33,554,432 bytes, or about 33.5 MB. P is the same size.
In the naive, unfused flow, each step is a separate kernel and each kernel reads its inputs from HBM and writes its outputs back. The steps are: (1) read Q and K, write S; (2) read S, write the softmax output P; (3) read P and V, write O. Counting only the N by N traffic, S is written once and read once, and P is written once and read once, which is four passes of 33.5 MB, or about 134 MB. Add about 4 MiB for Q, K, V and O, and the total is roughly 138 MB of HBM traffic for a single head. Real implementations add more passes for scaling, masking and dropout, so this is a lower bound for the naive flow.
At 3.35 TB/s, moving 138 MB takes about 41 microseconds. The arithmetic for the same head is two matrix multiplications of 2 x N x N x d floating point operations each, or 4 x 4096 x 4096 x 128, which is about 8.6 GFLOP. At the dense FP16 peak of 989.5 TFLOPS that takes about 8.7 microseconds if everything ran at peak. So the naive flow spends roughly four times longer waiting on memory than it would need to compute. Its arithmetic intensity is 8.6 GFLOP divided by 138 MB, about 62 FLOP per byte, well to the left of the 295 FLOP per byte ridge point.
Now note how the traffic scales. Doubling N quadruples the 134 MB term, while the compute also quadruples. The ratio stays the same, but the absolute memory footprint of S becomes a hard limit: at N = 32,768, S alone would be 2 GiB per head in FP16, and a model with 32 heads would need 64 GiB just for scores in a single layer if all were materialised at once. That is the practical reason long context was impossible with the naive flow.
The unfused flow as a sequence

Figure 2: Naive attention. Three kernels, and the N by N matrices make two full round trips to HBM.
The sequence below is the same flow as plain text, with a byte tally.
Kernel 1 matmul : HBM -> read Q,K (2 MiB) -> compute S = QK^T
S -> write to HBM (33.5 MB)
Kernel 2 softmax : HBM -> read S (33.5 MB) -> compute P = softmax(S)
P -> write to HBM (33.5 MB)
Kernel 3 matmul : HBM -> read P (33.5 MB), V (1 MiB) -> compute O = PV
O -> write to HBM (1 MiB)
Total N x N traffic : 4 x 33.5 MB = 134 MB (+ about 4 MiB for Q,K,V,O)
Each arrow into HBM is a round trip that adds latency, consumes bandwidth, and forces the next kernel to pay a launch cost and wait for the previous one to finish completely. Part 2 explained why launch gaps and the tail of each kernel leave SMs idle; unfused flows pay that cost three times.
What “fusion” actually means
Kernel fusion means writing a single kernel that performs the whole chain, so that intermediate values live in registers or shared memory and are never written to HBM. In the school analogy, the student does all four steps at the desk. In GPU terms, the fused kernel loads a tile of Q, loops over tiles of K and V, computes scores, applies softmax, accumulates into the output, and writes only O.
Fusion is not magic and it is not free. It does not reduce the arithmetic. It removes the stores and loads of intermediates, so it helps in proportion to how large those intermediates are compared with the inputs. For element-wise chains such as bias plus GELU plus dropout, fusion is straightforward, because each output element depends on one input element. For attention it is harder, because softmax needs a whole row of S before it can normalise, and the whole row is exactly the N-sized thing you cannot hold on chip when N is large. Solving that is the contribution of FlashAttention.
Tiling and Online Softmax: The FlashAttention Idea
The FlashAttention paper by Tri Dao, Daniel Fu, Stefano Ermon, Atri Rudra and Christopher Re, arXiv 2205.14135, describes the method as an IO-aware exact attention algorithm that uses tiling to reduce the number of memory reads and writes between GPU HBM and on-chip SRAM. The paper states that it requires fewer HBM accesses than standard attention and is optimal for a range of SRAM sizes. Two ideas make this work: tiling and an online softmax.
Tiling attention
Instead of forming all of S, the kernel splits Q into row blocks and K and V into column blocks. A block of Q stays in fast memory while the kernel streams through blocks of K and V. For each pair it computes a small tile of scores, say 128 by 64, entirely in registers and shared memory. As an illustration, a 128 by 128 FP16 Q tile is 32 KB, and 64 by 128 K and V tiles are 16 KB each, a total of 64 KB, comfortably inside the 228 KB of shared memory on an H100 SM. Production kernels choose tile sizes by experiment, so treat these as illustrative.
Online softmax: the trick that removes the row dependency
Softmax of a row is exp(x_i – max) divided by the sum of exp(x_j – max). The max subtraction is for numerical stability. The naive approach needs the full row’s maximum first, which seems to force a complete pass. The online normaliser, described by Milakov and Gimelshein of NVIDIA in arXiv 1805.02867, keeps a running maximum m and a running sum l, and rescales the sum whenever the maximum changes. They reported up to 1.3 times speedup for softmax and up to 5 times for fused softmax plus top-k, from reducing memory accesses.
FlashAttention extends the same recurrence to the output. For each new K and V tile with local score block s, the kernel updates:
m_new = max(m_old, rowmax(s))
l_new = exp(m_old - m_new) * l_old + rowsum(exp(s - m_new))
O_new = exp(m_old - m_new) * O_old + exp(s - m_new) @ V_tile
...after the last tile: O = O_new / l_new
The previous partial results are rescaled by exp(m_old – m_new), which is the correction for having used an older maximum. At the end, dividing by the final sum gives exactly the same result as full softmax attention, up to floating point rounding. This is why FlashAttention is exact, not approximate. Many other efficient-attention methods trade accuracy for speed; this one trades nothing.
The fused flow as a sequence

Figure 3: Fused attention. The N by N matrix never exists in HBM. Only Q, K, V are read and only O is written.
Load Q tile (HBM -> shared memory, once)
for each K,V tile j:
Load K_j, V_j (HBM -> shared memory)
s = Q_tile @ K_j^T # registers, never leaves the chip
update running m, l # registers
O += exp(s - m) @ V_j # registers, rescaled as needed
Write O tile (registers -> HBM, once)
Total HBM traffic: Q + K + V + O (about 4 MiB in our example, not 138 MB)
For our worked example the fused flow moves roughly 4 MiB of HBM traffic per head instead of roughly 138 MB, about a 32-fold reduction. At 3.35 TB/s that is about 1.25 microseconds of memory time, which is now well below the 8.7 microseconds of arithmetic at peak. The arithmetic intensity rises to about 2,000 FLOP per byte, far above the 295 FLOP per byte ridge point. The kernel has moved from the left side of the roofline to the right side: it is now limited by compute, which is exactly where you want an attention kernel to be. K and V tiles are re-read once per Q block, so real traffic is somewhat higher than the idealised figure, which is why the paper’s IO complexity involves the SRAM size.
IO complexity, without the symbols fear
The paper’s analysis gives standard attention HBM accesses on the order of N x d plus N squared. FlashAttention’s accesses are on the order of N squared times d squared divided by M, where M is the size of SRAM. Because d is typically 64 to 128 and M is on the order of 100 KB or more, d squared divided by M is much smaller than one, so the N squared term shrinks by a large constant factor. Two things follow. The work remains quadratic in N, so FlashAttention does not make attention linear. But the memory footprint becomes linear in N, because S is never stored, and that is what unlocked long context.
From FlashAttention to FlashAttention-3: What Changed on Each Generation
The original paper reported a 15 percent end-to-end wall-clock speedup on BERT-large at sequence length 512 against the MLPerf 1.1 training record, a 3 times speedup on GPT-2 at sequence length 1K, and a 2.4 times speedup on the long-range arena benchmark at lengths of 1K to 4K. Those are the authors’ figures and they depend on hardware and baseline, but the pattern is clear: the longer the sequence, the larger the win, because the N squared intermediate is what was being avoided.
FlashAttention-2 revisited work partitioning. Its abstract reports around 2 times speedup over the first version, up to 225 TFLOPs per second per A100 in end-to-end GPT-style training, and 72 percent model FLOPs utilisation. The three changes it lists are fewer non-matrix-multiply operations, parallelising across thread blocks along the sequence dimension as well as across heads and batch, and distributing work between warps to reduce shared memory traffic. Connect this to Part 1: it is a warp-level scheduling improvement as much as a memory one.
FlashAttention-3 targets Hopper. According to its abstract it reaches 1.5 to 2.0 times speedup over FlashAttention-2 with FP16, up to 740 TFLOPs per second, which is 75 percent utilisation of the H100, and close to 1.2 PFLOPs per second with FP8 while reducing numerical error by 2.6 times compared with a baseline FP8 attention. The paper notes that FlashAttention-2 reached only around 35 percent utilisation on H100. The techniques use Hopper-specific features: asynchronous Tensor Core instructions, TMA copies so that data movement overlaps computation, and warp specialisation, in which some warps act as producers loading data while others act as consumers computing. These map onto the hardware features we described above. Note the progression: version 1 removed unnecessary HBM traffic, version 2 fixed parallelism, version 3 hides the remaining latency by overlapping copy and compute.
A caution on comparing numbers
All of these figures come from the authors, on specific hardware, sequence lengths, head dimensions and precisions. They are not interchangeable with your end-to-end model speedup, because attention is only part of a transformer step. In decoding, other effects dominate, as the next section shows.
Why This Matters for LLM Inference Serving
Inference has two phases with opposite personalities. In prefill, the model processes the entire prompt at once, the matrices are large, and attention over the prompt benefits directly from FlashAttention-style fused kernels. In decode, the model emits one token at a time, and each step must read all the model weights and the accumulated key-value (KV) cache from HBM to produce a single token. The arithmetic intensity is very low, so decode is the textbook memory-bound workload.
Decode arithmetic
Consider a hypothetical 8-billion-parameter model stored in FP16, which is 16 GB of weights. At batch size 1, every generated token requires reading all 16 GB. Dividing by 3.35 TB/s gives about 4.8 milliseconds per token, an upper bound of roughly 209 tokens per second per GPU no matter how many Tensor Core TFLOPS it has. This is an idealised ceiling that ignores overheads, and real systems fall below it. It explains the central lever of modern serving: batch many requests together so that one read of the weights serves many tokens. That is why continuous batching in LLM inference is such a large gain: it raises arithmetic intensity without changing the hardware.
The KV cache has its own byte arithmetic. For a model with 32 layers, 8 key-value heads (grouped-query attention) and a head dimension of 128 in FP16, each token stores 2 (K and V) x 32 x 8 x 128 x 2 bytes, which is 131,072 bytes, or 128 KiB. At an 8,000-token context that is about 1 GiB per sequence. Sixty-four concurrent sequences would hold 64 GiB, which is most of an 80 GB card before weights are counted. This is why techniques covered in our KV cache optimisation guide are about bytes, not FLOPs. Decode attention kernels also use FlashAttention-like tiling, with variants that split the KV cache across thread blocks because there is only one query row.
Where the engines fit
Serving engines differ mostly in how well they exploit these facts: fused attention kernels, paged KV storage, batching policy, and quantised weights. Our vLLM, SGLang and TensorRT-LLM benchmark on H100 compares real stacks, and it is worth reading with this lens: ask for each result whether it moved fewer bytes or computed faster. If you are sharing a GPU between workloads, remember that HBM bandwidth and L2 are shared resources too; our guide to Kubernetes GPU sharing with MIG, time-slicing and MPS explains which sharing modes isolate memory bandwidth and which do not.
From H100 to iPhone: The Same Physics at the Edge
The hierarchy is universal; the numbers change. An iPhone’s Apple-designed GPU has its own registers, a small on-chip threadgroup (shared) memory, caches, and then main memory. Two differences matter. First, the phone uses unified memory: the CPU, GPU and Neural Engine share one pool of low-power DRAM (LPDDR), not a dedicated HBM stack, so GPU bandwidth is lower by a large factor and is contended by every other part of the system. Second, there is a power and thermal budget measured in a few watts, and moving a byte from DRAM costs far more energy than doing arithmetic on it. I do not quote specific bandwidth or on-chip capacity numbers for a particular iPhone here, because Apple does not publish them in a form I verified for this post; treat any figure you read elsewhere with caution.
The consequence is that on-device language models are even more bandwidth-bound than server models. With a fraction of the bandwidth and a model that must fit in a few gigabytes of RAM shared with the operating system, quantisation to 4 bits is not an optimisation but a requirement. The same fusion logic applies: avoid writing intermediates. Frameworks that target Apple GPUs, such as Metal-based runtimes, fuse operations for exactly this reason. For more on how edge runtimes handle this, see our post on ExecuTorch on-device LLM serving with batched KV cache and on INT4 vs INT8 vs FP8 quantisation for edge NPUs. The title of this series is not a metaphor: the iPhone and the H100 differ by orders of magnitude in capacity and bandwidth, but they obey the same inequality, bytes moved divided by bandwidth versus operations divided by throughput.
Trade-offs, Gotchas, and What Goes Wrong
Fusion has costs. Here are the failure modes worth knowing before you write or choose a fused kernel.
Register pressure and spills. A fused kernel holds more live values. If the compiler runs out of registers it spills to local memory, which lives in HBM-backed space, undoing the benefit. Tile size is therefore a balance between reuse and register budget.
Occupancy versus tile size. Large shared-memory tiles reduce how many blocks can reside on an SM at once. Fewer resident warps mean less latency hiding, the scheduling point from Part 2. The best tile is not the biggest tile.
Bank conflicts. Shared memory has 32 banks. If threads in a warp access different addresses in the same bank, the accesses serialise. Padding or swizzled layouts fix it, but only if you look for it.
Non-matmul throughput. The FlashAttention-3 authors note that the softmax exponential runs on special function units with far lower throughput than the Tensor Cores, so once memory is fixed, the exponentials can become the next bottleneck. I did not independently verify the exact ratio for this post; consult the paper for the figure. This is why FlashAttention-3 overlaps softmax with matrix multiplies.
Numerical subtleties. Online softmax is exact in real arithmetic, but floating point rounding differs from the unfused path, so outputs can differ in the last bits. In FP8 the stability concerns are larger, as the FP8 error figure above suggests.
Fusion is not always possible or worth it. If an intermediate is small, fusion buys little. If two operations need different parallel decompositions, as with a reduction across a whole row followed by a column-wise step, a single kernel may be awkward. Compilers such as TorchInductor and XLA fuse what they can prove safe; hand-written kernels cover the rest.
Second-order thinking: what breaks at scale
Every optimisation changes the system it lives in. Here are the assumptions behind this post and where they weaken.
Assumption: the fix is permanent. FlashAttention moved attention to the compute-bound side on H100. If the next hardware generation raises Tensor Core throughput faster than SRAM bandwidth and capacity, the ridge point moves right and the same kernel can become memory-bound again. The result is a treadmill, which is why each Flash version targets new hardware features. Likewise, widely discussed slow scaling of SRAM density with newer process nodes suggests on-chip memory will not grow as freely as logic, though you should check the current foundry data before relying on that.
Assumption: attention is the bottleneck. After attention is fused, the mixture-of-experts routing, communication across GPUs (see our Spectrum-X fabric article), KV cache capacity and weight reads dominate. Fusing attention harder yields diminishing returns once you cross over.
Assumption: more fusion is always better. Giant fused kernels are hard to debug, hard to port between architectures, and brittle: a change in head dimension can fall off the fast path. At scale, the maintenance burden becomes an engineering cost that never shows up on a roofline plot.
Assumption: HBM bandwidth is yours alone. In multi-tenant serving, noisy neighbours on the same GPU share L2 and HBM bandwidth, so a kernel’s measured speed in isolation overstates what it gets in production.
Practical Recommendations
Start with measurement, not intuition. Profile the kernel, compute its arithmetic intensity from bytes moved and operations performed, and compare it with the ridge point of your GPU. Only then choose a remedy.

Figure 4: A diagnostic flow. Measure intensity first, then reduce bytes by fusing, tiling or shrinking data types.
The flow in Figure 4 reads left to right. If intensity is below the ridge point, the kernel is memory-bound, so ask whether an intermediate is written and re-read (fuse), whether the same data is fetched repeatedly (tile), or whether the data type is wider than needed (quantise). If intensity is above the ridge, stop optimising bytes and look at occupancy, instruction mix and Tensor Core usage.
A short checklist:
- Use a library or compiler first. FlashAttention is integrated into major frameworks, and PyTorch exposes fused scaled dot-product attention; writing your own is rarely the right first step.
- Compute bytes on paper before profiling. If the measured traffic is far above your estimate, you have found an unfused intermediate.
- Check hardware counters for achieved HBM bandwidth against peak. Near peak means you are bandwidth-limited and need fewer bytes. Far below peak with low intensity points to access patterns or latency, such as uncoalesced loads.
- In serving, raise batch size before anything else for decode, and size the KV cache budget explicitly.
- Re-measure after each hardware upgrade. The best kernel on one generation may not be on the next.
- On edge devices, quantise first, then fuse, and budget memory for the whole system, not only for the model.
Frequently Asked Questions
What is the GPU memory hierarchy?
It is the layered set of storage a GPU uses, ordered from fastest and smallest to slowest and largest: registers, shared memory and L1 cache inside each SM, the shared L2 cache, and off-chip HBM or GDDR. On an H100 that means about 256 KB of registers per SM, up to 228 KB of shared memory per SM, 50 MB of L2 and 80 GB of HBM3. Good GPU code keeps reused data near the top.
What is the difference between HBM and SRAM on a GPU?
HBM is stacked DRAM beside the GPU die. It has large capacity, 80 GB on the H100 SXM5, and about 3.35 TB/s of bandwidth, but each access costs hundreds of nanoseconds. SRAM is the on-die memory that forms registers, shared memory, L1 and L2. It is much faster and has far higher aggregate bandwidth, but is tiny, tens of megabytes in total. HBM stores the data; SRAM is where the work happens.
What is kernel fusion and why does it speed up GPU code?
Kernel fusion merges several GPU kernels into one so that intermediate results stay in registers or shared memory instead of being written to HBM and read back. It does not reduce arithmetic. It reduces bytes moved and kernel launches. It helps most for memory-bound operations such as softmax, layer norm and activation chains, where runtime is set by memory traffic rather than computation.
Why is FlashAttention faster than standard attention?
Standard attention writes the N by N score matrix to HBM, reads it for softmax, writes the result, and reads it again. FlashAttention tiles the computation, keeps scores in on-chip SRAM, and uses an online softmax so the full matrix never exists in HBM. It computes the same exact result with far fewer HBM reads and writes. The original paper reports up to 3 times speedup on GPT-2 at sequence length 1K.
Does FlashAttention make attention linear in sequence length?
No. The arithmetic is still quadratic in sequence length, since every query is compared with every key. What FlashAttention changes is memory: because the score matrix is never stored, the extra memory grows linearly with sequence length instead of quadratically, and HBM traffic drops by a large constant factor. That makes long contexts feasible, but very long contexts remain expensive in compute and KV cache size.
Is LLM decoding compute-bound or memory-bound?
At small batch sizes decoding is memory-bound. Each new token requires reading all model weights and the KV cache from HBM while performing only a few operations per byte. For a 16 GB model on a 3.35 TB/s GPU, the weight read alone sets an upper bound near 209 tokens per second at batch size 1. Batching requests raises arithmetic intensity, which is why batching policy dominates serving throughput.
Conclusion: The Series in One Page
We began in Part 1 by learning who does the work: the SM, block, warp and thread hierarchy that lets a GPU run tens of thousands of threads. In Part 2 we saw how that work is launched and scheduled, and how kernels stall when data does not arrive in time. In this final part we explained why the data is slow to arrive, the physics of DRAM versus SRAM, and the engineering answer: tile to reuse, fuse to avoid intermediates, and quantise to shrink what must move. One inequality ties the three posts together. A kernel is fast when the time to compute is the dominant cost; it is slow when the time to move bytes is.
The practical takeaway is a habit: before asking how to make the arithmetic faster, ask how many bytes cross the boundary between HBM and the chip, and whether any of them have to. For further reading, see the serving benchmarks in vLLM vs SGLang vs TensorRT-LLM on H100, the sharing modes in Kubernetes GPU sharing, and the cache analysis in KV cache optimisation for LLM inference.
One-Page Cheat Sheet
| Concept | Plain meaning (school analogy) | H100 number | Rule of thumb |
|---|---|---|---|
| Register | Book open on your desk | About 256 KB per SM | Fastest; spills hurt |
| Shared memory SRAM | Shelf in your classroom | Up to 228 KB per SM, 227 KB per block | Use for data reused within a block |
| L2 cache | Cupboard in the corridor | 50 MB | Hardware managed, shared by all SMs |
| HBM3 | Library across the road | 80 GB, about 3.35 TB/s | Minimise trips |
| Ridge point | Break-even work per trip | About 295 FLOP per byte | Below it, you are memory-bound |
| Memory-bound kernel | Waiting at the library | Softmax, layer norm, decode | Move fewer bytes |
| Tiling | Fetch a chunk, finish it, fetch the next | Tile must fit in SRAM | Maximise reuse per tile |
| Kernel fusion | Do all steps at the desk | Removes intermediate HBM writes | Helps most when intermediates are large |
| Online softmax | Keep a running tally, correct as you go | Running max m and sum l | Makes exact tiled attention possible |
| FlashAttention | Fused, tiled attention | v3 reports up to 740 TFLOPs FP16 on H100 | Exact; memory linear in N; compute still quadratic |
| Worked example | N=4096, d=128, FP16, one head | About 138 MB naive vs about 4 MiB fused | Roughly 32 times less HBM traffic |
| Decode bound | Weights read per token | 16 GB at 3.35 TB/s is about 4.8 ms | Batch to amortise |
Further Reading
- GPU parallelism anatomy: SM, block, warp and thread hierarchy (Part 1)
- GPU kernel engineering: CUDA kernel execution and data starvation (Part 2)
- Kubernetes GPU sharing with MIG, time-slicing and MPS
- vLLM, SGLang and TensorRT-LLM benchmark on H100
- FlashAttention paper, Dao et al., arXiv 2205.14135
- NVIDIA Hopper Tuning Guide
By Riju — about
