GPU Parallelism Explained: SM, Block, Warp and Thread Hierarchy (Series Part 1)
Your phone shoots video at 120 frames per second and, for every frame, it denoises, aligns, tone-maps and sharpens millions of pixels before the next frame arrives 8.33 milliseconds later. A single fast processor core working through those pixels one at a time would miss that deadline by orders of magnitude. The reason your camera app does not stutter is GPU parallelism: thousands of simple arithmetic lanes, organised into a strict hierarchy, all doing the same operation on different pixels at the same moment.
This matters far beyond phones. The same hierarchy that develops a photo also trains a large language model on an NVIDIA H100, and the vocabulary (streaming multiprocessor, thread block, warp) is the shared language of every GPU performance conversation in 2026. Engineers who understand it write kernels that run at hardware speed; engineers who do not wonder why a “fast” GPU sits mostly idle.
You will leave with a mental model you can draw from memory: a school analogy that maps one-to-one onto the hardware, verified H100 numbers, an honest comparison with Apple’s GPU, and a one-page cheat sheet.
What this covers: the 5W1H of GPU parallelism, CPU vs GPU design philosophy, the chip to SM to block to warp to thread hierarchy, a school analogy, a sequence-level view of scheduling, an iPhone camera case study, failure modes, and a cheat sheet.
Series note. This is Part 1 of 3 in the series GPU Compute: From iPhone to H100. Part 1 (this post) covers the hardware hierarchy. Part 2, GPU kernel engineering: how a CUDA kernel executes and why data starvation kills performance, covers the software that runs on it. Part 3, GPU memory physics: DRAM vs SRAM, FlashAttention and kernel fusion, explains why memory, not arithmetic, is the real bottleneck.
Context and Background
For about three decades the easy way to make software faster was to wait for a faster clock. That ended in the mid-2000s, when power density, not transistor count, became the limit. Chip designers kept adding transistors but could no longer turn them into one dramatically faster core. They turned them into many cores instead, and the question became how to program thousands of them without losing your mind.
Two answers emerged. A Central Processing Unit (CPU) spends its transistors on making one thread of execution fast: large caches, out-of-order execution and branch prediction. A Graphics Processing Unit (GPU) spends them on making many threads run at once, accepting that any single thread is slower and simpler. Graphics was the first natural customer, because rendering a frame is millions of independent pixel computations. Deep learning later proved to be an even better fit, since a neural network layer is mostly a giant matrix multiplication, which is the same operation applied to enormous grids of numbers.
NVIDIA’s CUDA (Compute Unified Device Architecture) platform, introduced in 2007, exposed this hardware to general programming. Its documentation, the CUDA Programming Guide, defines the vocabulary used throughout this post. Apple, AMD and others use different names for similar structures, and we will flag where the analogy bends.
If you operate GPU fleets rather than write kernels, the practical consequences of this hierarchy show up in how devices are shared; our guide to Kubernetes GPU sharing with MIG, time-slicing and MPS builds directly on the SM concept introduced below.
The 5W1H of GPU parallelism
Before the detail, here is the whole topic in six questions. Every later section expands one of these answers.
What is it? Executing one program across thousands of data elements at once, using many simple lanes that march in lockstep groups. Why does it exist? Because for data-parallel work, throughput per watt and per dollar beats latency per task. Who uses it? Graphics, computational photography, scientific simulation, and above all machine learning. Where does it live? On a GPU die, organised as many Streaming Multiprocessors (SMs), inside phones, laptops, workstations and data-centre accelerators. When is it the right tool? When the same operation applies to a large amount of independent data, and the data can be fed fast enough. How does it work? Through a hierarchy: the chip holds SMs, each SM holds thread blocks, each block is split into warps of 32 threads, and each warp executes one instruction at a time across its lanes.
The Core Idea: One Program, Thousands of Threads
The short answer: a GPU is a throughput machine built from many Streaming Multiprocessors (SMs). Work is launched as a grid of thread blocks; each block lands on one SM; the SM slices the block into warps of 32 threads; a warp scheduler issues one instruction at a time to a warp, and 32 lanes execute it together.

Figure 1: The GPU thread hierarchy. A chip contains many SMs; each SM hosts blocks, splits them into warps, and uses warp schedulers to feed the thread lanes.
Figure 1 shows the containment relationships. The chip is the whole die. An SM is the largest independent unit of execution: it owns its own register file, its own on-chip shared memory, and its own warp schedulers. Blocks are the unit of work assigned to an SM. Warps are the unit of instruction issue. Threads are the individual lanes. Notice the two arrows pointing into warps: a warp is both a slice of a block and the thing a scheduler picks.
CPU vs GPU: two philosophies
The CPU vs GPU difference is not “slow versus fast”. It is a difference in what the silicon is optimised to hide. A CPU core hides latency: when it needs data that is not in cache, it uses out-of-order execution and speculation to find other useful work. A GPU hides latency by having so many threads that something else is always ready: when one warp waits for memory, the scheduler switches to another warp at essentially zero cost.

Figure 2: CPU vs GPU. The CPU spends transistors on caches and prediction to make one task finish early; the GPU spends them on arithmetic lanes and warp switching to finish many tasks per second.
The consequence is a different performance model. For a single dependent chain of calculations, the CPU wins by a large margin. For a million independent calculations, the GPU wins by an even larger one. A rule of thumb that holds up in practice: if you can describe your problem as “do this to every element of a big array” and the elements do not depend on each other, it is a GPU problem. If each step needs the previous step’s result, it is a CPU problem, no matter how big the data.
Why the numbers matter: the 104-day sanity check
A story circulating in a source video summary that we reviewed illustrates the scale gap well, though the figures are claimed in the source and we could not trace them to a primary document. The claim is that a phone GPU performs roughly one billion operations per second per core-equivalent, and that a task of 9 million independent problems, solved at one per second on a single sequential worker, would take 104 days.
The arithmetic is easy to verify. Nine million problems at one per second is 9,000,000 seconds. Dividing by 86,400 seconds per day gives 104.17 days. That part is simply correct. The point of the story is the contrast: spread the same 9 million problems over thousands of lanes and the wall-clock time collapses from months to seconds. Treat the “one billion operations per second” figure as illustrative rather than as a measured specification; real throughput depends on the instruction, precision and clock.
The hierarchy at a glance (ASCII view)
Here is the same structure as a text diagram you can paste into a code comment. Counts are for an NVIDIA H100 SXM5 (compute capability 9.0) as documented by NVIDIA.
H100 SXM5 chip (132 SMs enabled)
|
+-- SM 0 (one "department")
| | registers : 64K x 32-bit (256 KB)
| | shared mem: up to 228 KB (configurable with L1)
| | warp schedulers: 4 sub-partitions, each issues 1 warp/cycle
| |
| +-- Thread block A (up to 1024 threads, e.g. 256)
| | +-- Warp 0 : threads 0..31 <- one instruction, 32 lanes
| | +-- Warp 1 : threads 32..63
| | +-- ...
| | +-- Warp 7 : threads 224..255
| |
| +-- Thread block B (another resident block, same SM)
| +-- Warp 0 ... Warp 7
|
+-- SM 1 ... (same structure)
+-- ...
+-- SM 131
Limits per SM (CC 9.0): 2048 threads, 64 warps, 32 blocks resident
Read it top-down: each level is a pool of resources shared by everything below it. That is why one greedy block that uses too many registers can starve its neighbours, a theme we return to in the failure-modes section.
Deeper Analysis: The School That Computes
The best way to remember the hierarchy is the one Richard Feynman would have reached for: explain it with something a child already understands. Imagine a very large school.
The analogy, mapped one-to-one
The chip is the school campus. An SM is a department, say the mathematics department. It has its own classrooms, its own supply cupboard (registers and shared memory) and its own teachers. Departments do not share a cupboard; if the physics department runs out of paper, mathematics is unaffected.
A thread block is a classroom with up to 1,024 students; we will use 256 as a comfortable example. A classroom is assigned to exactly one department and stays there until the lesson ends. Students within a classroom can talk to each other through the whiteboard (shared memory) and can all stop and wait for each other at a checkpoint (a barrier, __syncthreads() in CUDA). Students in different classrooms cannot.
A warp is a row of 32 desks. Here the analogy needs one strict rule: the 32 students in a row all receive the same instruction from the teacher at the same instant. “Everyone write the number on your card, add 5, and write it back.” They each have a different card, so they produce different answers, but they do the same step together. A thread is one student in one desk, with a private notebook (its registers).
The warp scheduler is the teacher. She can only talk to one row at a time, but she has several rows in the room. When row 3 is waiting for a textbook to arrive from the library (a memory fetch), she does not stand idle; she turns to row 5, which is ready, and gives it the next instruction. When the textbook arrives, row 3 is ready again. The kernel is the homework: one set of instructions written once, which every student in every classroom executes on their own card.
| School concept | GPU concept | What it really is |
|---|---|---|
| Campus | Chip (die) | The whole GPU with all SMs and shared L2 cache |
| Department | Streaming Multiprocessor | Independent execution unit with its own registers and shared memory |
| Classroom | Thread block | Group of up to 1,024 threads scheduled onto one SM |
| Row of 32 desks | Warp | 32 threads issued one instruction together |
| Teacher | Warp scheduler | Hardware that picks a ready warp each cycle |
| Student | Thread | One lane with private registers |
| Homework | Kernel | The function every thread runs |
| Whiteboard | Shared memory | Fast on-chip scratchpad for one block |
| Library | Global memory | Large off-chip memory, slow to reach |
Why 32, and why lockstep
The number 32 is the warp size on every NVIDIA GPU generation to date, documented in the CUDA Programming Guide. It is a design trade-off. Issuing one instruction to 32 lanes amortises the cost of fetching and decoding that instruction across 32 units of arithmetic, which is how a GPU packs so many lanes into so little area and power. NVIDIA calls the execution model SIMT (Single Instruction, Multiple Threads): programmers write code as if each thread were independent, while the hardware batches threads into warps behind the scenes.
Where does the illusion leak? On branches. If half the students in a row must do task A and half task B, the teacher cannot instruct both groups at once. The warp executes A with half the lanes switched off, then B with the other half switched off. This is called warp divergence, and its cost is the sum of both paths rather than the larger one. A kernel with a fully divergent if can run at half speed or worse, which is why GPU code favours arithmetic that is the same for all lanes.
The scheduling flow, step by step
The analogy describes the structure; the sequence below describes the motion. When the host CPU launches a kernel, a grid of blocks is created. A global block scheduler hands blocks to SMs that have enough free resources. Inside the SM, warp schedulers do the fine-grained work.

Figure 3: Scheduling sequence. The host launches a grid, the block scheduler places blocks on SMs with free resources, and warp schedulers issue instructions, switching warps whenever one stalls on memory.
The important subtlety is step three: the SM reserves registers and shared memory for the entire block at the moment the block is admitted. Resources are not allocated lazily. That is why an SM can hold 64 warps in flight, and why switching between them is free: each warp already has its own registers sitting in the register file, so a switch changes which warp the scheduler points at, not what is stored where. A CPU context switch, by contrast, saves and restores state and costs thousands of cycles.
Verified H100 numbers
Marketing pages quote headline throughput; the numbers that explain behaviour are the per-SM limits. The following are from NVIDIA’s documentation for compute capability 9.0 (Hopper) and the Hopper architecture in-depth blog.
| Property | Value | Source |
|---|---|---|
| SMs on full H100 SXM5 | 132 | NVIDIA Hopper in-depth blog |
| FP32 (32-bit floating point) cores per SM | 128 | NVIDIA Hopper in-depth blog |
| Fourth-generation Tensor Cores per SM | 4 | NVIDIA Hopper in-depth blog |
| Warp size | 32 threads | CUDA Programming Guide |
| Max threads per block | 1,024 | CUDA Programming Guide |
| Max resident threads per SM | 2,048 | CUDA Programming Guide |
| Max resident warps per SM | 64 | CUDA Programming Guide |
| Max resident blocks per SM | 32 | CUDA Programming Guide |
| 32-bit registers per SM | 64K (256 KB) | CUDA Programming Guide |
| Max registers per thread | 255 | CUDA Programming Guide |
| Max shared memory per SM | 228 KB | CUDA Programming Guide |
| Max shared memory per block | 227 KB | CUDA Programming Guide |
| L2 cache | 50 MB | NVIDIA Hopper in-depth blog |
| Memory | 80 GB HBM3 at 3.35 TB/s | NVIDIA H100 product page |
| Max power (SXM) | up to 700 W, configurable | NVIDIA H100 product page |
From these we can do some arithmetic of our own. With 132 SMs and 2,048 resident threads each, a full H100 can keep 132 x 2,048 = 270,336 threads in flight at once. With 128 FP32 cores per SM, there are 132 x 128 = 16,896 FP32 lanes across the chip. A thread block of 256 threads has 8 warps; because an SM can hold at most 2,048 threads, it can host at most 8 such blocks resident, well under the 32-block limit, so here the thread limit binds first. A block of 1,024 threads is 32 warps, so only two fit in one SM.
One further caveat on a number you will see quoted: NVIDIA lists H100 SXM at 1,979 teraFLOPS for FP16/BF16 Tensor Core operations, but that headline figure is with structured sparsity; the dense figure is half. We mention it because it is a classic way a spec sheet number outruns real workloads.
Registers: the hidden budget
The register file is the quiet constraint that governs everything. An SM has 65,536 32-bit registers. If every thread in a kernel uses 64 registers, the SM can hold 65,536 / 64 = 1,024 threads, only half the 2,048 maximum. The fraction of the maximum resident warps that actually fit is called occupancy. At 128 registers per thread the SM holds only 512 threads, an occupancy of 25 percent.
Why does occupancy matter? Because latency hiding depends on having spare warps to switch to. Where does the trade-off bite? In complex kernels that want many registers for speed but then leave the scheduler with too few warps to hide memory waits. When should you worry? Whenever a profiler shows the SM stalling on memory while few warps are resident. How do you fix it? Reduce register use, shrink the block, or accept lower occupancy if each thread does more independent work (instruction-level parallelism). Higher occupancy is not always better; it is a means to latency hiding, not a goal.
Thread block clusters: Hopper’s new layer
Hopper added an optional level between block and grid: the thread block cluster. NVIDIA describes it as giving “programmatic control of locality at a granularity larger than a single thread block”. Blocks in a cluster are guaranteed to run at the same time on neighbouring SMs, and they can read and write each other’s shared memory through a feature called distributed shared memory. In school terms, two adjacent departments get a connecting door.
The reason for the feature is that some algorithms, notably large matrix tiles, want more fast shared memory than one SM provides. Rather than spilling to slow global memory, a cluster lets several SMs pool their scratchpads. It is a useful example of how the hierarchy itself evolves: the programming model grows layers as hardware grows. We pick this thread up in Part 2.
Case Study: A 120 FPS Camera Pipeline
Now apply the hierarchy to the opening problem. At 120 frames per second, each frame has a budget of 1 / 120 = 8.33 milliseconds from sensor readout to the next frame starting. Everything the camera does to improve the picture must fit inside that window, and it must do so continuously, because a missed deadline is a dropped frame the user can see.
Mapping pixels onto the hierarchy
Take a 1080p frame as an illustrative size: 1,920 x 1,080 = 2,073,600 pixels. At 120 frames per second that is about 249 million pixels per second for every processing pass. A computational photography pipeline runs several passes: align frames, estimate and remove noise, merge exposures, tone-map and sharpen. If each pixel needs even a few hundred arithmetic operations in total, the demand is tens of billions of operations per second. These counts are illustrative; real pipelines vary widely and the vendors do not publish them.
The mapping is natural. Slice the image into square tiles of 16 x 16 pixels. Each tile is 256 pixels, so each tile becomes one thread block of 256 threads, and each thread owns exactly one pixel. Eight warps cover one tile. Neighbouring pixels in a row are consecutive threads, which also means that when they read memory they read adjacent addresses, a pattern the memory system rewards (the subject of Part 3). A 1080p frame has 120 x 67.5 tiles, so about 8,100 blocks, rounded up to cover the edge.

Figure 4: The 120 FPS pipeline mapped onto the hierarchy. Each frame is split into tiles, tiles become thread blocks, pixels become threads, and the pipeline stages run in sequence inside the 8.3 ms budget.
Why does this work so well? Because every pixel in a denoise pass runs the same instructions; branches are rare, so warps rarely diverge; and tiles are independent, so the block scheduler can place them on any SM in any order. How does it break? At tile borders, where a blur needs pixels from the neighbour tile. Real implementations load a slightly larger “halo” region into shared memory so that each block has what it needs on its own whiteboard. That is the kind of decision the kernel engineer makes, and the subject of Part 2.
What the source video claims about the iPhone, and what we can verify
The video summary we were given as a source claims an “A20 Pro” iPhone chip has 7 GPU cores, which it equates with SMs. Checking: reporting from September 2026 on Apple’s chip announcement states the A20 Pro has a seven-core GPU, one more than the six-core GPU in the A19 Pro, with Apple claiming up to 40 percent faster graphics and a 50 percent memory bandwidth increase (Notebookcheck coverage). We could verify the 7-core figure from that report, but not the “one billion operations per second” figure or the 9-million-problem scenario, which remain claimed in the source.
The equation of “GPU core” with “SM” needs care. Apple calls them GPU cores, not SMs. They play a similar structural role, since each is a repeated unit with its own register file and on-chip memory. In Apple’s Metal programming model the analogue of a warp is the SIMD-group, and Apple’s documentation and third-party analyses (notably the community microarchitecture notes at philipturner/metal-benchmarks) describe SIMD-groups of 32 threads, with a threadgroup playing the role of a thread block. Those community numbers were derived by benchmarking earlier Apple chips such as the M1 and should be treated as informative, not authoritative, for the A20 Pro. Apple has not published an SM-style specification table for its phone GPUs.
So the analogy is not one-to-one. Apple’s GPU is a tile-based deferred renderer designed around power efficiency and unified memory shared with the CPU; an H100 is a discrete data-centre accelerator with dedicated high-bandwidth memory and a 700 W budget. The concepts transfer; the numbers do not. A seven-core Apple GPU is nothing like seven NVIDIA SMs in raw capability, because the H100 has 132 SMs, each with 128 FP32 lanes.
Trade-offs, Gotchas, and What Goes Wrong
The hierarchy is powerful, and every level also has a failure mode. Most GPU performance disasters are one of the five below.
Warp divergence. As described, divergent branches serialise the paths. The usual cause is a data-dependent if on a per-thread value. The usual fix is to restructure data so that threads in the same warp follow the same path, for example by sorting or bucketing work first.
Underfilled GPU. If you launch fewer blocks than there are SMs, some SMs sit idle. Launching 100 blocks on a chip with 132 SMs leaves 32 of them with no work at all. At the other extreme, a grid whose block count is just over a multiple of the SM capacity creates a “tail effect”: the last wave of blocks runs on a mostly empty chip. This is the second-order issue of scale: a kernel that is efficient on a small GPU can be inefficient on a bigger one, simply because the problem no longer fills it.
Resource starvation inside an SM. A block that demands 200 KB of shared memory permits only one resident block per SM, because the maximum is 228 KB. Latency hiding then depends on that single block’s warps. Register and shared memory use are the hidden knobs that decide occupancy.
Wrong abstraction at the wrong scale. Thread blocks cannot synchronise with each other within a kernel except through global memory and atomics (or, on Hopper, within a cluster). Algorithms that need a global reduction between steps must be split across several kernel launches. Teams porting CPU code often discover that their serial dependencies force exactly this.
Treating the GPU as always faster. Moving data across the PCIe bus to the GPU and back can cost more than the computation saved. For small problems, a CPU wins. This is the first place the memory story in Part 3 enters: arithmetic is cheap, moving data is expensive.
A second-order view: what breaks at scale
Which assumption fails when you go from one GPU to a fleet? The comfortable assumption is that “more SMs means proportionally more speed”. It holds only while the workload keeps all SMs busy and fed. As chips grow, two things happen: the amount of independent work required to fill the chip rises, and the distance data travels rises as well. Both push the cost from arithmetic toward scheduling and memory, which is exactly why the sharing techniques in our Kubernetes GPU sharing guide exist: Multi-Instance GPU (MIG) partitions a physical GPU along SM and memory boundaries so that small jobs do not waste a big chip. And as memory bandwidth becomes the scarce resource, the economics shift towards memory suppliers, a trend we analyse in the HBM memory supercycle.
Practical Recommendations
Start with the question, not the hardware: is the work data-parallel, large enough, and fed fast enough? If any answer is no, the GPU will not help. If all three are yes, design from the top of the hierarchy down. Choose a grid large enough to give every SM several blocks, choose a block size that is a multiple of 32 (128 or 256 are sensible starting points), and keep per-thread resource use modest enough that several blocks fit on each SM.
Then measure rather than guess. NVIDIA’s profilers report achieved occupancy, warp stall reasons and branch divergence directly, and they will tell you which of the failure modes above you actually have. Optimise the one the profiler names, then re-measure; the bottleneck will move.
A short checklist before you ship a kernel:
- Block size is a multiple of 32 and no larger than 1,024.
- Grid has at least several times as many blocks as the GPU has SMs.
- Threads within a warp follow the same branch most of the time.
- Registers per thread and shared memory per block leave room for at least two to four resident blocks per SM.
- Neighbouring threads access neighbouring memory addresses.
- You know the arithmetic-to-memory ratio of the kernel, and you have checked it is not bound by memory (Part 3 explains how).
- The CPU to GPU transfer cost has been counted in the total, not ignored.
One-Page Cheat Sheet
Pin this to your monitor. The H100 column uses the verified numbers from earlier; the analogy column is the school.
| Level | School analogy | What it is | Size or limit (H100, CC 9.0) | Shares what | Typical mistake |
|---|---|---|---|---|---|
| Chip | Campus | Whole GPU die | 132 SMs, 50 MB L2, 80 GB HBM3 | L2 cache and global memory | Launching too few blocks to fill all SMs |
| SM | Department | Independent execution unit | 128 FP32 cores, 4 Tensor Cores, 64K registers, up to 228 KB shared | Registers and shared memory among resident blocks | Over-using registers and killing occupancy |
| Cluster | Connected departments | Hopper only, group of blocks | Co-scheduled on neighbouring SMs | Distributed shared memory | Assuming it exists on older GPUs |
| Thread block | Classroom | Group scheduled onto one SM | Up to 1,024 threads, 32 blocks per SM | Shared memory and barriers | Block size not a multiple of 32 |
| Warp | Row of 32 desks | Lockstep issue group | 32 threads, 64 warps per SM | One instruction stream | Divergent branches |
| Warp scheduler | Teacher | Picks a ready warp each cycle | 4 per SM | Issue slots | Too few warps to hide stalls |
| Thread | Student | One lane | Up to 255 registers, 2,048 per SM | Nothing, private registers | Register spilling to slow memory |
| Kernel | Homework | Function run by every thread | Launched as a grid of blocks | Instructions | Serial dependency inside it |
Two memory rules to carry into Part 2: registers are the fastest storage and private to a thread; shared memory is fast and private to a block; global memory is large and slow, and everything in a grid can reach it.
Conclusion: From the School to the Kernel
The whole post reduces to one sentence. A GPU is a campus of departments (SMs), each running classrooms (blocks) of rows (warps) of students (threads), with a teacher (the warp scheduler) who hides every delay by always finding a ready row. Understand the limits at each level (32 threads per warp, 1,024 per block, 2,048 and 64 warps per SM, 132 SMs on an H100) and you can predict, rather than guess, why a kernel is fast or slow.
The iPhone and the H100 share the vocabulary but not the scale: seven Apple GPU cores serve a camera pipeline in a power budget measured in watts, while 132 NVIDIA SMs serve model training at up to 700 W. The idea transfers; the numbers do not.
Hardware is only half the story. The hierarchy tells you where work can run, but not how to write code that keeps it busy. In Part 2, GPU kernel engineering: how a CUDA kernel executes and why data starvation kills performance, we write the homework: how a kernel maps to this hierarchy, and why lanes that wait for data produce the real bottleneck. Part 3 then goes to the physics: DRAM vs SRAM, FlashAttention and kernel fusion. Read them in order and the series becomes one continuous argument: the arithmetic is abundant, and moving data is the cost.
Frequently Asked Questions
What is the difference between a CPU and a GPU?
A CPU has a few powerful cores tuned to finish one task with minimum latency, using large caches and speculation. A GPU has thousands of simpler lanes grouped into SMs, tuned to finish a huge number of independent tasks per second by switching between warps whenever one stalls. CPUs excel at branching, serial logic and operating-system work; GPUs excel at data-parallel work such as graphics, image processing and matrix multiplication.
What is a streaming multiprocessor (SM)?
A streaming multiprocessor is the repeated building block of an NVIDIA GPU. Each SM has its own register file, on-chip shared memory, warp schedulers and arithmetic units, and it executes thread blocks independently of other SMs. A full H100 SXM5 has 132 SMs, each with 128 FP32 cores and four fourth-generation Tensor Cores, according to NVIDIA’s Hopper architecture documentation.
What is a CUDA warp and why is it 32 threads?
A CUDA warp is a group of 32 threads that the hardware issues one instruction to simultaneously. The size is fixed at 32 across NVIDIA architectures. Sharing one instruction fetch and decode across 32 lanes saves area and power, which lets a chip carry far more arithmetic units. The cost is warp divergence: if threads in a warp take different branches, the paths run one after another.
How many threads can an H100 run at once?
Each SM can hold up to 2,048 resident threads (64 warps), so a full H100 SXM5 with 132 SMs can keep 270,336 threads resident. Not all of them issue an instruction in every cycle; the warp schedulers pick ready warps. The extra resident threads exist so that the schedulers always have something to run while other warps wait for memory.
Is an Apple GPU core the same as an NVIDIA SM?
No. Apple calls its units GPU cores and has not published an SM-style specification. They are similar in role, as each is a repeated block with its own registers and on-chip memory, and Apple’s Metal model uses SIMD-groups of 32 threads like warps. But the architecture, memory system and power budget differ greatly, so core counts cannot be compared directly. The A20 Pro was reported with a seven-core GPU; the H100 has 132 SMs.
Why do GPUs have thread blocks as well as warps?
Warps are the hardware unit of instruction issue, while thread blocks are the programmer’s unit of cooperation. Threads in one block can share fast on-chip memory and synchronise at barriers, so a block is the largest group that can work closely together. Because each block is confined to one SM and independent of other blocks, the same program scales across GPUs with different numbers of SMs without change.
Further Reading
- Next in the series: GPU kernel engineering and CUDA kernel execution (Part 2)
- Series finale: GPU memory physics, DRAM vs SRAM and FlashAttention (Part 3)
- Kubernetes GPU sharing with MIG, time-slicing and MPS
- Micron and the HBM AI memory supercycle
- External: NVIDIA CUDA Programming Guide and NVIDIA Hopper architecture in depth
By Riju — about

Pingback: GPU Kernel Engineering: Why C = A + B Is Not Enough (Series