The Memory Wall in 2026: Optical Interconnect, LPDDR6 and Physical AI

The Memory Wall in 2026: Optical Interconnect, LPDDR6 and Physical AI

The Memory Wall in 2026: Optical Interconnect, LPDDR6 and Physical AI

Here is the uncomfortable arithmetic behind modern inference: to generate one token, a large language model must stream essentially all of its active weights from memory into the compute units, and then do very little arithmetic with each byte. The processor sits idle waiting on data. That gap, the memory wall AI workloads keep slamming into, is why a chip with petaflops of peak compute can still emit only a few dozen tokens per second for a single user.

On October 1, 2026 two announcements landed on the same day and attack that wall from opposite ends. Micron reported record fiscal fourth-quarter revenue of $54.23 billion and said it has begun sampling LPDDR6 to physical AI markets, while the startup Volantis raised an $88 million Series A to link memory to processors with light instead of copper. This post explains why bandwidth, not FLOPs, bounds so much of AI, what each announcement actually changes, and how to reason about the memory tiers you will be choosing between over the next three years.

What this covers: the roofline arithmetic of memory-bound inference, the HBM capacity ceiling, why LPDDR6 matters for robots and edge devices, how an optical memory fabric is meant to work, and a practical framework for choosing a memory tier.

Context and Background

The phrase “memory wall” dates to a 1995 paper by Wulf and McKee, who observed that processor speed was improving far faster than DRAM access speed and that memory would eventually dominate performance. For decades caches and out-of-order execution hid the problem. AI inference removed the hiding places. A transformer decoder reads every active weight once per generated token, and unless many requests share that read, each byte fetched supports only a couple of floating point operations.

Architects responded with High Bandwidth Memory (HBM). HBM stacks DRAM dies vertically, connects them with through-silicon vias, and places the stacks beside the processor on a silicon interposer. The JEDEC HBM4 standard, published in April 2025, widens the interface to 2,048 bits, specifies speeds up to 8 Gb/s per pin and up to 2 TB/s per stack, and doubles independent channels to 32 per stack. Stacks of up to 16 dies with 32 Gb DRAM reach a stated maximum of 64 GB per cube. We covered the commercial side of that boom in our analysis of the Micron HBM AI memory supercycle.

The market has noticed. Micron’s fiscal 2026 results, reported October 1, show fourth-quarter revenue of $54.23 billion against $11.32 billion a year earlier, and full-year revenue of $133.19 billion against $37.38 billion for fiscal 2025. Guidance for the next quarter is $61.5 billion, plus or minus $1.5 billion. Those are company-reported figures from the Micron fiscal 2026 earnings release. Revenue growth of that magnitude is a market price on the scarcity of memory bandwidth and capacity.

Two structural pressures keep the wall tall. First, model sizes and context lengths grow faster than per-stack capacity. Second, the physical perimeter of a processor die, its “shoreline”, limits how many HBM stacks can touch it. Both pressures push toward new memory tiers, and that is the lens for the rest of this post.

Why Bandwidth, Not Compute, Bounds Inference

Direct answer: autoregressive decoding is memory-bound because each generated token requires reading all active model weights (plus the key-value cache) from memory, performing only about two operations per parameter. Token rate is therefore roughly memory bandwidth divided by bytes read per token, until batching raises arithmetic intensity enough to become compute-bound.

Memory wall AI roofline: model weights streamed from memory each token leave compute units idle and cap tokens per second

Figure 1: The memory-bound inference loop. Weights are read from memory for every token, so bytes moved per second, not FLOPs, set the token rate at low batch sizes.

The diagram shows the causal chain. Weights live in memory. Each decode step pulls them across the memory interface. The interface delivers a fixed number of bytes per second, and at low batch size the compute units finish their small share of work and wait. Tokens per second is set by the bytes, not the math.

The roofline in four lines of arithmetic

The roofline model, introduced by Williams, Waterman and Patterson in 2009, bounds attainable performance by the lesser of peak compute and the product of memory bandwidth and arithmetic intensity (operations per byte moved). The numbers below are illustrative, first-principles arithmetic for a hypothetical accelerator, not measurements of any product.

Take a model with 70 billion active parameters stored at 8-bit precision. That is 70 GB of weights. Suppose the accelerator delivers 8 TB/s of memory bandwidth. The upper bound on single-stream decode is 8,000 GB/s divided by 70 GB, about 114 tokens per second, before counting the KV cache and any overhead. Real systems land well below that bound, because achieved bandwidth is rarely the peak figure.

Now count compute. Each token needs about 2 operations per parameter, so 140 billion operations. If that same accelerator has 2,000 TFLOPS at 8-bit, compute alone would allow 2,000 trillion divided by 140 billion, about 14,000 tokens per second. Compute headroom exceeds the bandwidth bound by roughly two orders of magnitude. The processor is idle more than 99 percent of the time at batch one in this illustration.

Batching rescues throughput but not latency

Batching amortizes the weight read. Serving 64 requests per step reads the weights once but produces 64 tokens, raising arithmetic intensity 64-fold. Aggregate throughput climbs toward the compute roof, which is why datacenter operators run large batches and why the economics in AI inference cost optimization revolve around keeping batches full.

Batching has two limits. A single user still sees per-token latency set by the bandwidth bound, so interactive agents and voice systems feel the wall directly. And every request in the batch brings its own KV cache, which grows with context length and competes for the same capacity and bandwidth. Long-context agent workloads, as discussed in AI agent memory systems, push the bytes-per-token number up, not down.

The KV cache is the second wall

The KV cache stores per-layer keys and values for every previous token so they need not be recomputed. Its size is roughly 2 times layers times KV heads times head dimension times bytes per element, per token. For a hypothetical model with 80 layers, 8 KV heads of dimension 128 and 16-bit values, that is 2 x 80 x 8 x 128 x 2 bytes, about 327 KB per token (illustrative). A 128,000-token context then holds roughly 42 GB of cache for one sequence. Ten concurrent long sequences would hold more than the weights themselves.

Grouped-query attention, multi-query attention, quantized caches and paged allocation all reduce this footprint, and they exist precisely because the cache pressures both capacity and bandwidth. This is the first original point of this post: the memory wall has two dimensions, bandwidth and capacity, and the industry’s three current responses, HBM, LPDDR6 and optical memory fabrics, trade between those dimensions in different ways.

The HBM Capacity and Shoreline Ceiling

HBM is the incumbent answer and it is excellent at what it does. A single HBM4 stack can, per the JEDEC specification, deliver up to 2 TB/s across a 2,048-bit interface, and it sits millimeters from the compute die. The problem is not that HBM is slow. The problem is how few stacks a processor can attach.

Memory wall AI constraint: a processor die with limited shoreline connects only a handful of HBM stacks through the interposer

Figure 2: The shoreline and capacity ceiling. Only a few HBM stacks fit around one die, so total capacity and bandwidth scale with perimeter, not with demand.

Three coupled limits

Shoreline. HBM connects through a very wide, short-reach parallel interface. Each stack consumes a stretch of the die edge, and the interposer has a maximum size set by lithography reticle limits and packaging yield. The practical result is that leading accelerators carry a small number of stacks. Volantis, in its launch coverage, contrasted its design with the eight memory chips around Nvidia’s best GPU today; that comparison is the company’s framing, but the order of magnitude, single-digit stacks per package, matches what the industry ships.

Capacity per stack. Even at the 64 GB per-cube maximum in the HBM4 standard, eight stacks give 512 GB at most, and shipping parts typically use smaller configurations. A 20-trillion-parameter model at 8 bits needs 20 TB for weights alone, which means dozens of packages just to hold the parameters, before KV cache. Capacity therefore forces model sharding across many devices, and sharding moves the bottleneck from memory to the network.

Cost and supply. HBM requires stacked dies, TSV processing, advanced packaging and known-good-die testing. Capacity of the whole supply chain, not only the DRAM fabs, is constrained. This is part of why memory vendors report the revenue growth described above, and why buyers sign multi-year supply commitments.

Why more bandwidth per stack is not enough

HBM generations double interface width or raise pin speed, and each generation helps. But accelerator compute grows faster. Bytes of memory bandwidth per FLOP has fallen across recent accelerator generations, which is the roofline restated: the ridge point, the arithmetic intensity at which a workload becomes compute-bound, moves to the right. Workloads that were balanced a few years ago become memory-bound on new hardware without any change in the software.

The SRAM alternative and its own wall

A different family of designs sidesteps DRAM by keeping weights in on-chip SRAM across many chips. SRAM offers enormous bandwidth and low latency but very limited capacity per square millimeter, so serving a large model needs many chips wired together, and the interconnect between them becomes the constraint. Volantis explicitly positions itself between these poles: more capacity than SRAM-centric approaches, more attachable memory than HBM. That claim is a company target, and we will treat it as such below.

LPDDR6: The Other Direction, Down to the Edge

While HBM climbs toward the datacenter, a quieter evolution is happening in low-power DRAM. LPDDR was built for phones and has become the memory of choice for laptops, automotive systems, and a growing class of embedded AI devices. Micron’s fiscal fourth-quarter release states: “We began sampling 1-gamma based LPDDR6 products to multiple physical AI markets.” That sentence is the entire official claim in the release; it gives no bandwidth, capacity or customer detail, so everything quantitative below comes from the JEDEC standard and press coverage of it.

What the LPDDR6 standard changes

JEDEC published the LPDDR6 standard (JESD209-6) on July 10, 2025. According to press coverage of the specification, data rates run from 10,667 to 14,400 MT/s, and the channel organization moves to 24-bit sub-channels, which lowers latency and improves concurrency relative to the 16-bit organization of LPDDR5. The standard lowers operating voltage and adds Dynamic Voltage Frequency Scaling for Low Power (DVFSL) so that devices save energy at low-frequency operation. Qualcomm, MediaTek, Samsung, Micron and SK hynix were named as supporters at publication. See the Tom’s Hardware report on the LPDDR6 standard for the summarized parameters.

A quick bandwidth illustration, which is arithmetic and not a product specification: a hypothetical 192-bit-wide LPDDR6 system running at 10,667 MT/s moves 192 x 10,667 / 8, about 256 GB/s. A single HBM4 stack at the 2 TB/s ceiling delivers roughly eight times that. LPDDR6 will not match HBM on bandwidth per package. Its argument is energy and cost per gigabyte, and the ability to attach a lot of it on a conventional board.

Why physical AI wants LPDDR

Physical AI means models that perceive and act in the world: robots, autonomous machines, drones, industrial vision systems and vehicles. Their memory problem differs from the datacenter’s in four ways.

  • Power envelope. A mobile robot or vehicle computer runs on a battery or a constrained thermal budget, so joules per bit matter more than peak bandwidth.
  • Latency over throughput. A control loop serves one stream, batch one, which is exactly the memory-bound regime of Figure 1. Bandwidth per watt is the figure of merit.
  • Capacity on a board. Vision-language-action models, world models and perception stacks hold several billion parameters plus sensor buffers, which fit in tens of gigabytes of LPDDR but cannot justify an HBM package.
  • Cost and supply. Automotive and industrial programs need multi-year supply and qualified parts, not the scarcest memory in the world.

This connects to the digital-twin and robotics themes this site covers elsewhere. A simulation-trained policy deployed onto an edge computer inherits the batch-one memory wall, so a robot’s reaction time is bounded by how quickly its weights can stream from LPDDR, not by its TOPS rating. Reading a vendor’s “TOPS” number without the memory bandwidth beside it tells you almost nothing about decode or policy-inference speed.

Edge-side mitigations that work today

Before buying new memory, there are software levers: weight quantization to 4 bits cuts bytes per token roughly in half relative to 8 bits, mixture-of-experts models read only the active experts, speculative decoding verifies several draft tokens per weight read, and small distilled policies fit in on-chip caches. Each lever raises tokens per byte. The hardware tiers below raise bytes per second, and a good system uses both.

Optical Memory Fabrics: What Volantis Is Actually Claiming

The third response is the most speculative and the most architecturally interesting. On October 1, 2026, Volantis announced an $88 million Series A, bringing total funding to $97 million including seed, according to Runtime Wire’s coverage. Co-leads were Lachy Groom and Abstract Ventures, with participation from John Doerr, VXI Capital, Triatomic and Susa Ventures. The CEO is Tapabrata Ghosh, and co-founder and CTO Roy Meade previously led Micron’s HBM program.

Optical interconnect sequence: a processor die sends a read request through an electro-optical driver and waveguide to a memory chip using VCSEL light

Figure 3: The conceptual optical memory path. Electrical requests become light, cross the interposer in a waveguide, and return as data. The component labels are generic; Volantis has not published its circuit details.

The reported claims

Coverage attributes the following to the company, and all of it should be read as targets rather than demonstrated results. The A-1 system targets models exceeding 20 trillion parameters and up to 10,000 tokens per second per user. The links use optical waveguides and custom micro-VCSELs (vertical-cavity surface-emitting lasers). They are described as spanning more than 200 millimeters across an interposer at under one picojoule per bit. The design is reported as taped out, with first customer deliveries planned for 2027. The headline figure is up to 220 memory chips around a processor, versus eight for Nvidia’s best today.

I could not find an independent measurement of any of these numbers, and no customer results have been reported. Treat the 10,000 tokens-per-second figure in particular as a design target whose assumptions (model sparsity, precision, batch, speculative decoding) were not detailed in the coverage I reviewed.

Why light helps the shoreline problem

Electrical links trade distance against bandwidth density and energy. A short parallel electrical bus, as in HBM, is dense and efficient but cannot reach far. Longer copper traces need equalization and burn energy that grows with distance. Optical links are nearly distance-insensitive over interposer and package scales, so memory can sit farther from the die without paying a steep energy penalty. If the optical I/O is compact enough, a die edge can host far more bandwidth than the same edge hosts as HBM PHY, and memory chips no longer have to hug the processor.

VCSELs are a known technology, widely used in short-reach data center optics and in consumer sensing. Their appeal here is low threshold current, easy two-dimensional array fabrication, and efficient direct modulation at short reach. The open engineering questions are the ones that have historically slowed co-packaged optics: thermal management (lasers dislike heat and processors produce plenty), alignment and packaging yield, reliability over a multi-year lifetime, and the cost of adding electro-optical conversion at both ends of every link.

What 220 chips would mean

Capacity is the reason the number matters. If each memory chip stored a few gigabytes, 220 of them would hold on the order of a terabyte per processor (illustrative, since Volantis has not disclosed the chip type or density in the sources I read). That is the right scale to hold a 20-trillion-parameter model at low precision within a small number of packages, instead of spreading it across dozens. Fewer packages means fewer network hops during decode, which is where multi-device sharding loses a great deal of its time.

The second consequence is bandwidth aggregation. If aggregate bandwidth scales with the number of optical lanes and memory chips, a processor could in principle read much more per second than eight HBM stacks allow. Whether that aggregate arrives at the processor depends on the on-die network that fans lanes into compute, and on whether commodity-class DRAM dies can keep the lanes busy. Both are unproven at scale.

Where this sits against an existing roadmap

Optical interconnect is not new to AI systems. Optical links already connect racks and, increasingly, switches through co-packaged optics. Companies including Ayar Labs, where Meade previously served as vice president according to the coverage, have pursued optical I/O for chip-to-chip links. What distinguishes the Volantis pitch is aiming the technology at the memory interface specifically and reorganizing the system around many memory chips instead of a few stacks. For the broader hardware market context, see our Micron HBM supercycle analysis, which explains why memory vendors currently hold pricing power.

A Decision Framework: Matching Memory Tier to Workload

The three responses are not competitors in most deployments. They occupy different points on a bandwidth, capacity, energy and cost plane, and a system architect picks per workload.

Memory tier selection for AI: HBM for bandwidth, LPDDR6 for energy at the edge, optical memory fabrics for capacity, feeding a workload-based tier choice

Figure 4: Three tiers, three optimization targets. The right choice depends on model size, latency target and power budget rather than on a single winner.

Dimension HBM3E / HBM4 LPDDR6 Optical memory fabric (Volantis-style)
Primary strength Bandwidth density next to the die Energy and cost per GB Capacity per processor
Standard status JEDEC HBM4 published April 2025 JEDEC JESD209-6 published July 2025 Proprietary, first deliveries targeted for 2027
Typical home Datacenter accelerators Edge, automotive, robotics, laptops Large-model inference (target)
Main risk Shoreline, supply, cost Lower bandwidth per package Unproven at scale, packaging and thermals
Maturity Shipping in volume Sampling to customers Design reported as taped out

Worked sizing examples

Case A: a voice assistant on a robot. A 7-billion-parameter model at 4 bits is about 3.5 GB. With an illustrative 100 GB/s of effective LPDDR bandwidth, the bandwidth bound is roughly 28 tokens per second at batch one, which is adequate for interactive speech. Cost and watts dominate, and LPDDR is the natural tier.

Case B: a 400-billion-parameter MoE with 40 billion active parameters in the cloud. At 8 bits the active weights are 40 GB per token. Eight TB/s gives a bound near 200 tokens per second per replica before cache and overhead. Capacity, not bandwidth, forces the sharding: the full 400 GB must reside somewhere, and HBM nodes are the proven path.

Case C: a multi-trillion-parameter sparse model with long agent contexts. Here capacity and KV footprint dominate everything. If an optical fabric delivers on its claims, it addresses exactly this regime. If it does not, the same workload is served by large HBM clusters at higher network cost.

A note on confidential workloads

Memory tiers also change the security boundary. Encrypting memory traffic is straightforward when the link is one on-package bus and becomes more complicated when memory is distributed across many chips and optical paths. If you plan confidential inference, read our guide to confidential AI inference with GPU trusted execution environments and ask any memory-fabric vendor how encryption, integrity protection and attestation extend across their links. I did not find published answers for Volantis on this point.

Measuring the Wall in Your Own System

Vendor claims are cheap and profiling is not. Before committing to a hardware tier, measure whether your workload is on the bandwidth slope or the compute plateau of the roofline. The method is simple enough to run in an afternoon.

Step 1: compute the theoretical bound

For decode, divide achievable memory bandwidth by bytes read per token. Bytes per token is active parameter count times bytes per parameter, plus the KV cache bytes read for the current context. Use measured bandwidth from a streaming benchmark, not the datasheet peak, because achieved bandwidth on real kernels commonly falls short of peak. If your observed token rate is within a modest fraction of this bound, the system is memory-bound and more FLOPs will not help.

Step 2: separate prefill from decode

Prefill processes the prompt in parallel and has high arithmetic intensity, so it tends to be compute-bound. Decode is the memory-bound phase. Many production stacks now run the two phases on different pools of hardware, which is a direct architectural response to the wall. If your traffic is dominated by long prompts and short answers, you are buying compute; if it is short prompts and long generations, such as agents and reasoning models, you are buying bandwidth.

Step 3: price the byte, not the FLOP

A useful internal metric is cost per terabyte moved per second, rather than cost per peak TFLOPS. Compare it across tiers. HBM wins on bandwidth per package but carries a high price per gigabyte. LPDDR is cheaper per gigabyte and lower per package bandwidth. For a latency-insensitive batch workload, buying many cheap devices may beat buying few expensive ones; for a batch-one control loop, the cheapest watt per byte wins. The practical guidance in AI inference cost optimization applies directly: raise tokens per byte first, then buy bytes per second.

Step 4: model the tail of the context distribution

Average context length hides the cases that break capacity planning. Plot the distribution of context lengths in production and size memory for the 95th or 99th percentile, because those requests hold the largest KV caches and decide how many concurrent sessions fit. Techniques such as KV cache quantization, sliding-window attention, and offloading cold cache to a slower tier all trade a little quality or latency for a lot of capacity. A slower, larger tier for cold cache is exactly where a fabric with many chips would slot in, if it arrives.

What to ask a memory-fabric vendor

Because optical memory fabrics are early, a buyer should ask specific questions. What is the sustained, not peak, bandwidth to a single compute die, and how does it vary with access pattern? What is the loaded read latency compared with HBM, since pointer-chasing and small reads suffer from added conversion stages? What is the failure model: if one of 220 chips or one laser fails, does the system degrade gracefully or lose the package? What is the error-correction scheme across the optical hop? How is refresh and thermal throttling handled when lasers and DRAM share a package? Which software stack, compilers and kernels, maps tensors onto a many-chip memory pool? Honest answers to these questions separate a platform from a prototype.

The system-level view: where tokens per second actually come from

It helps to decompose a quoted tokens-per-second figure into its multiplicative parts. Single-user tokens per second equals effective bandwidth divided by effective bytes per token. Effective bandwidth is peak bandwidth times utilization. Effective bytes per token is the model’s active footprint divided by whatever the software saves through sparsity, quantization and speculation. When a startup targets 10,000 tokens per second per user, the claim could be met by a very large bandwidth, by a very small effective footprint, or by a speculative-decoding multiplier, and the three have very different credibility. A target stated without its model, precision and batch assumptions cannot be audited. That is not a criticism of any one company; it is a reminder to ask for the denominator.

Consider the arithmetic with illustrative numbers. If a sparse model activates 100 billion parameters at 4 bits, each token reads 50 GB. A 10,000 tokens-per-second stream would then need 500 TB/s of effective weight bandwidth, far beyond any single package today, unless speculative decoding verifies many tokens per weight read or the active footprint is far smaller than assumed. Whatever the company’s actual assumptions are, a statement of this kind implies aggressive combinations of sparsity, speculation and aggregate bandwidth. It is exactly the type of claim whose first public benchmark in 2027 will be more informative than the announcement.

Trade-offs, Gotchas, and What Goes Wrong

Optical is not free energy. The reported figure of under one picojoule per bit is attractive, but system energy includes laser bias, driver and receiver circuits, serialization, and thermal control. Whether a headline per-bit figure holds at the wall plug under real traffic is an open question until measured. Likewise, adding conversion stages adds latency, and latency-sensitive random reads can lose what bandwidth gains.

More chips means more failure surface. The probability that at least one of 220 memory chips or its optical lane fails in a given period is much higher than for eight stacks, even if each chip is individually reliable. Operating such a system requires redundancy, sparing and error handling designed in from the start. Datacenter operators who have lived with HBM reliability and yield issues will want field data before trusting a first-generation fabric.

Packaging yield compounds. Advanced packaging yield multiplies across components. A package containing a processor, optical engines and hundreds of memory die is a large yield-risk object, and a single bad element can scrap an expensive assembly unless known-good-die testing and repair are strong.

Software maturity lags hardware. Frameworks assume a flat, fast, small memory with a clear hierarchy. A pool of hundreds of chips with non-uniform access patterns needs new allocators, compilers and scheduling. New hardware with immature tooling often delivers a fraction of its theoretical performance in year one.

LPDDR6 is not a miracle. Sampling is a long way from volume shipment in automotive or robotics programs, which require qualification cycles. The Micron release provides no performance numbers for its LPDDR6 parts, and the standard’s maximum data rates will not all be available on day one. Plan edge designs around parts you can actually procure, and keep a fallback to LPDDR5X.

Do not mistake funding for validation. An $88 million round and a reported tape-out are meaningful signals about investor conviction and engineering progress. They are not benchmark results. The same caution applies to the revenue numbers: Micron’s record quarter shows demand and pricing power, which tells you the wall is expensive, not that any particular solution will break through it. This post does not offer investment advice, and nothing here should be read as a recommendation to buy or sell any security.

Software can shrink the wall faster than hardware. Quantization to 4 bits and below, sparse and mixture-of-experts architectures, multi-token prediction and speculative decoding have each delivered large effective gains at no hardware cost. A decision to wait for new memory hardware can be wrong if a software change on current hardware delivers the same improvement sooner.

Practical Recommendations

Start from the workload, not the headline. If you serve interactive agents or voice, you live in the batch-one regime and bandwidth per user is your metric. If you serve high-volume batch jobs, capacity and cost per token dominate and batching hides much of the wall.

For datacenter inference planning through 2027, keep HBM-based systems as the baseline, design your serving layer to separate prefill and decode pools, and treat optical memory fabrics as an option to evaluate with a pilot when silicon is available, not as a roadmap dependency. For edge and robotics designs, evaluate LPDDR6 availability and qualification timelines now, but build the first product on parts you can source today and abstract the memory interface in your software.

A short checklist:

  • Compute the bandwidth-bound token rate for your model on your hardware and compare it with your observed rate.
  • Profile prefill and decode separately and place them on hardware that suits each.
  • Size memory for the 99th percentile context length, not the mean.
  • Apply quantization, caching and speculation before buying more bandwidth.
  • Ask memory-fabric vendors for sustained bandwidth, loaded latency, failure model and encryption story.
  • Treat every startup performance figure as a target until an independent benchmark exists.
  • Keep a fallback memory tier in every edge design.

Frequently Asked Questions

What is the memory wall in AI?

The memory wall is the widening gap between how fast processors can compute and how fast memory can supply data. In AI inference, generating each token requires reading most of the model’s active weights, so token rate is limited by memory bandwidth divided by bytes read, not by peak FLOPs. The term originates from a 1995 paper by Wulf and McKee and has become central to large-model serving.

Why is LLM inference memory-bound instead of compute-bound?

At low batch sizes, each weight fetched from memory is used for only about two operations, so arithmetic intensity is very low compared with the thousands of operations per byte that modern accelerators can sustain. Compute units finish quickly and wait for data. Large batches reuse each weight across many requests, which is why batching moves workloads toward the compute-bound regime, though it cannot lower single-user latency.

What is LPDDR6 and why does it matter for physical AI?

LPDDR6 is the JEDEC low-power DRAM standard JESD209-6, published July 2025, with reported data rates of 10,667 to 14,400 MT/s and a new 24-bit sub-channel organization. Micron says it has begun sampling 1-gamma LPDDR6 to physical AI markets. It matters because robots and edge machines run batch-one inference under tight power budgets, where energy and cost per gigabyte count more than peak package bandwidth.

How does optical interconnect help with the memory wall?

Optical links carry data as light, which loses far less energy over package-scale distances than electrical traces. That lets memory sit farther from the processor and lets more memory chips attach than the die edge permits with HBM. Volantis reports a design using micro-VCSELs and waveguides targeting under one picojoule per bit and up to 220 memory chips, with deliveries planned for 2027. These are company targets, not independent measurements.

Will optical memory replace HBM?

Unlikely in the near term. HBM is standardized, shipping in volume and supported by mature packaging and software. Optical memory fabrics are early, with reported first deliveries in 2027, and face open questions on thermals, packaging yield, reliability and tooling. A more plausible outcome is coexistence: HBM for bandwidth-critical working sets, and a larger, slower tier for capacity if optical fabrics prove out. This is an informed opinion, not a forecast.

How can I reduce memory-bound latency without new hardware?

Raise tokens per byte. Quantize weights to 4 bits or lower where quality allows, use mixture-of-experts models so only active experts are read, apply speculative decoding to verify several tokens per weight read, quantize and page the KV cache, and separate prefill from decode on different hardware pools. Each technique cuts bytes moved per token, which directly lifts the bandwidth-bound token rate.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *