Mamba State Space Models vs Transformers: How Hybrid Architectures Actually Work
Every token a Transformer generates makes the next one more expensive. The key-value (KV) cache grows linearly with context, attention reads all of it on every step, and at 128k tokens the cache for a 7B-class model can outweigh the model weights several times over. Mamba state space models attack that problem at the root: they compress the whole past into a fixed-size state, so memory per generated token stays constant no matter how long the context gets.
That sounds like a free lunch, and it is not. A fixed-size state cannot hold arbitrary detail, so pure SSMs lose to attention on exact recall and in-context copying. The interesting engineering result of 2024 and 2025 is that you rarely have to choose: hybrids that keep a few attention layers inside a mostly-SSM stack recover most of the quality at a fraction of the memory.
This article gives you the mechanism (S4, then Mamba, then Mamba-2), worked memory arithmetic for a 7B-class model at 128k context, the evidence on where SSMs fail, how hybrids like Jamba are built, and a decision matrix for choosing.
What this covers: the S4 to Mamba to Mamba-2 lineage, the selective scan and why it needs a custom kernel, KV-cache versus constant-state arithmetic, the recall weakness, hybrid designs, a runnable reference implementation, and when to pick which architecture.
Context and Background
For most of the 2017 to 2023 period, “sequence model” meant “Transformer”. Self-attention gives every token direct access to every earlier token, trains in parallel across the sequence, and scales well on GPUs. Its price is quadratic compute in sequence length during training and prefill, and a linear-in-context cache during decoding.
The cache cost is not abstract. If you want to see exactly why decoding is limited by memory bandwidth rather than arithmetic, our piece on GPU memory physics, DRAM versus SRAM and FlashAttention kernel fusion walks through the hardware side. The short version is that each decoded token must stream the entire KV cache from high-bandwidth memory, so cache size translates directly into latency and into how many concurrent requests fit on a GPU.
Researchers have long tried to escape this. Linear attention replaces the softmax with a kernel feature map so that attention can be computed as a recurrence with a fixed-size state. Recurrent networks and gated convolutions made similar bets. The Mamba paper by Albert Gu and Tri Dao states the diagnosis plainly: these subquadratic models were faster but trailed attention on language because they lacked content-based reasoning, meaning the ability to decide what to remember based on what the input actually is (Gu and Dao, arXiv:2312.00752).
State space models come from a different tradition, continuous-time control theory. A linear SSM maps an input signal to an output through a latent state governed by a matrix equation. The deep-learning version, introduced as S4 (Structured State Spaces for Sequence Modeling), made that idea trainable at scale. Per its authors, S4 was the first model to solve the Path-X task in the Long Range Arena benchmark at sequence length 16,384, where prior models performed at chance (Gu et al., arXiv:2111.00396).
S4 was strong on long-range signals but still trailed Transformers on language perplexity. Mamba closed most of that gap, Mamba-2 made it faster to train, and hybrids made it practical. That is the arc this article follows.
Why the cache, not the FLOPs, is the real bottleneck
Training cost gets the headlines, but most deployed LLM spend is inference. During decoding, arithmetic intensity is low: each step multiplies a handful of vectors against large matrices and a large cache. The GPU spends its time waiting on memory.
Two consequences follow. First, anything that shrinks per-request state lets you batch more requests per GPU, which directly lowers cost per token. Second, long-context workloads such as repository-level code analysis, long-document question answering, and agent loops with growing transcripts are exactly where the cache dominates.
This is the economic case for SSMs. It has nothing to do with beating Transformers on a leaderboard and everything to do with the memory a request occupies while it is alive.
How Mamba State Space Models Work: From S4 to Selective Scan
Short answer: a state space model keeps a fixed-size hidden state h that is updated once per token by a linear recurrence, then reads the output from that state. Mamba makes the update parameters depend on the current token (the selective state space), so the model can choose what to write into and forget from its state. A custom parallel scan makes this fast on GPUs.

Figure 1: One step of a selective SSM. B, C and the step size are computed from the current input, which is what separates Mamba from S4.
Figure 1 shows the per-token data flow. The long-description version: a token embedding is projected into the SSM input, and three input-dependent quantities (B, C and a step size delta) are computed from it. The step size discretizes the continuous matrices A and B, the state is updated, and the output is read out through C before a gate and an output projection.
The continuous system and its discretization
A linear SSM is defined by h'(t) = A h(t) + B x(t) and y(t) = C h(t). To process discrete tokens you pick a step size delta and discretize, typically with zero-order hold, giving h_t = A_bar h_(t-1) + B_bar x_t and y_t = C h_t. The bars denote the discretized matrices.
In S4, A, B, C and delta are fixed after training. That makes the whole system linear time-invariant (LTI). An LTI system can be unrolled into a single long convolution kernel, so training uses FFT-based convolution in parallel while inference switches to the recurrence with constant memory and constant work per step. The S4 authors report roughly 60 times faster generation than a vanilla Transformer on CIFAR-10 and WikiText-103, while noting that the language model still trailed in perplexity.
S4’s other ingredient was initialization. The A matrix is initialized from the HiPPO family, which is constructed so the state approximates a running projection of the input history onto orthogonal polynomials. The paper cites earlier work showing that swapping a random matrix for a HiPPO matrix lifted sequential MNIST accuracy from 60 percent to 98 percent.
Why time-invariance is the limitation
An LTI system treats every token identically. The same A and B apply whether the token is a crucial name or a stray comma. It cannot ignore noise or hold on to a particular fact, because the dynamics do not depend on content.
Gu and Dao illustrate this with two synthetic tasks. Selective copying requires remembering specific tokens while skipping filler at variable positions. Induction heads require completing a pattern seen earlier in context. LTI models cannot solve either well, since both need the update rule to vary with the input.
The selective state space
Mamba’s central change is to make B, C and delta functions of the input x_t. A large delta makes the model focus on the current token and overwrite state, while a small delta makes it largely ignore the token and preserve the existing state. This acts like a learned gate, and the connection to the gating in RNNs is deliberate.
The price is that the system is now time-varying, so the convolution trick no longer applies. The model must be computed as a recurrence, and a naive sequential loop over thousands of tokens is hopelessly slow on a GPU. This is the engineering problem the rest of the Mamba paper solves.
The hardware-aware scan
The recurrence h_t = a_t h_(t-1) + b_t is associative when written as a composition of affine maps, so it can be computed with a parallel prefix scan in logarithmic depth. Mamba combines that with the same memory-hierarchy discipline used by FlashAttention. The expanded state, with shape batch by length by channels by state size, is never materialized in slow high-bandwidth memory. Instead the kernel loads the compact inputs into fast on-chip SRAM, performs discretization and the scan there, and writes only the final outputs back.
The backward pass recomputes intermediate states instead of storing them, trading cheap arithmetic for scarce memory. The authors report this fused kernel is the reason the model is practical: the paper claims inference throughput about five times that of Transformers and linear scaling in sequence length, with quality still improving on sequences up to a million tokens (arXiv:2312.00752). Those are the authors’ claims on their evaluated setups, not a guarantee for your serving stack.
The Mamba block
Mamba also simplified the network. A standard Transformer layer alternates an attention sub-layer with an MLP sub-layer. The Mamba block merges these ideas into one homogeneous unit: an input projection that expands the width, a short causal depthwise convolution, a nonlinearity, the selective SSM, and a multiplicative gate branch, followed by an output projection. Stacking the same block repeatedly gives the full model with no attention and no separate MLP.
The paper reports that Mamba-3B outperforms Transformers of the same size and matches Transformers twice its size on pretraining perplexity and downstream evaluation. Treat this as a result at the scales tested, roughly up to the low billions of parameters, rather than proof that pure SSMs win at frontier scale.
From Mamba to Mamba-2: State Space Duality
Mamba’s selective scan is fast relative to a naive implementation, but it still does not use tensor cores, the matrix-multiply units that provide most of a modern GPU’s throughput. Mamba-2 was designed to fix that.

Figure 2: Lineage of the architectures. Mamba-2 sits between linear attention and SSMs, which is what the duality result formalizes.
The Mamba-2 paper, “Transformers are SSMs”, argues that SSMs and attention variants are closely related, connected through different decompositions of structured semiseparable matrices, a result the authors call state space duality (SSD) (Dao and Gu, arXiv:2405.21060). The practical outcome is a refined core layer that the abstract reports as 2 to 8 times faster than Mamba’s selective SSM while remaining competitive with Transformers on language modeling. The abstract does not give benchmark conditions, so treat that range as an upper-level claim.
What the duality means in practice
The same sequence transformation can be computed two ways. As a recurrence, it costs linear time with a small state, ideal for decoding. As a masked matrix multiplication, it looks like attention with a structured mask, which maps onto tensor cores and suits training and prefill.
Mamba-2 uses a chunked algorithm that blends both: matrix multiplication inside each chunk, and a recurrence passing state between chunks. To make the matmul form possible, Mamba-2 restricts the A matrix to a scalar times identity per head, a simplification relative to Mamba’s richer structure. That restriction is what buys the larger state sizes and the speed.
Why duality matters beyond speed
It reframes linear attention and SSMs as one family. Ideas that matured on the Transformer side, such as multi-head structure, tensor parallelism, and sequence parallelism, transfer to SSMs more cleanly. It also explains hybrids: if both layer types are points in one design space, interleaving them is less exotic than it first appears.
Deeper Analysis: Memory Arithmetic, Recall, and Where SSMs Lose
Claims about linear-time inference are easy to make and easy to misread. This section puts numbers on the memory difference and then looks at the evidence for where pure SSMs fall short.

Figure 3: Decode-time memory behavior. Transformer cache grows with every token, while the SSM state is overwritten in place.
Worked example: a 7B-class model at 128k context
The following numbers are illustrative. They use a generic 7B-class configuration, not any specific released model, with 16-bit cache values and a single sequence.
Transformer with full multi-head attention. Assume 32 layers, 32 KV heads, head dimension 128. Per token, the cache stores keys and values: 2 x 32 layers x 32 heads x 128 dims x 2 bytes = 524,288 bytes, or 0.5 MiB. At 131,072 tokens that is 64 GiB for a single sequence, against roughly 13 to 14 GiB for 7B weights in 16-bit.
Transformer with grouped-query attention (GQA). Reduce to 8 KV heads. Per token cost drops to 2 x 32 x 8 x 128 x 2 = 131,072 bytes, 128 KiB. At 128k tokens that is 16 GiB per sequence. GQA is the standard mitigation and already makes long context affordable, but the cost still scales linearly with length and with concurrent users.
Pure SSM (Mamba style). Assume model width 4,096, expansion factor 2 (inner width 8,192), state size N = 16, and 64 Mamba blocks (the usual way to match a 32-layer Transformer’s parameter count). The recurrent state per block is 8,192 x 16 = 131,072 values. At 4 bytes (SSM states are often kept in higher precision) that is 512 KiB per block, 32 MiB across 64 blocks. Add a small convolution state (an inner width times a kernel of 4, a few hundred KiB per block at most) and the total is on the order of tens of MiB. Crucially it is the same at 1k or 1M tokens.
Hybrid with one attention layer in eight. Of 32 layers, 4 are attention (GQA with 8 KV heads) and 28 are SSM. Cache per token is 4/32 of the GQA figure, so 16 KiB, and 128k tokens costs 2 GiB. The SSM state adds tens of MiB. That is an 8-fold reduction versus the GQA Transformer and 32-fold versus full multi-head attention.
| Architecture (7B-class, illustrative) | Per-token cache | At 128k tokens, one sequence |
|---|---|---|
| Full MHA Transformer | 0.5 MiB | 64 GiB |
| GQA Transformer, 8 KV heads | 128 KiB | 16 GiB |
| Hybrid, 1 attention in 8, GQA | 16 KiB | about 2 GiB plus tens of MiB state |
| Pure SSM | none | tens of MiB, constant |
A sanity check against published data: the Jamba paper reports a KV cache at 256k tokens in 16-bit of 4 GB for Jamba, versus 32 GB for Mistral and Mixtral and 128 GB for LLaMA-2 (Lieber et al., arXiv:2403.19887). The ratios are in the same range as the arithmetic above, which is what you would expect from a 1:7 attention-to-Mamba ratio.
What the constant state buys in serving
On an 80 GB GPU holding 14 GB of weights, about 60 GB remains for caches. At 16 GiB per sequence the GQA Transformer fits three concurrent 128k requests. The hybrid fits roughly 25 by the cache arithmetic alone. The pure SSM is limited by activations, not state.
These are upper bounds that ignore activation memory, fragmentation, and framework overhead. They also assume you can actually keep all those requests compute-bound; decode batch size gains are real but saturate. Still, the order-of-magnitude difference explains why inference providers care.
Where the compute goes during prefill
Linear-time does not mean free. During prefill, attention cost per layer is quadratic in length, but with FlashAttention-class kernels the constant is small until contexts reach tens of thousands of tokens. SSM prefill is linear but each token does more per-layer work than a single attention head. At short contexts, below a few thousand tokens, a well-tuned Transformer is often as fast or faster.
The crossover depends on hardware, kernel quality and model width. Any SSM speedup figure quoted without context length and batch size should be read with suspicion, including those in this article that originate from authors’ abstracts.
The recall weakness
An SSM compresses history into a fixed number of values. If the task needs an exact string from 80k tokens earlier, a Transformer can look it up directly. An SSM must have decided to store that string, in a form that survives every subsequent state update.
NVIDIA’s “An Empirical Study of Mamba-based Language Models” tested this under controlled conditions: 8B-parameter Mamba, Mamba-2, Transformer and hybrid models trained on the same data, up to 3.5T tokens. The pure SSMs matched or exceeded the Transformer on many tasks but lagged on tasks requiring strong copying or in-context learning, with 5-shot MMLU and a Phonebook lookup as named examples, and on long-context reasoning (Waleffe et al., arXiv:2406.07887).
This is the most important finding for practitioners. The gap is not general intelligence; it is specifically lookup and copying. That is also the capability that retrieval-augmented generation, tool use, and document question answering lean on hardest.
Why attention fixes it cheaply
Attention is an exact, content-addressable memory. A handful of attention layers is enough to supply that capability to the whole network, because later SSM layers can consume what the attention layers retrieved. That observation is what makes hybrids work: you do not need attention everywhere, only enough to provide lookup.
The same NVIDIA study evaluated Mamba-2-Hybrid, with 43 percent Mamba-2 layers, 7 percent attention layers and 50 percent MLP layers. Per the abstract, it beat the 8B Transformer on all 12 standard tasks evaluated, by 2.65 points on average, and is predicted to generate tokens up to 8 times faster at inference. On 23 long-context tasks with variants extended to 16K, 32K and 128K, it matched or exceeded the Transformer on average. The checkpoints and code are released in Megatron-LM.
These are one lab’s results at one scale with one data mix. They are strong evidence that a small attention fraction suffices, and weaker evidence about any particular frontier-scale recipe.
Hybrid LLM Designs: Jamba and Mamba-2-Hybrid
A hybrid LLM interleaves attention layers with SSM layers, often with MoE or MLP blocks as well. The design question is the ratio and placement, and the published evidence converges on “mostly SSM, a little attention”.

Figure 4: A schematic hybrid stack. Runs of Mamba layers are separated by occasional attention layers, with MoE or MLP blocks interleaved. Layer counts are illustrative, not a specific model.
Jamba: what the paper actually specifies
AI21 Labs’ Jamba is a hybrid Transformer-Mamba mixture-of-experts model. From the paper: 52B total parameters with 12B active, a 1:7 ratio of attention to Mamba layers, 8 layers per block with 4 blocks stacked, and MoE replacing the MLP in every other layer with 16 experts and the top 2 used per token. The authors report the configuration fits on a single 80 GB GPU and supports up to 256K tokens of context (arXiv:2403.19887).
The paper does not specify the Mamba state size in the portions I could verify, so I will not quote one. The reader-relevant point is the stack: MoE buys capacity without proportional compute, Mamba layers buy constant-memory sequence mixing, and the sparse attention layers buy recall.
Follow-up work, Jamba 1.5, shipped Large and Mini variants with 94B and 12B active parameters respectively, an effective 256K window, and a quantization technique called ExpertsInt8 that lets the Large model serve 256K contexts on a single machine with eight 80 GB GPUs, per the authors (arXiv:2408.12570). Check the current model card for license terms and size details before building on it, as these change between releases.
Mamba-2-Hybrid: the controlled comparison
Jamba is a product; NVIDIA’s study is an experiment. Because Mamba, Mamba-2, Transformer and hybrid were trained on identical data at 8B parameters, it isolates the architecture. The headline is that 7 percent attention was enough to erase the recall gap on the evaluated tasks.
That ratio is lower than Jamba’s 1 in 8 (12.5 percent of layers), which is a reminder that the optimum is not universal. It depends on data mix, scale, how the MLP layers are distributed, and which tasks you weigh. Anyone designing a hybrid should treat the attention fraction as a hyperparameter to sweep, not a constant to copy.
Placement and design choices that matter
Where attention sits in the stack is a real design variable. Putting a single attention layer too early gives it only shallow features to retrieve over. Putting all attention at the end limits how much later computation can use the retrieved content. Spreading attention evenly, as Jamba’s block structure does, is a defensible default.
Attention heads in hybrids can often be fewer or use more aggressive KV sharing, since they are not the only mixing mechanism. Positional encoding is another subtlety. Because SSM layers carry order information implicitly through the recurrence, some hybrids drop explicit positional embeddings in their attention layers, which can ease length extrapolation. Verify this per model rather than assuming it.
Linear attention versus SSMs
Linear attention deserves a note because it shares the fixed-state property. It replaces softmax(QK^T)V with a kernelized form phi(Q)(phi(K)^T V), where the running sum phi(K)^T V is the state. Mathematically it is a recurrence with a matrix-valued state and a data-independent decay of one.
Mamba differs by adding selective, input-dependent decay and by a different parameterization, and the SSD result says that restricted forms of both land in the same family. In practice the line is blurry: modern linear-attention variants add gating and decay, and modern SSMs look more like attention. The useful distinction for a buyer is not the name but the state size, the gating mechanism, and the quality of the training kernel.
Reference Implementation: A Selective Scan You Can Run
The following is a minimal, sequential reference of a selective SSM in PyTorch. It is for understanding, not performance: it loops over time in Python, which is exactly what the fused CUDA kernel avoids. It also shows why decoding needs only a fixed-size state.
import torch
import torch.nn as nn
import torch.nn.functional as F
class TinySelectiveSSM(nn.Module):
"""Sequential reference of a Mamba-style selective SSM (single block core)."""
def __init__(self, d_inner=64, d_state=16, dt_rank=8):
super().__init__()
self.d_inner, self.d_state = d_inner, d_state
# A is diagonal and negative real; stored as log for stability
A = torch.arange(1, d_state + 1, dtype=torch.float32).repeat(d_inner, 1)
self.A_log = nn.Parameter(torch.log(A)) # (D, N)
self.D = nn.Parameter(torch.ones(d_inner)) # skip connection
# input-dependent parameters: delta, B, C come from x
self.x_proj = nn.Linear(d_inner, dt_rank + 2 * d_state, bias=False)
self.dt_proj = nn.Linear(dt_rank, d_inner, bias=True)
self.dt_rank = dt_rank
def forward(self, x, state=None):
# x: (batch, length, d_inner)
b, L, D = x.shape
A = -torch.exp(self.A_log) # (D, N)
dt, Bm, Cm = torch.split(
self.x_proj(x), [self.dt_rank, self.d_state, self.d_state], dim=-1)
dt = F.softplus(self.dt_proj(dt)) # (b, L, D)
h = torch.zeros(b, D, self.d_state, device=x.device) if state is None else state
ys = []
for t in range(L):
dA = torch.exp(dt[:, t, :, None] * A) # (b, D, N) discretized A
dB = dt[:, t, :, None] * Bm[:, t, None, :] # (b, D, N) simplified B
h = dA * h + dB * x[:, t, :, None] # state update
y = (h * Cm[:, t, None, :]).sum(-1) + self.D * x[:, t]
ys.append(y)
return torch.stack(ys, dim=1), h # h is the whole memory
if __name__ == "__main__":
torch.manual_seed(0)
m = TinySelectiveSSM()
x = torch.randn(2, 1024, 64)
y_full, h_full = m(x)
# streaming: process in two halves, carrying only h
y1, h1 = m(x[:, :512])
y2, h2 = m(x[:, 512:], state=h1)
print("max diff:", (torch.cat([y1, y2], 1) - y_full).abs().max().item())
print("state bytes per sequence:", h_full[0].numel() * 4) # constant in L
Running the script prints a max difference near zero, confirming that chunked streaming with a carried state reproduces the full-sequence result. The state size printed (64 x 16 x 4 = 4,096 bytes) does not depend on sequence length, which is the property that makes decoding cheap.
The discretization of B here uses the simplified Euler form dt x B, which the Mamba authors also use in their reference. The production implementation fuses all of this into one kernel and handles the causal convolution, gating and projections.
Using the official library
For real use, the official mamba-ssm package from the authors provides fused kernels and requires a CUDA GPU. A minimal block looks like this:
import torch
from mamba_ssm import Mamba
block = Mamba(
d_model=1024, # model width
d_state=16, # SSM state size N
d_conv=4, # local conv width
expand=2, # inner width multiplier
).to("cuda")
x = torch.randn(2, 4096, 1024, device="cuda")
y = block(x) # same shape as x
Check the repository README for the current Mamba-2 class and version requirements, since the API and CUDA compatibility change between releases. Hybrids such as Jamba are also supported in the Hugging Face Transformers library, which is usually the easier on-ramp for inference experiments.
Inference-time behavior to test yourself
Do not trust any single throughput claim. Benchmark your own workload with three variables swept: context length (1k, 8k, 32k, 128k), batch size (1, 8, 32), and phase (prefill versus decode). Record tokens per second, peak memory, and time to first token.
Expect SSM and hybrid advantages to appear mainly at long contexts and larger batches, where cache pressure dominates. At short context and batch size 1, a highly tuned Transformer stack with speculative decoding or quantization may still win. The only honest answer comes from your own sweep.
Trade-offs, Gotchas, and What Goes Wrong
Recall and copying. This is the main failure mode and it is structural. Needle-in-a-haystack tests pass or fail depending on whether the model chose to store the needle, and the failure is silent: the model confidently generates something plausible. If your product depends on verbatim retrieval from context, a pure SSM is the wrong tool, and even a hybrid needs a recall-specific evaluation.
Fixed state means a capacity ceiling. State size is a design knob. Mamba-2 allows larger states than Mamba because of its matmul-friendly form, but a larger state raises per-token compute and memory. There is no setting where a fixed state matches unbounded lookup.
Tooling maturity. The Transformer ecosystem has paged KV caches, continuous batching, speculative decoding, prefix caching, and years of kernel tuning. SSM support in serving engines is improving but lags. Prefix caching is a good example: Transformers can reuse a stored cache for a shared system prompt, while SSMs must store and restore a state snapshot at the right boundary, and the engine has to support that.
Quantization sensitivity. Recurrent state accumulates error over time, which is why implementations often keep it in higher precision. Aggressive low-bit quantization of SSM parameters can behave differently than for attention weights. Validate on long sequences, not only short benchmarks.
Training and fine-tuning quirks. Tensor-parallel and sequence-parallel strategies for SSMs are less standardized. Parameter-efficient fine-tuning recipes designed for attention projections may not map directly to SSM blocks. Expect more experimentation, and check that your framework version supports the model you chose.
Benchmarks overstate the win. Many headline numbers are throughput at long context with large batches, which is the most flattering regime. Reported quality parity often comes from models of a few billion parameters. Treat frontier-scale claims as unproven until you see controlled comparisons with matched data.
Anti-pattern: pure SSM for RAG-heavy products. Retrieval-augmented systems stuff passages into context and ask the model to quote them. That is precisely the lookup workload where SSMs are weakest. If you want SSM economics here, use a hybrid and evaluate on your own documents.
Anti-pattern: copying the 1:7 ratio. The Jamba and NVIDIA results use different attention fractions. The right number depends on your task mix, so sweep it if you are training, and test it if you are buying.
Practical Recommendations
Start from the workload, not the architecture. The question is whether memory per request or recall fidelity is your binding constraint.
If you serve long contexts at high concurrency and cache memory limits your batch size, evaluate a hybrid. The 4 GB versus 32 GB to 128 GB cache comparison in the Jamba paper shows why. If your requests are short and latency-sensitive, stay with an optimized Transformer; you are not paying the cache tax and the tooling is better.
If your product depends on exact retrieval from context, require evidence on recall-specific benchmarks and on your own data. If you are doing streaming, signal processing, genomics, or other very long non-text sequences, pure SSMs remain a strong fit, as the S4 and Mamba papers’ domains suggest.
Decision matrix, based on the evidence discussed above:
| Use case | Transformer (GQA) | Pure SSM | Hybrid |
|---|---|---|---|
| Short chat, low concurrency | Best tooling, strong | Fine, little gain | Little gain |
| 128k+ document QA with quoting | Strong, memory heavy | Risky on recall | Strong candidate |
| High-concurrency long-context serving | Cache limits batch | Best memory, recall caveat | Best balance |
| Streaming or very long signals | Costly | Strong fit | Optional |
| Mature fine-tuning and serving stack | Best | Limited | Improving |
Checklist before adopting:
- Measure your median and 95th percentile context length, and your concurrency target.
- Compute KV cache per request with your exact head and layer counts.
- Run a recall test on your own documents, with the needle at several depths.
- Sweep context, batch size and phase on your target GPU.
- Confirm your serving engine supports the architecture, including prefix caching.
- Confirm the license and the current model card for any open weights.
If you also care about physical-world sensor streams, the same fixed-state logic applies to long telemetry, the kind of problem seen in conjunction assessment architectures for tracking space debris, where streaming state estimators have always preferred recurrences over re-reading history.
Frequently Asked Questions
What are Mamba state space models?
Mamba state space models are sequence models built on a selective state space layer. Each token updates a fixed-size hidden state through a linear recurrence whose parameters depend on the current input, letting the model choose what to remember or forget. Introduced by Albert Gu and Tri Dao in December 2023, Mamba computes this with a hardware-aware parallel scan, giving compute and memory that scale linearly with sequence length and constant memory per generated token.
Is Mamba better than a Transformer?
Not universally. The authors report Mamba-3B beating same-size Transformers and matching ones twice its size in their evaluations. But NVIDIA’s controlled 8B study found pure Mamba and Mamba-2 lag Transformers on tasks needing strong copying or in-context learning. Mamba wins on memory and long-sequence throughput; Transformers win on exact recall and tooling maturity. Hybrids aim to capture both.
What is selective scan in Mamba?
Selective scan is the algorithm that computes Mamba’s input-dependent recurrence efficiently. Because the parameters change per token, the model cannot use a fixed convolution, so it uses a parallel prefix scan fused into one GPU kernel. The expanded state stays in fast on-chip SRAM instead of high-bandwidth memory, and the backward pass recomputes states rather than storing them, following the same logic as FlashAttention.
What is the difference between Mamba and Mamba-2?
Mamba-2 reformulates the selective SSM using state space duality, which shows SSMs and attention variants are related through structured semiseparable matrices. This lets the core layer use matrix multiplications that map onto tensor cores. The authors report the core layer is 2 to 8 times faster than Mamba’s while staying competitive on language modeling, partly by restricting A to a scalar times identity per head.
What is a hybrid LLM and why use one?
A hybrid LLM mixes attention layers with SSM layers, sometimes plus mixture-of-experts. Attention supplies exact recall; SSM layers supply cheap, constant-memory sequence mixing. Jamba uses one attention layer per seven Mamba layers and reports a 4 GB KV cache at 256K tokens, versus 32 GB for Mixtral. NVIDIA’s Mamba-2-Hybrid used about 7 percent attention layers and matched or beat an 8B Transformer on evaluated tasks.
Does constant state mean unlimited context?
No. Constant memory means the cost does not grow, not that quality holds. The state is a lossy summary, so details can be overwritten. Jamba reports strong results up to 256K tokens and Mamba’s authors show improvement up to million-token sequences on selected data, but real long-context reliability must be tested on your task, especially for retrieval of exact facts.
Further Reading
- GPU memory physics: DRAM vs SRAM and FlashAttention kernel fusion for the memory-hierarchy ideas Mamba’s scan reuses.
- Space debris tracking and conjunction assessment architecture for streaming state estimation in a physical system.
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces, Gu and Dao.
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, Dao and Gu.
- An Empirical Study of Mamba-based Language Models, NVIDIA and collaborators.
- Jamba: A Hybrid Transformer-Mamba Language Model, AI21 Labs.
By Riju — about
