ML-KEM and ML-DSA on Constrained IoT Microcontrollers: Memory, Speed and Migration

ML-KEM and ML-DSA on Constrained IoT Microcontrollers: Memory, Speed and Migration

ML-KEM ML-DSA IoT: Memory, Speed and Migration on Constrained Microcontrollers

A Cortex-M4 sensor node can run ML-KEM and ML-DSA on IoT hardware in well under three kilobytes of stack, yet the same node may still fail its post-quantum migration. The reason is rarely the arithmetic. It is the bytes: a 1,184-byte ML-KEM-768 public key and a 2,420-byte ML-DSA-44 signature are 18 to 38 times larger than their elliptic-curve predecessors, and they land on radios whose frames hold 127 bytes or fewer.

This matters now because NIST finalized FIPS 203 (ML-KEM) and FIPS 204 (ML-DSA) in August 2024, procurement rules for national-security systems already reference them, and industrial devices fielded today will still be running in 2040. Harvest-now-decrypt-later attacks mean confidentiality keys need protection first, while firmware-signing roots of trust need a plan before the factory line is frozen.

You will leave with exact size tables, measured cycle and memory figures from the open pqm4 benchmark project, a handshake-cost model for constrained links, and a migration order that protects what is hard to patch.

What this covers: the standards and parameter sizes, the reference architecture for a PQC-capable device, measured Cortex-M4 costs, handshake and fragmentation arithmetic, the stateful hash-based option for firmware, failure modes, and a practical rollout checklist.

Context and Background

Public-key cryptography on embedded devices has settled on a small toolkit for a decade: ECDH over Curve25519 or P-256 for key agreement, and ECDSA or Ed25519 for signatures, sometimes RSA where legacy PKI demands it. These schemes are compact. An X25519 public key is 32 bytes, a P-256 ECDSA signature is 64 bytes, and a full TLS 1.3 handshake with an elliptic-curve certificate chain fits in a few kilobytes. A Shor-capable quantum computer would break all of them, which is why the industry is moving to lattice-based replacements.

NIST published three post-quantum standards on 13 August 2024: FIPS 203 (ML-KEM, derived from CRYSTALS-Kyber), FIPS 204 (ML-DSA, derived from CRYSTALS-Dilithium) and FIPS 205 (SLH-DSA, derived from SPHINCS+). NIST’s FIPS 204 publication page lists the 13 August 2024 date. A separate stateful hash-based track, NIST SP 800-208, standardizes LMS/HSS and XMSS/XMSS-MT for contexts such as firmware signing. For the broader strategy question of inventory, crypto-agility and sequencing across an enterprise, see our companion piece on post-quantum cryptography migration and crypto-agility. This article deliberately narrows the lens to the engineering numbers that decide whether a constrained device can carry the new algorithms.

Two threat clocks drive device design. Confidentiality has a harvest-now-decrypt-later clock: traffic recorded today could be decrypted later, so key establishment for long-lived secrets (utility metering, medical telemetry, industrial process data) should move first. Authenticity has a forgery clock that only matters once a cryptographically relevant quantum computer exists, but a device that cannot be updated cannot change its root of trust later. That asymmetry is the first original observation of this article: key exchange is cheap to migrate and urgent, while signature roots of trust are expensive to migrate and urgent for a different reason, namely that they are frozen at manufacture.

NSA’s CNSA 2.0 guidance sets the tone for regulated sectors. As summarized by postquantum.com’s CNSA 2.0 FAQ, software and firmware signing is on the earliest schedule: support and prefer CNSA 2.0 algorithms by 2025 and use them exclusively by 2030, with LMS and XMSS as the preferred near-term firmware choices and ML-DSA-87 approved once validated implementations are available. Treat that secondary summary as a pointer and read the NSA text itself before committing a product roadmap.

What “constrained” means in numbers

RFC 7228 defines classes of constrained devices, but practical engineering uses a simpler test: how much RAM remains after the application, the network stack and the TLS or DTLS state machine have taken their share. A Class 1 device might have around 10 KB of RAM and 100 KB of flash; a mainstream Cortex-M4 industrial gateway sensor often has 128 to 640 KB of RAM and 512 KB to 2 MB of flash. The pqm4 test board, an STM32L4R5ZI, has 640 KB of SRAM and 2 MB of flash, which makes it a generous target. Results on it are an upper bound on what a Cortex-M0+ part will manage and a lower bound on stack needs only when the same code paths are used.

Reference Architecture: A PQC-Capable Constrained Device

The reference architecture answers one question: which post-quantum primitive goes where on a device, and what does each place cost? The device needs three cryptographic roles. A key-establishment role (session security with a gateway or cloud), an authentication role (proving device identity during the handshake), and a firmware-verification role (checking signed images at boot and during over-the-air updates).

ML-KEM ML-DSA IoT device cryptographic roles and the primitives that fill them

Figure 1: Three cryptographic roles on a constrained device and the post-quantum primitive typically assigned to each.

Figure 1 shows the split. Session key establishment uses ML-KEM-768, usually inside a hybrid with X25519. Handshake authentication uses ML-DSA-44 or ML-DSA-65 certificates, or a pre-shared or raw-public-key shortcut where the deployment allows it. Boot-time firmware verification uses LMS or XMSS, or ML-DSA, with the verification key anchored in immutable storage. The three roles have different latency tolerances, different memory budgets and different upgrade stories, so treating “PQC” as a single switch is the most common planning error.

Direct answer: what do ML-KEM and ML-DSA cost on a Cortex-M4?

On the pqm4 benchmark for a Cortex-M4 at 24 MHz measurement conditions, ML-KEM-768 with the stack-optimized implementation needs about 2.8 KB of stack and roughly 13 KB of flash, with encapsulation near 665,000 cycles. ML-DSA-44 verification needs about 2.7 KB of stack in the stack-optimized build and about 1.4 million cycles in the speed-optimized build. Signing is variable and several times more expensive.

Standards-defined sizes

The sizes below are fixed by the standards, so they are the one part of this article that does not depend on any implementation. ML-KEM-768 sizes (public key 1,184 B, secret key 2,400 B, ciphertext 1,088 B) are corroborated by the Kyber overview on Wikipedia and the hybrid TLS draft discussed later. ML-DSA public key and signature sizes are corroborated by Go’s standard-library crypto/mldsa constants. The remaining ML-KEM-512/1024 and ML-DSA secret-key sizes are standard FIPS 203 and 204 values that I did not re-fetch from the PDFs in this run, so verify them against the standards before quoting them in a datasheet.

Scheme Security category Public key (B) Secret key (B) Ciphertext or signature (B)
ML-KEM-512 1 800 1,632 768
ML-KEM-768 3 1,184 2,400 1,088
ML-KEM-1024 5 1,568 3,168 1,568
ML-DSA-44 2 1,312 2,560 2,420
ML-DSA-65 3 1,952 4,032 3,309
ML-DSA-87 5 2,592 4,896 4,627
SLH-DSA-SHA2-128s 1 32 64 7,856
SLH-DSA-SHA2-128f 1 32 64 17,088
X25519 (for contrast) classical 32 32 32 shared
ECDSA P-256 (for contrast) classical 64 32 64

The SLH-DSA signature sizes come from the PQ Crypto Registry summary of FIPS 205. Two points follow from the table. First, ML-KEM-768 replaces a 32-byte X25519 share with 1,184 bytes in one direction and 1,088 in the other, a factor of roughly 34 to 37. Second, the secret keys are bigger than they look, but both standards allow storing a compact seed and regenerating the expanded key on demand. FIPS 204 permits a 32-byte seed for ML-DSA (confirmed by the Go library’s seed-only private key), and FIPS 203 keys can similarly be reconstructed from a 64-byte seed. For flash-constrained or secure-element-constrained parts, seed storage turns a 2.5 KB private key into 32 bytes at the cost of a key-expansion step.

Why the sizes, not the cycles, set the design

An ML-KEM-768 encapsulation on a Cortex-M4 costs a few hundred thousand cycles, which is about the same order as the elliptic-curve scalar multiplications that already run on such devices. A modest 64 MHz core completes it in roughly ten milliseconds. The arithmetic is not the obstacle. The obstacle is that a 1,184-byte key share does not fit in one IEEE 802.15.4 frame (127-byte maximum PHY payload), does not fit in one LoRaWAN uplink at low data rates, and pushes a TLS ClientHello past the single-packet size that middleboxes and small-MTU radios handle well.

The hybrid pattern in TLS 1.3

The current IETF approach pairs a classical and a post-quantum exchange so that the session stays safe if either one falls. The IETF draft draft-ietf-tls-ecdhe-mlkem-05 is a Standards Track Internet-Draft that defines X25519MLKEM768 (codepoint 4588, 0x11EC), SecP256r1MLKEM768 (4587) and SecP384r1MLKEM1024 (4589). For X25519MLKEM768 the client key share is 1,216 bytes (1,184 for ML-KEM-768 plus 32 for X25519) and the server share is 1,120 bytes (1,088 plus 32); the combined shared secret is 64 bytes. These figures are as read from the draft in this run. Because it is still a draft with an expiry date, codepoints and details should be rechecked before you ship, even though the codepoints have been widely deployed in browsers and servers.

For constrained devices, hybridization has a cost that is easy to overlook: it keeps the X25519 code and adds ML-KEM, so flash is the sum of both. If a device supports only ML-KEM with no classical component, it loses the belt-and-braces property that regulators and cautious architects currently want.

Measured Cost on a Cortex-M4: What pqm4 Actually Reports

The pqm4 project is the community benchmarking and test framework for post-quantum schemes on the ARM Cortex-M4, run on an STM32L4R5ZI. Its README states that cycle counts were obtained at 24 MHz to avoid flash wait states, and that the compiler was the Arm GNU Toolchain 11.3.Rel1 (arm-none-eabi-gcc 11.3.1). The numbers below are read from the project’s benchmarks.md and benchmarks.csv as retrieved on the day of writing; the project updates them as implementations improve, so treat them as a snapshot and re-pull the file when you size a product. They are not my own measurements.

pqm4 reports several implementation flavors. For ML-KEM, “clean” is the portable reference C, “m4fspeed” is assembly-optimized for speed, and “m4fstack” minimizes stack at a small speed penalty. For ML-DSA, “clean”, “m4f” (speed) and “m4fstack” play the same roles. The choice among them is the single biggest lever you control.

ML-KEM: cycles, stack and flash

ML-KEM ML-DSA IoT Cortex-M4 cost map from pqm4 cycles, stack and flash

Figure 2: How the pqm4 implementation flavors trade cycles against stack and flash for ML-KEM and ML-DSA on a Cortex-M4.

Figure 2 summarizes the trade-off structurally; the exact figures are in the tables. For ML-KEM-768 the pqm4 mean cycle counts are:

ML-KEM-768 flavor KeyGen (cycles) Encaps (cycles) Decaps (cycles) Stack, worst op (B) Flash (B)
clean 988,722 1,138,225 1,387,984 14,480 5,120
m4fspeed 642,096 658,754 707,827 6,468 16,016
m4fstack 644,195 664,654 714,194 2,860 13,320

Two things stand out. First, the stack-optimized build costs under one percent more cycles than the speed build (644,195 versus 642,096 for key generation) while cutting peak stack from about 6.5 KB to under 2.9 KB. For a constrained device that is nearly a free lunch, and it is why m4fstack is the default recommendation for memory-poor parts. Second, the “clean” reference code is smaller in flash (about 5 KB) but uses roughly five times the stack of m4fstack and is about 1.5 to 2 times slower, so on a part with little RAM and plenty of flash, hand-optimized assembly wins on every axis but flash.

Scaling across parameter sets is gentle. In the m4fstack flavor, ML-KEM-512 encapsulation is 392,864 cycles, ML-KEM-768 is 664,654 and ML-KEM-1024 is 1,037,953, and stack stays between about 2.3 KB and 3.4 KB. Choosing ML-KEM-1024 instead of 768 for a CNSA 2.0 requirement costs roughly 1.56 times the encapsulation cycles and 512 extra bytes of stack, which is affordable. What it does not spare you is the wire size: a 1,568-byte public key and ciphertext.

Converting cycles to time

pqm4 reports cycles, not milliseconds, because wall time depends on the core clock and flash wait states. The conversion is plain division, and it is worth doing for your own part. Using the pqm4 means:

Operation (flavor) Cycles At 24 MHz At 64 MHz At 120 MHz
ML-KEM-768 encaps (m4fspeed) 658,754 27.4 ms 10.3 ms 5.5 ms
ML-KEM-768 decaps (m4fspeed) 707,827 29.5 ms 11.1 ms 5.9 ms
ML-DSA-44 verify (m4f) 1,421,623 59.2 ms 22.2 ms 11.8 ms
ML-DSA-44 sign mean (m4f) 3,943,121 164 ms 61.6 ms 32.9 ms
ML-DSA-44 sign max seen (m4f) 17,009,165 709 ms 266 ms 142 ms

The 64 and 120 MHz columns are my arithmetic and assume cycle counts scale with clock, which breaks down on parts where flash wait states increase at higher frequencies. They are illustrative, not measured. At 24 MHz the table is the closest to what pqm4 measured.

ML-DSA: the signing distribution is the story

ML-DSA signing rejection loop and its effect on worst-case latency on a microcontroller

Figure 3: ML-DSA signing uses rejection sampling, so latency is a distribution with a long tail rather than a constant.

ML-DSA verification is a fixed amount of work, but signing is not. The algorithm samples a candidate, tests whether the result leaks information about the secret key, and retries if it does. Each attempt costs a similar number of cycles, and the number of attempts is random. pqm4 therefore reports minimum, average and maximum over 1,000 executions. For ML-DSA-44 in the m4f flavor the figures are:

ML-DSA-44 m4f Min (cycles) Mean (cycles) Max (cycles)
KeyGen 1,379,650 1,426,025 1,466,529
Sign 1,812,557 3,943,121 17,009,165
Verify 1,420,738 1,421,623 1,422,362

The maximum signing time over 1,000 runs is about 4.3 times the mean and more than nine times the minimum. For ML-DSA-65 m4f the mean is 6,193,171 cycles and the observed maximum 26,008,621; for ML-DSA-87 m4f the mean is 7,947,380 and the maximum 29,357,607. If your protocol has a timeout, size it for the tail, not the mean. This is the second original observation here: on constrained devices, plan ML-DSA signing as a bounded-retry latency distribution, and put signing on the gateway or backend where possible, because a verify-heavy workload is far friendlier to MCUs than a sign-heavy one. Most IoT designs happen to be verify-heavy on the device (check firmware, check the server certificate) and sign-heavy only for device authentication, which is one signature per session.

The memory flavors differ even more for ML-DSA. For ML-DSA-44, verification needs 36,308 bytes of stack in clean, 8,912 in m4f and 2,712 in m4fstack. Signing needs 51,976, 44,816 and 5,080 bytes respectively. That is a striking spread: the speed build of signing wants almost 45 KB of stack, which is out of reach for many MCUs, while the stack build needs about 5 KB. But the stack build pays heavily in time: its mean signing cost for ML-DSA-44 is 12,134,284 cycles, roughly three times the m4f mean, and its verification is 3,242,333 cycles versus 1,421,623, about 2.3 times. The max signing observation for m4fstack ML-DSA-44 is 65,357,843 cycles, which is about 2.7 seconds at 24 MHz.

ML-DSA-44 flavor Verify cycles Verify stack (B) Sign mean (cycles) Sign stack (B) Flash (B)
clean 2,063,096 36,308 7,925,955 51,976 8,212
m4f 1,421,623 8,912 3,943,121 44,816 19,592
m4fstack 3,242,333 2,712 12,134,284 5,080 24,844

Key generation stack is also a surprise: 38,296 bytes for m4f (and 4,408 for m4fstack). Key generation is normally run once, at provisioning, so a provisioning-time budget (or a factory station that does it for the device) removes the pressure entirely.

What the flash numbers mean

ML-KEM-768 code in m4fspeed is about 16 KB and m4fstack about 13 KB (pqm4’s total column); ML-DSA-44 is roughly 19.6 KB in m4f and 24.8 KB in m4fstack. Add the Keccak permutation, which ML-KEM and ML-DSA both depend heavily on (SHAKE and SHA-3 are a large share of the runtime), and you should budget around 40 to 45 KB of flash for a device that does both, plus a hybrid X25519 implementation if you keep one. These flash totals are my sum of pqm4’s figures and an allowance; pqm4 reports per-scheme sizes, so check whether shared symbols such as the hash code are counted twice before relying on the sum.

Caveats on the benchmark

Three limits apply. pqm4 targets Cortex-M4 with its DSP-style instructions; a Cortex-M0+ (ARMv6-M) lacks the DSP-style packed multiply-accumulate instructions the optimized assembly exploits, so cycles will be noticeably worse there and I do not have verified pqm4-grade M0+ numbers to quote. The measurements are at 24 MHz without peripherals active, so interrupt jitter and DMA contention on a real device add latency. And side-channel resistance is a separate axis: constant-time behavior of the reference and optimized code is a design goal, but power and fault attacks on a physically accessible device need countermeasures that cost more cycles and memory than any figure above.

Handshake Impact: Bytes on the Wire, Fragments and Airtime

A handshake transfers key shares, certificates and signatures. Post-quantum algorithms inflate all three, and constrained links punish size in ways that cycle counts never reveal. This section builds the arithmetic you can apply to your own radio.

ML-KEM ML-DSA IoT hybrid handshake message flow with sizes and fragmentation on a constrained link

Figure 4: Hybrid ML-KEM handshake with ML-DSA authentication, showing where each large message forces fragmentation on a constrained link.

The byte budget of a hybrid TLS or DTLS handshake

Start with key exchange. In the X25519MLKEM768 hybrid, the client sends 1,216 bytes of key share and the server returns 1,120 bytes. Classical X25519 needs 32 bytes each way. The key exchange alone therefore adds about 2.2 KB over the classical baseline.

Now authentication. If the server certificate chain uses ML-DSA-44 end to end, each certificate carries a 1,312-byte public key and a 2,420-byte signature from its issuer, so each certificate is at least 3,732 bytes before names, extensions and ASN.1 overhead. A two-certificate chain (leaf plus intermediate) is therefore at least about 7.5 KB, and the CertificateVerify signature adds another 2,420 bytes. The authentication payload from the server lands near 10 KB. The same chain with P-256 ECDSA is roughly 1 KB of certificate content and a 64-byte signature, plus overhead. These totals are my own derivation from the standard sizes; real certificates are larger by their metadata.

Handshake component Classical (X25519, ECDSA P-256) Hybrid KEM, ML-DSA-44 chain
Client key share 32 B 1,216 B
Server key share 32 B 1,120 B
Leaf certificate (key + issuer sig, min) about 128 B plus overhead at least 3,732 B plus overhead
Two-cert chain (min) about 256 B plus overhead at least 7,464 B plus overhead
CertificateVerify signature 64 B 2,420 B
Rough server flight, key share plus chain plus verify under 1 KB typical about 10.5 KB or more

Estimates like these explain why many constrained deployments stop sending full chains at all. Options include raw public keys (RFC 7250), pre-installed trust anchors with a single device-to-server hop, certificate compression, or an application-layer token that proves possession without X.509. Each of those removes bytes the radio would otherwise carry.

IPv6 requires a minimum link MTU of 1,280 bytes, and 6LoWPAN (RFC 4944) fragments packets across IEEE 802.15.4 frames, which carry at most 127 bytes at the PHY. After link-layer, security and compressed IPv6 and UDP headers, the usable payload per frame is on the order of 80 to 100 bytes, depending on addressing and security mode. At that rate a 1,216-byte ClientHello key share needs roughly 12 to 15 fragments, and a 10 KB server flight roughly 100 to 125. Both counts are my estimates under that assumed payload range.

Fragmentation matters because losing one fragment loses the whole datagram. With a per-frame loss probability p and n fragments, the datagram survives with probability (1 – p) raised to the n. At 2 percent frame loss, 14 fragments succeed about 75 percent of the time, while 110 fragments succeed about 11 percent. DTLS retransmission timers then fire and retransmit entire flights, and the cost compounds on duty-cycled radios. DTLS 1.3 (RFC 9147) improves matters by supporting handshake-message fragmentation and retransmission at message granularity, but it cannot shrink the bytes.

LoRaWAN is a harder case. Regional parameters allow application payloads as small as 51 bytes at the lowest data rates in EU868, and a 1 percent duty-cycle limit applies in many sub-bands. A 1,184-byte key would need at least 24 uplinks at 51 bytes, and a regulated duty cycle forces long silences between them. I did not verify the exact regional-parameter numbers in this run, so check the LoRaWAN Regional Parameters document for your region. The practical conclusion stands: do not run an in-band PQC handshake over LoRaWAN uplinks. Use the network’s built-in key hierarchy and put post-quantum protection at the join server, the backhaul and the application server, where bandwidth is plentiful.

NB-IoT and LTE-M are friendlier because they carry IP packets of normal size, but they have their own cost: each extra transmission wakes a modem that dominates the energy budget. A handshake that grows from 2 KB to 12 KB on a battery device is an energy problem long before it is a CPU problem. Quantifying it requires your modem’s per-byte energy, which I do not have a verified figure for, so measure it on your hardware.

Session resumption is the pressure valve

TLS 1.3 session resumption with pre-shared keys (PSK) and DTLS Connection IDs (RFC 9146) let a device pay the post-quantum handshake once and reuse the result. If the device wakes every ten minutes to report, a full hybrid handshake each time is wasteful. A resumed session uses a PSK derived from the earlier exchange, and when it is paired with a PSK-with-ephemeral-key-exchange mode it still gains forward secrecy, at the price of another ML-KEM operation but not another certificate chain. Resumption tickets are symmetric and therefore already quantum-resistant against the attack that breaks key exchange, provided the initial exchange was safe.

Where to terminate

A pragmatic pattern is to terminate the constrained link at a nearby gateway with plenty of resources, do post-quantum TLS from there to the cloud, and use a lighter, quantum-safe symmetric or pre-shared-key scheme on the last hop. This does not remove the need for ML-KEM on the device when the device itself must prove end-to-end security, but it reduces the number of devices that must carry full public-key stacks.

Firmware Signing: Stateful Hash-Based Signatures versus ML-DSA

Firmware verification differs from handshake authentication in four ways. Signatures are checked once per image, not once per session. The verification key is baked into immutable storage. The signing happens in a controlled factory or release environment, not on the device. And the device cannot easily be fixed in the field if the algorithm turns out to be wrong. That mix favors conservative primitives.

LMS and XMSS

LMS/HSS (RFC 8554) and XMSS/XMSS-MT (RFC 8391) are stateful hash-based signature schemes standardized for US government use in NIST SP 800-208. Their security rests only on hash functions, which is the strongest conservative argument available. The catch is state: each one-time key must be used at most once. Reusing a one-time key breaks security, so the signing system must reliably track a counter, ideally inside a hardware security module with tamper-resistant monotonic state.

An LMS public key is compact. Per RFC 8554’s structure it is 56 bytes (a 4-byte LMS type, a 4-byte LM-OTS type, a 16-byte identifier and a 32-byte root), a derivation from the RFC format rather than a number I re-fetched. Signature size depends on the tree height and the Winternitz parameter. An arXiv benchmark study of stateful hash-based signatures, Stateful Hash-Based Signature Benchmark Data for XMSS and LMS, reports LMS signature sizes from 1.27 KB to 9.11 KB, and XMSS from 1.46 KB to 27.04 KB, with smaller LMS signatures from larger Winternitz parameters. That study measured validation on x86 with liboqs and reports no Cortex-M benchmarks, so I cannot quote an embedded verification time for LMS. Verification cost is dominated by hash evaluations and is generally cheap on devices with a SHA-256 accelerator, but you should measure it on yours.

For a firmware update, a signature of 1.3 to 9 KB is negligible against an image of hundreds of kilobytes. This is the key insight: a signature that is painful in a handshake is trivial when attached to a firmware image, so the “big signature” objection against hash-based and lattice signatures largely disappears for the firmware role.

ML-DSA versus LMS for firmware

ML-DSA needs no state management, which removes the largest operational risk of LMS and XMSS. Its verification costs about 1.4 million cycles on a Cortex-M4 in the m4f flavor, or 2,712 bytes of stack in the m4fstack flavor, per pqm4. LMS and XMSS offer smaller public keys and a purely hash-based security argument, but a finite signature budget fixed at key generation, for example a tree of height 20 yielding about a million signatures, and the need for a stateful signing service.

Criterion LMS or XMSS ML-DSA-44 or 65 SLH-DSA
Standard NIST SP 800-208, RFC 8554 and 8391 FIPS 204 FIPS 205
Public key tens of bytes 1,312 to 1,952 B 32 to 64 B
Signature about 1.3 to 9 KB 2,420 to 3,309 B 7,856 B and up
State management required, high risk none none
Verify on MCU hash-bound, measure about 1.4 to 2.4 million cycles on M4 hash-bound, slower to sign
Best fit root of trust, offline signing general firmware and handshake conservative stateless root

For a device with a ten- to twenty-year life, a reasonable approach is dual: ship an LMS or XMSS root key in immutable memory for the first-stage boot loader, and use ML-DSA for later stages and for update packages, with a documented path to rotate the second-stage trust anchor. The CNSA 2.0 summary above reflects this thinking, with LMS and XMSS preferred for firmware roots of trust now and ML-DSA acceptable once validated implementations exist.

Trade-offs, Gotchas, and What Goes Wrong

Stack overruns from picking the wrong flavor. The ML-DSA-44 signing routine in the speed-optimized build wants about 44.8 KB of stack according to pqm4. Dropped onto an RTOS task with a 4 KB stack, it corrupts memory silently. Always measure stack high-water marks on the real target and use the stack-optimized build where RAM is scarce.

Timeouts sized for the mean. Because ML-DSA signing is a retry loop, a watchdog or protocol timer tuned to the average will fire on the long tail. pqm4’s observed maximum of 17 million cycles against a 3.9 million mean for ML-DSA-44 is a normal outcome, not an anomaly. Budget for at least the observed maximum with headroom, or move signing off the device.

Fragmentation amplification. One lost frame discards a whole fragmented datagram. A migration that looks fine in a lab with a wired link degrades in the field on a lossy mesh, and retransmitted flights multiply airtime and battery drain. Test with realistic loss rates, not zero.

Stateful signature misuse. Reusing an LMS or XMSS one-time key is a catastrophic, silent failure. Backups of a signing server, virtual machine snapshots and cloned HSM partitions can all rewind state. Operational controls, such as keeping the key inside a hardware module that guarantees monotonic counters, matter more than the algorithm.

Side channels and fault injection. Lattice schemes involve secret-dependent operations over small coefficient ranges and large amounts of hashing. Constant-time reference code protects against timing leaks but not against power analysis or fault attacks on a device an adversary can hold. Masked implementations exist but cost cycles and RAM; no pqm4 figure above includes them, so a product with physical-attack requirements needs a separate budget.

Hybrid is not free. Keeping X25519 beside ML-KEM adds flash and a second code path to test, and the hybrid combiner in the specification must be implemented exactly. Do not invent your own combiner by concatenating secrets without the standard’s key-derivation structure.

Unvalidated implementations. A fast assembly implementation from a research repository is not a FIPS 140-3 validated module. If your sector requires validated cryptography, plan for a validated library or a secure element, which may have different performance and memory characteristics than the pqm4 results. I did not verify the current validation status of any specific embedded library for this article.

Standard drift. The hybrid TLS specification is an Internet-Draft. Codepoints and structures can change, and FIPS errata exist; NIST’s FIPS 204 page notes an errata spreadsheet. Pin versions, and keep the cryptographic suite behind an abstraction so it can be replaced.

Certificates and revocation. Larger certificates inflate OCSP responses, CRLs and certificate transparency proofs too. A device-side revocation strategy built on 64-byte signatures does not scale unchanged to 2.4 KB signatures; shorten chains, use short-lived credentials and avoid revocation lists that the device must download.

Practical Recommendations

Start from the roles in Figure 1 and move them in order of urgency and difficulty. Key establishment moves first, because it protects confidentiality against recorded traffic and can be updated by firmware. Firmware-verification roots move early at the design stage, because they freeze at manufacture. Handshake authentication moves last, because certificate-chain infrastructure has to change on the server side too, and because many constrained devices can avoid heavy X.509 entirely.

For a Cortex-M4 or better with at least 32 KB of RAM free, deploy hybrid X25519MLKEM768 for session key establishment using the stack-optimized ML-KEM implementation, and reserve roughly 3 KB of stack and 13 to 16 KB of flash for it. Terminate constrained links at a gateway where radio bandwidth is the limit. For firmware, put an LMS or XMSS public key in immutable memory with an offline, HSM-held signing key and a documented state-management procedure, and add an ML-DSA path for later boot stages and update packages. For device authentication, prefer raw public keys or pre-provisioned trust anchors to long X.509 chains, and sign on the device only once per session.

A short checklist for the next hardware spec review:

  1. Inventory every place the device uses RSA, ECDH or ECDSA, and classify it as key exchange, session authentication or firmware verification.
  2. Measure real RAM, stack high-water mark and flash on the target for ML-KEM-768 and your chosen signature scheme, not on a reference board.
  3. Compute handshake bytes and fragment counts for your radio and your measured loss rate.
  4. Set protocol timeouts from ML-DSA’s tail latency, not its average.
  5. Store keys as seeds where the standards allow it, and protect them with a secure element if one exists.
  6. Reserve flash for a second signature scheme and for algorithm agility; make the cipher suite a field in the update manifest, not a constant.
  7. Decide where stateful signing state lives and who audits it.
  8. Pin the draft and library versions, and schedule a re-check against the standards before release.

Finally, align with the enterprise-level plan: the sequencing and inventory practices in our post-quantum migration and crypto-agility guide apply to fleets of devices as much as to servers, and the device-level numbers in this article feed straight into it.

Frequently Asked Questions

Can ML-KEM and ML-DSA run on a Cortex-M4 microcontroller?

Yes. The pqm4 benchmark shows ML-KEM-768 in about 2.8 KB of stack and 13 KB of flash with the stack-optimized build, with encapsulation near 665,000 cycles. ML-DSA-44 verification needs about 2.7 KB of stack in its stack-optimized build and about 1.4 million cycles in its speed build. Signing is slower and variable. The constraint is usually radio bandwidth and RAM left over for the protocol stack, not raw compute.

How big are ML-KEM-768 and ML-DSA-44 keys and signatures?

ML-KEM-768 has a 1,184-byte public key, a 2,400-byte secret key and a 1,088-byte ciphertext. ML-DSA-44 has a 1,312-byte public key and a 2,420-byte signature. Compared with X25519 (32 bytes) and ECDSA P-256 (64-byte signature), that is roughly 34 to 38 times larger on the wire. Both standards let devices store a compact seed instead of the expanded private key, which reduces secure storage requirements.

Why is ML-DSA signing slow and unpredictable on microcontrollers?

ML-DSA signing uses rejection sampling: it generates a candidate signature, checks it against bounds that protect the secret key, and retries on failure. The number of attempts is random. In pqm4 results for ML-DSA-44, signing ranged from about 1.8 million to 17 million cycles over 1,000 runs, with a mean near 3.9 million. Verification is deterministic. Design timeouts around the tail, or sign on a gateway.

Should IoT firmware signing use LMS, XMSS or ML-DSA?

LMS and XMSS are standardized in NIST SP 800-208, rely only on hash security and have tiny public keys, but require careful state management. ML-DSA, standardized in FIPS 204, is stateless and easier to operate. Because firmware signatures are checked rarely and attached to large images, size matters little. Many designs use LMS or XMSS in the immutable root and ML-DSA for later stages. Follow your sector’s guidance, such as CNSA 2.0.

Does post-quantum TLS fit on constrained links like LoRaWAN or 802.15.4?

Not directly. A hybrid X25519MLKEM768 client key share is 1,216 bytes, and an ML-DSA certificate chain can add roughly 10 KB, while 802.15.4 frames carry 127 bytes at the physical layer. Fragmentation multiplies loss and airtime. Use session resumption, raw public keys, certificate compression or gateway termination, and avoid in-band PQC handshakes over LoRaWAN uplinks entirely.

Is hybrid ML-KEM with X25519 required for IoT devices?

Hybrid is recommended by the IETF draft for TLS because it stays secure if either component fails. It adds flash for both algorithms and about 1.2 KB of extra key-share bytes. NIST standardizes ML-KEM itself, and some regulated environments require ML-KEM-1024 on its own, so check your sector’s requirements. For devices that cannot afford both, a gateway-terminated design is often the better answer.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *