IoT Protocol Latency Benchmarks: MQTT vs CoAP vs AMQP vs HTTP/3 (Updated April 2026)

IoT Protocol Latency Benchmarks: MQTT vs CoAP vs AMQP vs HTTP/3 (Updated April 2026)

MQTT vs CoAP vs AMQP vs HTTP/3: IoT Protocol Latency Benchmarks and How to Run Your Own

Last Updated: October 2026

Most published comparisons of MQTT vs CoAP vs AMQP vs HTTP/3 crown a single winner with a table of milliseconds. Those tables are almost always the product of one rig, one network, and one set of defaults, and they do not transfer. A protocol that wins at 5 ms of round-trip time on a lab switch can lose by hundreds of milliseconds on a cellular link with 2% loss, because the dominant term in IoT latency is not the wire format. It is the number of round trips the protocol needs, and how it recovers when a packet disappears.

This rewrite replaces the April 2026 version, whose measurement tables I could not trace to a reproducible source and have withdrawn. What replaces them is more useful: a round-trip and byte-level model you can verify against the RFCs, the few published measurements that are real, and a harness you can run on your own devices and network.

What this covers: the latency budget for a single IoT message, cold-start and warm-path round trips for each protocol, header and security overhead with sourced figures, loss recovery timers, where MQTT over QUIC fits, a reproducible benchmark method, and a decision guide by constraint.

What Changed for October 2026

  • The April benchmark tables are withdrawn. The earlier post presented specific latency, throughput and microjoule-per-message figures from an ESP32-S3 rig. I could not tie them to published data or a repeatable artifact, so they are removed rather than updated. Treat any copy of those numbers elsewhere as unverified.
  • Errors in the earlier analysis are corrected. It credited QUIC with BBR congestion control as if that were part of the protocol. RFC 9002 specifies a NewReno-style default; BBR is an implementation choice. It also listed broker version numbers that I could not match to real release lines.
  • The method changed from “who wins” to “why”. Every latency figure here is either derived from round-trip counts in the RFCs, taken from a cited study, or labelled illustrative arithmetic.
  • MQTT over QUIC is now a first-class section. Vendor material from EMQX dated January 2026 pitches it for connected vehicles, and it changes the answer for roaming devices. It is still vendor-reported and not an OASIS-standardised transport, so I flag that clearly.
  • Security overhead is quantified. A peer-reviewed Computer Networks study measured how much DTLS inflates a CoAP exchange. Those numbers reshape the “CoAP is lightest” intuition.
  • New reproducible harness. The benchmark section now includes netem commands and measurement code, plus the pitfalls (coordinated omission, clock sync, cold versus warm) that invalidate most published tables.

Context and Background

Four protocols dominate the conversation about moving small messages from devices to a backend, and they come from different design lineages.

MQTT, the Message Queuing Telemetry Transport, is a publish-subscribe protocol over a long-lived TCP connection to a broker. MQTT 5.0 became an OASIS standard in March 2019 and added session expiry, reason codes, topic aliases and flow control. If you want the feature walk-through, our MQTT 5 features deep dive and the MQTT protocol technical guide cover them.

CoAP, the Constrained Application Protocol, is specified in RFC 7252. It is a compact request-response protocol modelled on REST, usually carried over UDP, with its own retransmission logic for reliability. Observe (RFC 7641) adds subscription-like notifications, block-wise transfer (RFC 7959) handles larger payloads, and RFC 8323 defines CoAP over TCP, TLS and WebSockets. Our CoAP deep dive goes further.

AMQP 1.0, the Advanced Message Queuing Protocol, is an OASIS standard (also published as ISO/IEC 19464) that defines a symmetric, link-based messaging layer with explicit credit-based flow control and delivery settlement. It is broker-agnostic by design. The AMQP complete guide and the AMQP vs MQTT comparison cover the model.

HTTP/3, defined in RFC 9114, is HTTP semantics mapped onto QUIC (RFC 9000), a UDP-based transport with integrated TLS 1.3 (RFC 9001). It is not an IoT protocol by origin, but it matters because constrained-device stacks increasingly reuse web infrastructure, and because QUIC’s properties are exactly what MQTT is now borrowing.

These are not four interchangeable options. They sit at different layers and make different promises, which is why the first step is to separate the transport and security handshake from the messaging exchange. A broader view of how they fit together lives in our IoT and industrial communication protocols architecture guide.

For external grounding, the primary documents are RFC 7252 (CoAP), RFC 9000 (QUIC), and the OASIS MQTT 5.0 specification.

The Latency Budget: Why Round Trips Beat Wire Format

The short answer: end-to-end latency for one IoT message equals the sum of connection setup round trips, security handshake round trips, the protocol’s own acknowledgement exchange, any loss-recovery wait, and broker queueing. On most real links the first four terms are multiples of the round-trip time (RTT), so protocol choice shifts latency by whole RTTs, not by microseconds of serialisation.

MQTT vs CoAP vs AMQP HTTP/3 IoT latency budget stages from sensor event to end-to-end percentile latency

Figure 1: The latency budget for one IoT message. Each stage is either a multiple of RTT, a timer, or a processing cost.

The diagram is the organising idea of this post. Serialisation of a 30-byte message at even 250 kbit/s takes about one millisecond. A single extra round trip over a cellular link with 80 ms RTT costs 80 ms. Protocols differ far more in the number of round trips and the timers they use than in how compactly they encode a header.

Cold start versus warm path

Every protocol has two latencies that benchmarks routinely blur together. The cold path is the first message after power-up, reconnection, or a long sleep. The warm path is a message on an already-established, already-authenticated session.

Devices that wake every few minutes and send one reading spend most of their lives on the cold path. A mains-powered gateway streaming telemetry spends nearly all of its time on the warm path. Comparing a warm MQTT connection to a cold HTTP/3 request, as some tables do, is comparing different questions.

Three components that do not scale with payload

Handshake round trips are fixed per connection. Acknowledgement round trips are fixed per message under reliable delivery. Retransmission timers are fixed by protocol constants and the estimated RTT. None of these grow with payload for small messages, which is why the 128-byte, 1 KB and 10 KB columns in many tables look almost identical, and why that sameness is a sign the table measures RTT rather than protocol.

For payloads that exceed a single packet, fragmentation and block-wise transfer start to matter. CoAP’s block-wise transfer sends one block per round trip by default in the simple case, so a 10 KB payload in 1,024-byte blocks costs roughly ten exchanges unless the implementation pipelines. TCP-based and QUIC-based protocols stream the bytes under congestion control instead. That difference, not header size, is where large payloads diverge.

Where processing time does and does not matter

Broker and server processing is real but usually small relative to a WAN RTT. It matters at scale, under fan-out, and when persistence is on, which is a throughput topic rather than a single-message latency topic. For broker capacity planning, see our EMQX Kubernetes cluster tutorial.

Round Trips by Protocol: Cold Start and Warm Path

This section derives round-trip counts from the specifications. These are counts of network round trips before the first acknowledged application message, assuming no packet loss, no connection reuse on the cold path, and TLS 1.3 or its equivalent. They exclude processing time and are my arithmetic from the RFCs, not measurements.

MQTT over TCP and TLS 1.3

A cold MQTT connection pays for TCP (one round trip for SYN, SYN-ACK), then TLS 1.3 (one round trip for the full handshake), then MQTT CONNECT and CONNACK. The MQTT specification allows a client to send further control packets immediately after CONNECT without waiting for CONNACK, so a well-written client can pipeline the first PUBLISH in the same flight as CONNECT.

MQTT vs CoAP vs AMQP vs HTTP/3 cold start sequence for MQTT over TCP and TLS 1.3 with pipelined publish

Figure 2: Cold-start sequence for MQTT over TCP and TLS 1.3. With a pipelined PUBLISH, the first PUBACK arrives after three round trips; without pipelining, four.

Once the session is warm, a QoS 0 publish needs no acknowledgement and is delivered to the broker after half an RTT. A QoS 1 publish completes in one RTT, when the PUBACK returns. QoS 2 uses a four-packet exchange (PUBLISH, PUBREC, PUBREL, PUBCOMP), costing two RTTs for full completion, which is the price of exactly-once semantics.

The practical MQTT latency lever is therefore session lifetime. A device that holds its TCP and TLS connection open pays the three-RTT cost once. A device that sleeps and reconnects pays it every time, though MQTT 5 helps: with a non-zero Session Expiry Interval and Clean Start set to 0, the broker keeps subscriptions and queued QoS 1 and 2 messages across the gap, avoiding re-subscribe round trips.

CoAP over UDP, DTLS and OSCORE

Plain CoAP has no connection. A Confirmable (CON) request is sent, and the server replies with an ACK, either empty and followed later by a separate response, or piggybacked with the response itself. The cold and warm costs are the same: one RTT. That is the source of CoAP’s reputation for low latency.

Secured CoAP changes this. With DTLS 1.3 (RFC 9147), a full handshake adds at least one round trip, and an additional one if the server demands a stateless cookie to defend against address spoofing. After that, DTLS Connection ID (RFC 9146) lets a device keep its session across NAT rebinding. An alternative is OSCORE (RFC 8613), which protects the CoAP message itself at the application layer. OSCORE needs no handshake round trips once the security context is established, which makes it attractive for sleepy devices, at the cost of key-management design.

So the honest CoAP numbers are: roughly one RTT for unsecured requests, about two to three RTTs cold for DTLS 1.3, and one RTT warm. The security-free figure is the one most benchmark tables quote, and it is the one least likely to match production.

AMQP 1.0 over TCP, TLS and SASL

AMQP 1.0 is the chattiest cold starter, because it layers several negotiation steps. After TCP and TLS, the client sends the 8-byte protocol header and, typically, performs a SASL exchange for authentication. It then sends Open, Begin (to create a session) and Attach (to create a link), and waits for the peer’s matching frames before it has link credit to transfer.

Depending on how much an implementation pipelines, that is roughly six to eight round trips before a first transfer is settled. I am labelling that as an estimate: it varies with library and broker, and it is exactly the kind of figure you should measure. Once warm, a transfer is settled with a Disposition frame in about one RTT for acknowledged delivery, or none if the sender uses pre-settled (at-most-once) transfers.

The upside of AMQP’s chatter is expressiveness. Link credit gives the receiver precise back-pressure, settlement modes distinguish at-most-once from at-least-once from exactly-once flows, and links can be long-lived and multiplexed on one session. The cost is a cold-start penalty that makes AMQP a poor fit for devices that reconnect frequently.

HTTP/3 over QUIC

QUIC merges transport and TLS 1.3 setup into a single round trip. A cold HTTP/3 request therefore costs about two RTTs: one for the QUIC handshake and one for the request and response. If the client holds a session ticket from an earlier connection, 0-RTT early data lets it send the request in the first flight, so the response returns after one RTT.

0-RTT has a security catch that matters for IoT. Early data is not protected against replay (RFC 8446, section 8), so a POST that triggers an actuator could be executed twice. RFC 8470 defines the Early-Data header and the 425 Too Early status so servers can refuse unsafe requests. For telemetry that is idempotent or deduplicated downstream, 0-RTT is safe. For commands, it is not, unless the application adds its own replay protection.

On a warm HTTP/3 connection, a request is one RTT, and QUIC’s independent streams mean a lost packet on one stream does not stall the others. Headers are compressed with QPACK (RFC 9204), but even a minimal POST carries tens of bytes of header data that a CoAP or MQTT message does not.

The comparison table

Scenario (no loss) MQTT 5 / TCP+TLS 1.3 CoAP CON / UDP CoAP / DTLS 1.3 AMQP 1.0 / TLS+SASL HTTP/3
First acked message, cold 3 RTT pipelined, 4 otherwise 1 RTT about 2-3 RTT about 6-8 RTT (estimate) 2 RTT
First acked message, with resumption 3 RTT (TLS resumption does not remove TCP and CONNECT) 1 RTT 1-2 RTT about 5-7 RTT (estimate) 1 RTT with 0-RTT
Warm acked message 1 RTT at QoS 1 1 RTT 1 RTT 1 RTT 1 RTT
Warm fire-and-forget 0.5 RTT delivery 0.5 RTT (NON) 0.5 RTT 0.5 RTT pre-settled not applicable to request-response

Now apply illustrative arithmetic. At a 30 ms RTT, a cold MQTT publish takes about 90 to 120 ms before processing, cold AMQP roughly 180 to 240 ms, and a cold unsecured CoAP request about 30 ms. At 100 ms RTT those become 300-400 ms, 600-800 ms and 100 ms. These are products of the table above, not measurements, and real figures add loss recovery and processing.

The warm path is a tie at the protocol level: one RTT for acknowledged delivery everywhere. That is why persistent-connection benchmarks look so similar across protocols, and why the differences that matter are in cold starts, loss, and overhead.

Bytes on the Wire: Headers, Security and Stack Layers

Latency and bytes are linked on constrained links. On a 250 kbit/s radio or a metered cellular plan, every byte costs air time and energy, and on NB-IoT or LTE-M, extra uplink packets can wake the radio for longer. Byte overhead is the second-order effect after round trips, and it is where CoAP and MQTT earn their reputations.

MQTT vs CoAP vs AMQP vs HTTP/3 protocol stack layers showing transport and security beneath each messaging protocol

Figure 3: The four stacks side by side. The messaging layer is the thin top; the transport and security layers underneath set most of the round-trip and byte cost.

Worked example: one 20-byte reading

Take a 20-byte sensor payload sent to a short topic or path. These byte counts come from the framing rules in the specifications and cover application framing only, before transport headers.

  • MQTT 5.0 PUBLISH at QoS 1: a 2-byte fixed header (type and flags, plus a one-byte remaining length), a 2-byte topic length plus the topic name (4 bytes for something like s/t1), a 2-byte packet identifier, a 1-byte properties length of zero, and the payload. That is 31 bytes. The PUBACK is 4 bytes, because MQTT 5 lets the reason code and properties be omitted when the result is success.
  • CoAP CON POST: the 4-byte fixed header, a token (2 bytes here), two Uri-Path options at roughly 5 bytes total, a 1-byte payload marker, and the 20-byte payload. About 32 bytes. An empty ACK is 4 bytes.
  • HTTP/3 POST: framing depends on QPACK state. The first request on a connection carries more literal header bytes than later ones, because QPACK builds a dynamic table. Expect tens of bytes of header data on top of the payload, plus HTTP/3 frame headers.
  • AMQP 1.0 Transfer: an 8-byte frame header, a performative with several fields (handle, delivery-id, delivery-tag and flags), then the message sections with their own descriptors. It is larger than MQTT for the same reading, though the exact size depends on the encoder and which message sections are present.

Below the application layer, TCP adds a 20-byte minimum header against UDP’s 8 bytes. TLS 1.3 adds a 5-byte record header, a one-byte inner content type and a 16-byte authentication tag per record for the common AEAD suites. QUIC short-header packets carry a connection ID and a packet number and also a 16-byte tag. Over IPv4, there are another 20 bytes for IP.

The total for one QoS 1 MQTT publish over TLS is therefore around 31 plus 22 plus 20 plus 20, or roughly 93 bytes in the data direction, plus acknowledgement packets. A CoAP request over plain UDP is around 32 plus 8 plus 20, or 60 bytes. Over DTLS 1.3, add the record overhead. These are my arithmetic, and they ignore TCP options, delayed ACK behaviour and Ethernet framing.

What security does to the “lightweight” claim

Security is where the lightweight ranking shifts. A 2021 study in Computer Networks, “Performance evaluation of CoAP and MQTT with security support for IoT environments”, measured bandwidth usage with and without protection. For an unsecured CoAP exchange in the best case, the paper reports 134 bytes. With DTLS using a pre-shared key it reports 1,769 bytes, and with certificate-based DTLS 2,299 bytes. That is roughly a 13-fold and a 17-fold increase for a single exchange, dominated by the handshake.

The lesson is not that CoAP is heavy. It is that comparing unsecured CoAP with TLS-secured MQTT, which many tables do, compares different security postures. A fair comparison secures both, or secures neither and says so. The same paper found security substantially raised CPU use, with certificate-based suites costing more than pre-shared keys.

Session resumption and long-lived sessions amortise this. A DTLS session that survives for hours, kept alive with Connection ID, spreads its handshake over thousands of messages. A sleepy device that must re-handshake every wake-up does not benefit, which is why OSCORE, or a TCP/TLS session kept warm by a gateway, often wins in battery designs.

MQTT-SN and the constrained end

For the very constrained end of the spectrum, MQTT has a sibling, MQTT-SN, a separate specification that runs over UDP and replaces long topic strings with short numeric topic identifiers. It is not part of the OASIS MQTT 5.0 standard, so support varies by broker and gateway. If you are weighing MQTT against CoAP on a sleepy radio, MQTT-SN is a more honest comparison to CoAP than TCP-based MQTT. For LPWAN radio choices underneath these protocols, see our LoRaWAN vs NB-IoT vs LTE-M comparison.

When Packets Disappear: Loss Recovery and Tail Latency

The p50 latency of a protocol on a clean link is dull. The p99 on a lossy link is where protocols diverge, and it is set by retransmission timers.

Timer constants

  • TCP: the initial retransmission timeout is 1 second under RFC 6298, and Linux enforces a 200 ms minimum RTO by default. Fast retransmit triggers on three duplicate ACKs and avoids the timeout when enough packets follow the lost one, which is often not the case for sparse IoT traffic.
  • CoAP: RFC 7252 sets ACK_TIMEOUT to 2 seconds and ACK_RANDOM_FACTOR to 1.5, so the first retransmission of a Confirmable message waits a random time between 2 and 3 seconds. The timeout doubles on each retry, with MAX_RETRANSMIT of 4. A single lost packet therefore costs two to three seconds of latency by default.
  • QUIC: RFC 9002 uses a probe timeout (PTO) derived from the RTT estimate, and recommends an initial RTT of 333 ms before any sample exists. Once RTT is measured, recovery is much faster than CoAP’s fixed defaults.

The consequence is counter-intuitive. On a lossy link, a CoAP CON request with default parameters can have worse tail latency than MQTT over TCP, despite CoAP’s one-RTT best case. CoAP’s timers are tunable, and RFC 7252 explicitly allows deployments to change them, but the defaults assume very slow networks.

Sparse traffic breaks TCP’s fast recovery

Fast retransmit needs later packets to generate duplicate ACKs. A sensor that sends one message every ten seconds has no later packets, so a lost segment waits for the RTO. For this traffic pattern, TCP’s loss recovery is closer to CoAP’s than the textbook suggests, and tail latency is set by timers, not RTT.

QUIC’s tail-loss probes address this directly by sending a probe when no acknowledgement arrives, rather than waiting a full RTO. That is a structural advantage for sparse, loss-prone traffic, independent of any 0-RTT benefit.

Head-of-line blocking

MQTT uses one TCP byte stream. If a segment carrying a large retained message or a big payload is lost, later packets for unrelated topics queue behind the retransmission, even if they have arrived. HTTP/3 and MQTT over QUIC both give each stream independent delivery, so loss on one stream does not delay others. AMQP 1.0 multiplexes sessions and links, but over TCP it inherits the same head-of-line blocking at the transport layer.

This matters for mixed workloads, for example a gateway that publishes both 50-byte telemetry and occasional 500 KB diagnostic dumps over a single connection. The dump delays the telemetry under TCP. Separate connections, or QUIC streams, remove the coupling.

MQTT over QUIC: The Hybrid That Changes the Roaming Case

MQTT over QUIC keeps MQTT’s pub-sub semantics and runs the packets over QUIC instead of TCP and TLS. EMQX documents it as a supported transport, and its open-source embedded broker NanoMQ and the NanoSDK client library also support it, with a fallback to TCP and TLS when QUIC is blocked. As far as I can tell it remains a vendor-led extension rather than part of the OASIS MQTT 5.0 specification, so interoperability across vendors is not guaranteed.

What QUIC contributes

Based on the vendor’s own technical description, there are four relevant properties. First, a combined transport and TLS 1.3 handshake of one round trip, or zero round trips on resumption. Second, connection migration, so a client that moves from Wi-Fi to cellular keeps its connection instead of reconnecting. Third, independent streams that remove cross-topic head-of-line blocking. Fourth, better behaviour under loss, because of QUIC’s per-packet loss detection and tail-loss probes.

For a roaming asset, a vehicle, or a handheld that changes networks, migration is the biggest win. Under TCP and TLS, an IP change kills the connection, and the client pays the full three-RTT reconnect plus any re-subscription. Under QUIC, the connection ID carries the session across the change.

Vendor-reported measurements, with caveats

EMQX published a test on EMQX 5.0 on a single AWS EC2 node with 5,000 clients. The vendor reports that at a 30 ms RTT, QUIC connections were significantly faster than TLS, and that in a mass-reconnect scenario CPU was about 65% for QUIC against about 75% for TLS, and memory about 9 GB against 12 GB. It also reports that with 10% packet loss and 20% out-of-order packets injected, QUIC kept delivering while TLS showed congestion and message loss.

Read those as directional. They are vendor numbers from a single setup, I did not reproduce them, and the blog post gives no confidence intervals. The same post notes real costs: QUIC needs UDP to pass through middleboxes, the initial handshake uses more bandwidth (the vendor measured a 100 Mb peak, attributed to path MTU validation), and 0-RTT early data is disabled by default because of replay risk. The vendor itself says QUIC suits mobile and frequently switching networks better than stable, low-latency ones.

That last sentence is the right framing. On a stable wired link with a persistent connection, QUIC buys little over TCP and TLS 1.3. On a roaming or lossy link, it removes whole classes of reconnect and stall cost.

Operational realities

QUIC runs in user space, so CPU cost per packet has historically been higher than kernel TCP with offload, although the vendor’s reconnect test reports lower broker CPU. UDP is often rate-limited or blocked on corporate and industrial firewalls, and some operators apply short NAT timeouts to UDP flows, which are carrier-dependent and should be measured. A QUIC deployment needs a TCP fallback path, and your benchmark should include the fallback case.

Designing a Reproducible Benchmark

A benchmark is only useful if someone else can re-run it and get the same shape of answer. Most published IoT tables fail on at least one of the points below.

Principles

  1. Separate cold and warm. Report them as different experiments with different questions.
  2. Control the network, not just the protocol. Emulate RTT, jitter and loss on the path with tc netem, and record the settings. A clean LAN measures stack overhead, not protocol design.
  3. Secure every stack equally. Use TLS 1.3, DTLS 1.3 or OSCORE, or none, consistently, and state which.
  4. Measure round-trip time at the client with a monotonic clock. For acknowledged delivery, the client sends, the ACK returns, and the same clock measures both ends. This avoids clock-sync error. One-way latency needs synchronised clocks, such as PTP, and an error budget.
  5. Avoid coordinated omission. If you send the next message only after the previous completes, a stall hides the delays other messages would have seen. Use an open-loop schedule that sends on a fixed timetable and records intended send time.
  6. Report percentiles from a proper histogram. p50 alone hides everything. Report p95, p99 and p99.9 with sample counts, and use a histogram library such as HdrHistogram.
  7. Pin versions and publish configs. Record broker, client library, kernel, TLS library and their settings. Defaults differ: Nagle, delayed ACK, keep-alive, QoS and CoAP timers all change results.

The network profile

On the Linux host acting as the path, add delay and loss with netem. The command below adds 40 ms each way (80 ms RTT), 5 ms jitter and 1% random loss on the interface facing the devices.

# Emulate a mediocre cellular path on the test interface
sudo tc qdisc replace dev eth0 root netem delay 40ms 5ms distribution normal loss 1%

# Inspect, and later remove
tc qdisc show dev eth0
sudo tc qdisc del dev eth0 root

Run each protocol across a small grid: RTT of 10, 40 and 120 ms; loss of 0%, 1% and 5%; and payloads of 32 bytes and 4 KB. That is 18 cells per protocol per mode, which is manageable and tells you more than a single configuration.

A minimal open-loop MQTT probe

The Python sketch below measures warm QoS 1 round-trip latency using the paho-mqtt 2.x client. It sends on a fixed schedule and records latency from the intended send time, which is how you avoid coordinated omission. It is a starting point rather than a production harness.

import time, threading
import paho.mqtt.client as mqtt

BROKER, PORT, TOPIC = "broker.example.net", 8883, "bench/t1"
RATE_HZ, DURATION_S, PAYLOAD = 20, 60, b"x" * 32
lat_us, intended = [], {}
lock = threading.Lock()

def on_publish(client, userdata, mid, reason_code, properties):
    now = time.monotonic_ns()
    with lock:
        t0 = intended.pop(mid, None)
        if t0 is not None:
            lat_us.append((now - t0) // 1000)

c = mqtt.Client(mqtt.CallbackAPIVersion.VERSION2, protocol=mqtt.MQTTv5)
c.tls_set()                      # TLS 1.3 when broker and OpenSSL allow
c.on_publish = on_publish
c.connect(BROKER, PORT, keepalive=60)
c.loop_start()

period_ns = int(1e9 / RATE_HZ)
start_ns = time.monotonic_ns()
for i in range(RATE_HZ * DURATION_S):
    due_ns = start_ns + i * period_ns
    wait = (due_ns - time.monotonic_ns()) / 1e9
    if wait > 0:
        time.sleep(wait)
    with lock:
        # record the INTENDED send time so queueing delay is counted
        info = c.publish(TOPIC, PAYLOAD, qos=1)
        intended[info.mid] = due_ns

time.sleep(2)
c.loop_stop()
lat_us.sort()
for p in (50, 95, 99, 99.9):
    print(p, lat_us[int(len(lat_us) * p / 100) - 1], "us")

For CoAP, the aiocoap library provides a CON client; for HTTP/3, aioquic provides an HTTP/3 client; for AMQP 1.0, Apache Qpid Proton or the AMQP 1.0 client in your broker’s ecosystem works. Keep the load shape identical across clients. The cold-start experiment is different: close the connection, wait, and time from “connect call” to “first ACK received”, repeated a few hundred times.

Energy: measure it, do not model it

Energy per message depends on radio state, not on the protocol alone. A device that keeps the radio on for a second to wait for a TCP ACK spends far more than the cost of the bytes. Measure with a current profiler (a Nordic Power Profiler Kit or a source-measure unit) over complete wake cycles including connect, send, wait and sleep. Report joules per cycle and cycles per battery life, not a micro-joule-per-message figure from a lab radio held in transmit mode. I do not publish an energy table here because I have no reproducible measurements to cite.

Throughput is a separate experiment

Single-message latency and sustained throughput are different questions. For throughput, hold payload and QoS constant, ramp the number of concurrent clients, and report the knee where p99 latency departs from p50. Broker implementation and configuration dominate that result more than the protocol does. For an example of an end-to-end pipeline with time-series storage, see our real-time asset tracking tutorial with MQTT 5.

Trade-offs, Gotchas, and What Goes Wrong

Benchmarking the broker instead of the protocol. If the MQTT test uses one broker, the AMQP test another and the HTTP/3 test a web server, you have measured three servers. Differences in persistence, fsync policy, and thread model can dwarf the protocol. Where possible, use one system that speaks multiple protocols, or acknowledge the confound.

Defaults that quietly decide the winner. Nagle’s algorithm with delayed ACK can add tens of milliseconds to small TCP writes unless TCP_NODELAY is set, which most MQTT clients do by default but not all. CoAP’s 2-3 second first retransmission is a default, not a law. Keep-alive intervals determine how quickly a dead connection is noticed. Report every value you change.

Idle-connection death. NAT and firewall state expires. A warm MQTT connection behind a carrier NAT can silently die between keep-alives, so the “warm” 1-RTT latency becomes a cold 3-RTT one plus a timeout while the client discovers the failure. Keep-alive intervals must be shorter than the shortest middlebox timeout on the path, which you should measure, not assume. UDP-based protocols often face shorter NAT bindings than TCP, so CoAP and QUIC need more frequent keep-alives or a Connection ID strategy.

0-RTT replay. QUIC and TLS 1.3 early data can be replayed by an attacker who captured it. For idempotent telemetry that is acceptable. For commands that actuate equipment, it is a safety issue. Use RFC 8470’s Early-Data mechanism, or design commands with sequence numbers and deduplication. This is also why EMQX ships MQTT over QUIC with early data off by default.

Fan-out and subscription cost. The latency of a publish that reaches 10,000 subscribers is a function of the broker’s delivery path, queue depth and persistence settings, not of the client protocol. MQTT 5 shared subscriptions and flow control through Receive Maximum help but also change ordering. For design guidance on this layer, see our DDS vs MQTT vs OPC UA comparison, and for payload semantics, Sparkplug B vs plain MQTT topics.

Treating HTTP/3 as an IoT messaging protocol. HTTP/3 gives you request-response over fast transport. It does not give you broker-mediated fan-out, retained messages, last-will, or session state. If you build those on top, you are reinventing a broker, and your latency story changes. HTTP/3 is a good fit when devices already talk to cloud APIs and you want QUIC’s connection behaviour without a new protocol.

Constrained-device library limits. Many microcontroller stacks implement QUIC poorly or not at all, and DTLS 1.3 support is still uneven across embedded TLS libraries. Check that your chosen library and hardware crypto accelerator support the exact suites you plan to use, because falling back to software crypto changes the CPU and energy picture.

Anti-pattern: one-shot connect-publish-disconnect. Opening a TLS connection per reading is the most expensive way to use MQTT or AMQP, and it is common in sleepy-device firmware. If a device must sleep, consider CoAP with OSCORE, MQTT-SN through a gateway, or a long-lived connection held by a gateway on the device’s behalf. A gateway pattern is covered in our IIoT edge gateway architecture guide.

Practical Recommendations

Choose by the constraint that dominates, and verify with your own network profile. Do not choose by a published winner table.

MQTT vs CoAP vs AMQP vs HTTP/3 decision guide mapping device constraints to recommended IoT protocol

Figure 4: Pick by the dominant constraint. This is a starting hypothesis to validate with a benchmark, not a verdict.

For sleepy, battery-powered devices on UDP-friendly radios, CoAP is the natural fit, ideally with OSCORE or DTLS 1.3 with Connection ID, and with tuned retransmission timers. For always-on fleets that talk through a broker, MQTT 5 over TLS 1.3 is the pragmatic default because of its ecosystem, session handling and tooling. For enterprise integration where settlement semantics, link credit and interoperability with message brokers matter more than cold-start time, AMQP 1.0 earns its overhead. For roaming or lossy links, test MQTT over QUIC, or HTTP/3 if you are API-centric, and keep a TCP fallback.

Checklist before you commit:

  • [ ] Record the real RTT, jitter, loss and NAT timeout of your deployment network.
  • [ ] Decide whether the dominant workload is cold start, warm streaming, or bursty reconnects.
  • [ ] Secure every candidate stack identically and state the suites.
  • [ ] Run the 3x3x2 grid of RTT, loss and payload, with an open-loop load generator.
  • [ ] Report p50, p95, p99 and p99.9, with sample counts and histogram files.
  • [ ] Measure energy over whole wake cycles with a current profiler.
  • [ ] Test the failure path: broker restart, network handover, and QUIC-to-TCP fallback.
  • [ ] Decide how to protect commands from replay before enabling any 0-RTT feature.

Frequently Asked Questions

Which is faster for IoT, MQTT or CoAP?

It depends on the path, not the protocol. On a warm session both complete an acknowledged message in about one round trip. CoAP over plain UDP has the cheapest cold start at one round trip, but once you add DTLS the gap narrows, and on lossy links CoAP’s default 2-3 second retransmission can give it worse tail latency than MQTT over TCP. Tune timers and measure on your own network before deciding.

Is HTTP/3 good for IoT devices?

It is good for devices that already speak HTTP to cloud APIs and need fast reconnects or network migration. QUIC’s one-round-trip handshake and optional 0-RTT resumption reduce cold-start cost, and independent streams reduce head-of-line blocking. It lacks broker features like retained messages and fan-out, and constrained microcontrollers often lack mature QUIC libraries, so check your hardware first.

Does MQTT over QUIC replace MQTT over TCP?

No, it complements it. It helps most on roaming, lossy, or frequently reconnecting links, because of connection migration and fast resumption. On stable, persistent links it adds little. It is a vendor-led extension supported by EMQX and NanoMQ rather than a part of the OASIS MQTT 5.0 specification, and UDP can be blocked by firewalls, so you need a TCP fallback.

Why are AMQP 1.0 connections slow to establish?

AMQP 1.0 layers several negotiation steps after TCP and TLS: protocol header exchange, usually SASL authentication, then Open, Begin and Attach frames to create the connection, session and link before credit allows a transfer. Depending on pipelining that is roughly six to eight round trips by my estimate. It is fine for long-lived links and a poor fit for devices that reconnect often.

Is 0-RTT safe for IoT commands?

Not by default. TLS 1.3 and QUIC early data can be replayed by an attacker, so a command that triggers an actuator could run twice. RFC 8470 provides the Early-Data header and a 425 Too Early response so servers can reject unsafe requests. Use 0-RTT for idempotent telemetry, and add sequence numbers or deduplication before using it for commands.

How should I benchmark IoT protocols fairly?

Emulate your real network with tc netem, secure all stacks equally, and measure round trips at the client with a monotonic clock. Use an open-loop sender to avoid coordinated omission, separate cold from warm runs, and report p50 to p99.9 from a histogram. Record versions and configs, and measure energy over complete wake cycles with a current profiler.

Further Reading

Related posts on this site:

External primary sources:

Sources I could not verify directly: the full text of the ACM 2025 paper “Comparative Analysis and Implementation of HTTP3, MQTT, and CoAP for IoT Applications” (access was blocked), so I do not cite its figures.

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *