TinyML Microcontrollers: TensorFlow Lite Micro, INT8 Quantization and CMSIS-NN Inference
A vibration sensor on a pump costs a few dollars, runs for years on a coin cell, and produces thousands of samples a second. Streaming that raw signal to the cloud kills the battery and the radio budget long before it tells you anything useful. The alternative is to run the neural network on the sensor node itself, on a chip with a few hundred kilobytes of memory and no operating system. That is the promise of TinyML microcontrollers: inference measured in milliwatts, decisions made where the data is born, and only the verdict leaving the device.
The promise is real but the engineering is unforgiving. A model that trains happily in float32 on a laptop will not fit, will not link, or will quietly lose ten points of accuracy once it is squeezed into 8-bit integers. Every kilobyte of flash and SRAM has to be accounted for before you flash the board, not after.
This tutorial walks the whole path, from a Keras model to bytes in flash to a running interpreter, and explains why each step exists.
What this covers: the memory model of a microcontroller, how TensorFlow Lite Micro works, INT8 quantization arithmetic, post-training quantization versus quantization-aware training, tensor arena sizing, CMSIS-NN, Ethos-U and ESP-NN acceleration, MLPerf Tiny, runnable conversion and C++ code, failure modes, and a decision matrix against Edge Impulse, ExecuTorch and microTVM.
Context and Background
TinyML is the practice of running machine learning inference on devices whose total memory is measured in kilobytes and whose power budget is measured in milliwatts. The usual target is a 32-bit microcontroller (MCU) built around an Arm Cortex-M core, an Espressif Xtensa or RISC-V core, or similar. These parts typically have between tens of kilobytes and a few megabytes of flash, between tens and a few hundred kilobytes of SRAM, clock speeds from tens to a few hundred megahertz, and no memory management unit, no filesystem, and often no operating system.
That is a different universe from the one covered in our guide to edge AI inference on NVIDIA Jetson, Intel Movidius and Arm NPUs. A Jetson-class module has gigabytes of RAM, runs Linux, and can host a full PyTorch or TensorRT runtime. A microcontroller cannot host a general-purpose framework at all. It needs a runtime that allocates nothing dynamically, links only the operators the model uses, and keeps every tensor in a block of memory that was sized at build time.
The reference runtime for this world is TensorFlow Lite for Microcontrollers, usually shortened to TensorFlow Lite Micro or TFLM. Google’s documentation, which now brands the wider stack as LiteRT for Microcontrollers, states that the core runtime fits in 16 KB on an Arm Cortex-M3, needs no operating system, no standard C or C++ libraries and no dynamic memory allocation, and is written in C++17 for 32-bit platforms. The same page is blunt about the limits: only a subset of TensorFlow operations is supported, only a limited set of devices is supported, the low-level C++ API requires manual memory management, and on-device training is not supported. The open source project lives at tensorflow/tflite-micro under the Apache-2.0 licence, and its continuous integration targets Cortex-M, RISC-V, Hexagon and Xtensa.
Why does anyone accept these constraints? Three forces push inference down to the MCU. The first is energy: transmitting a bit over a radio typically costs far more energy than computing on it locally, so a node that classifies locally and transmits only events can last orders of magnitude longer than one that streams. The second is latency and availability: a keyword spotter or a safety interlock cannot wait for a round trip. The third is privacy: audio and video that never leave the device cannot leak from a server.
The field is also standardising. MLPerf Tiny, introduced by a community of more than 50 industry and academic organisations, defines four tasks (keyword spotting, visual wake words, image classification and anomaly detection) and measures accuracy, latency and energy. We return to it later. For now the key point is that TinyML is no longer a hobbyist curiosity: it has a reference runtime, vendor-optimised kernels, and a benchmark.
This post assumes you can train a small model in Keras. If you want a prior grounding in what precision does to accuracy on larger edge accelerators, read INT4 vs INT8 vs FP8 on edge NPUs first; the numeric ideas carry over, but the microcontroller version is stricter because there is no floating point fallback to hide behind.
Reference Architecture: From Keras Model to Bytes in Flash

Figure 1: The TinyML microcontrollers pipeline. A float model is converted to an INT8 flatbuffer, embedded in firmware as a C array, and executed by the micro interpreter against a statically allocated tensor arena.
The pipeline in Figure 1 has eight stages, but only three decisions matter: what the model looks like, how it is quantized, and how memory is laid out. Everything else is mechanical. A trained float32 Keras model is passed to the TensorFlow Lite converter together with a small representative dataset, which the converter uses to measure activation ranges. The converter emits a .tflite file, a FlatBuffer that holds the operator graph, the quantization parameters and the weights. That file is turned into a C array with a tool such as xxd -i, compiled into the firmware image, and stored in flash. At run time the micro interpreter reads the flatbuffer directly from flash, with no parsing copy, and executes operators one at a time against activations held in SRAM.
The memory model decides what fits
On a server, “does the model fit” means one number. On a microcontroller it means three, because flash and SRAM are separate resources with different jobs.

Figure 2: Where each artefact lives. Weights and code sit in read-only flash. The tensor arena, interpreter state, stack and sensor buffers compete for the same small SRAM.
Flash holds the program: your application, the TFLM runtime, the kernels for the operators you registered, and the model weights as a constant array. SRAM holds everything that changes at run time: the tensor arena (intermediate activations and scratch buffers), the interpreter’s persistent bookkeeping, the stack, and your sensor and feature buffers. A common beginner mistake is to count only the model file against flash and forget that the runtime and kernels, plus a data-logging stack or Bluetooth stack, may consume as much again. The reverse mistake is to count only the weights against SRAM when the real consumer is the largest pair of activation tensors alive at the same moment.
In practice SRAM is the binding constraint for convolutional models and flash is the binding constraint for dense or recurrent ones. A depthwise-separable convolutional network for audio might have only tens of thousands of parameters, so its weights fit easily, while its first layer’s feature maps dominate memory. A fully connected autoencoder has tiny activations but many weights. Knowing which regime you are in tells you which knob to turn: shrink spatial resolution and channel counts for the first, shrink width and depth for the second.
Why the interpreter is a good fit for MCUs
TFLM keeps the interpreter model, a loop over a list of operators, rather than generating one monolithic C function per network. This looks wasteful on a device where every byte counts, but it buys three things. The model is data, so you can update it over the air without relinking the firmware, as long as the operator set is unchanged. Operators are shared, so a network that reuses Conv2D eight times pays for the kernel code once. And the memory planner can reason about the whole graph, reusing arena space for tensors whose lifetimes do not overlap.
The cost is a small interpretation overhead per operator and the need to register operators explicitly. You declare a MicroMutableOpResolver and add only the kernels the model uses, so unused operators never link. This is how a small model keeps the code footprint low. Code generation approaches exist and we compare them in the decision matrix, but the interpreter remains the default because it is the path every silicon vendor tests.
Operator coverage is the first compatibility check
Before any optimisation, check that every operator in your .tflite file has a TFLM kernel. The supported set is a subset of TensorFlow Lite: standard convolution, depthwise convolution, fully connected, pooling, softmax, reshape, add, mul, logistic, tanh, quantize and dequantize, and a growing list of others. Exotic layers, dynamic shapes, control flow and most recurrent cells either are missing or are only supported in restricted forms. The practical method is to list the operators in the converted model with a flatbuffer viewer, compare them to the TFLM operator list in the repository, and redesign the network where there is a gap. Replacing a Swish activation with ReLU6, or a GlobalAveragePooling that the converter lowers to unsupported reductions with an explicit AveragePool2D, is typical. Do this check before you spend days training.
INT8 Quantization: The Arithmetic That Makes It Fit
Microcontrollers in the Cortex-M family often have no floating-point unit at all (Cortex-M0, M0+, M3), or only a single-precision one (M4F, M7, M33). Even with an FPU, float32 weights are four times the size of INT8 weights and float arithmetic costs more cycles and more energy. INT8 quantization is therefore not an optimisation on this class of hardware; it is the baseline.
The affine mapping
TensorFlow Lite uses an affine quantization scheme. A real value is represented as real = scale * (q - zero_point), where q is an 8-bit integer. The scale is a positive float that sets the step size, and the zero_point is the integer that represents real zero. For INT8 the integer range is -128 to 127. The TensorFlow Lite quantization specification fixes the conventions used by the converter and the kernels: weights are quantized symmetrically (zero point of 0) and, for convolutions, per output channel; activations are quantized asymmetrically per tensor.

Figure 3: The INT8 integer-only path. Inputs are quantized, multiplied in INT8, accumulated in INT32, biased, then requantized to INT8 using a fixed-point multiplier and shift, with no floating point at run time.
Consider why per-channel matters. In a convolution with 64 output channels, the weight ranges of different filters can differ by an order of magnitude. A single per-tensor scale would be set by the largest filter, leaving the small filters represented by only a handful of integer levels. Per-channel scales give each filter its own range. The cost is storage for one scale per channel and a slightly more complex requantization, which the kernels already handle. For small, depthwise-heavy networks, per-channel quantization is frequently the difference between a tolerable accuracy loss and an unusable model.
How integer-only inference works
Look at the path in Figure 3. The input activation is already INT8. The kernel multiplies INT8 weights by INT8 activations and accumulates the products in INT32, because the sum of many 8-bit products overflows 8 bits immediately. The INT32 bias, pre-scaled to the product of the input and weight scales, is added. Then the accumulator has to return to the INT8 output scale. The ratio input_scale * weight_scale / output_scale is a real number, but the kernel never touches a float: the converter encodes it as a 32-bit fixed-point multiplier plus a right shift, and the kernel performs a rounding multiply-and-shift followed by a clamp. That is the whole trick. The entire network runs on integer multiply-accumulate instructions, which is exactly what the DSP and SIMD extensions on Cortex-M cores accelerate.
A worked example makes the memory saving concrete. Take a hypothetical convolution layer with 3×3 kernels, 32 input channels and 64 output channels. It has 3 x 3 x 32 x 64 = 18,432 weights. In float32 that is 73,728 bytes; in INT8 it is 18,432 bytes, plus 64 per-channel scales and 64 INT32 biases. The layer is four times smaller in flash before you do anything clever. (These sizes are arithmetic on an illustrative layer, not measurements from a specific model.)
Post-training quantization versus quantization-aware training
There are two routes to an INT8 model. Post-training quantization (PTQ) takes a finished float model and a few hundred representative samples, runs them through the network to record the min and max of each activation, and derives scales and zero points. It needs no retraining and takes minutes. Quantization-aware training (QAT) inserts fake-quantization nodes into the training graph so the network learns weights that survive rounding; the TensorFlow Model Optimization Toolkit provides this for Keras.
For most MCU models, start with PTQ. Small convolutional networks for keyword spotting or gesture recognition often lose only a small amount of accuracy under full-integer PTQ with a good calibration set. Move to QAT when PTQ drops accuracy more than your tolerance, which happens most often in three situations: networks with very few channels, where each weight matters; networks with activation ranges dominated by outliers; and regression or autoencoder outputs, where small numerical errors shift the reconstruction error you are thresholding on. Treat the claim “PTQ loses under one point” as something you measure on your own validation set, not a rule. The right habit is to evaluate the float model, the PTQ model and, if needed, the QAT model with the same held-out data through the same metric.
Calibration data is a design decision
The representative dataset is the most underestimated input in the pipeline. The converter chooses each activation’s range from the samples you give it. If the calibration set contains only clean, in-distribution recordings, the ranges will be too tight, and a noisy production signal will saturate the INT8 range and clip. If it contains rare extreme outliers, the ranges will be too wide, and ordinary signals will use only a few quantization levels. Use a few hundred samples that span the real operating conditions: every machine state, every noise condition, every sensor orientation you expect. Resist the temptation to calibrate on training data that has been augmented beyond what the sensor will ever see.
Deeper Analysis: Arena Sizing, Kernels and a Runnable Walk-through
Sizing the tensor arena
The tensor arena is one contiguous, aligned byte array that you declare in firmware, for example alignas(16) static uint8_t tensor_arena[40 * 1024];. TFLM carves everything the interpreter needs out of it: the activation tensors, per-operator scratch buffers requested by kernels, and the interpreter’s own persistent structures. Nothing is allocated from a heap, which is why the runtime can promise deterministic memory behaviour.
The activation portion is determined by the graph’s liveness. At each operator, the input tensors and the output tensor must coexist, and any tensor consumed by a later operator must survive until then. TFLM’s memory planner packs tensors with non-overlapping lifetimes into the same addresses. The peak is therefore not the sum of all activations but the maximum over time of the live set. For a plain feed-forward chain, that peak is roughly the largest adjacent input-plus-output pair. For a network with skip connections, such as a residual block, the skip tensor stays alive across the block and raises the peak.
A hypothetical audio model shows the arithmetic. Suppose the input is a 49 by 10 matrix of mel-frequency cepstral coefficients, one channel, stored as INT8: 490 bytes. A first convolution with 64 filters, stride 2, produces roughly a 25 by 5 by 64 map: 8,000 bytes. The two tensors coexist during that operator, so the peak so far is about 8.5 KB. Later layers with the same channel count and smaller spatial size are cheaper, so the first layer sets the peak. Add scratch space for the convolution kernel, which depends on the optimised kernel in use, plus a few kilobytes of persistent overhead, and the arena lands in the low tens of kilobytes. (Illustrative arithmetic, not a measurement.) Estimates like this are good for design; the final number comes from the interpreter.
The practical procedure is empirical. Start with a generous arena, call AllocateTensors(), then read interpreter.arena_used_bytes() and shrink the array to that value plus a small safety margin. If AllocateTensors() returns an error, the arena is too small or an operator is missing; the error reporter prints which. Because the arena is statically sized, an oversized one wastes SRAM that the rest of your application needs, and an undersized one fails at start-up rather than at run time, which is the failure mode you want.
Two levers shrink the arena without changing accuracy. The first is operator ordering and graph shape: for branches, the converter’s output order affects lifetimes. The second is reducing the first layer’s output: larger strides, fewer filters, or pooling early all cut the peak linearly. A third lever is more radical: restructure computation to stream. Streaming models, such as the 1D depthwise-separable network that MLPerf Tiny v1.3 added for wake-word detection in continuous audio, process audio in small chunks and keep a small state, so the arena does not scale with the window length.
From Keras to a quantized flatbuffer
The conversion script below defines a small depthwise-separable keyword-spotting style network, converts it with full-integer quantization, and checks the result. Replace the random arrays with your own features and labels. The calls to tf.lite.TFLiteConverter and representative_dataset are standard TensorFlow Lite APIs.
import numpy as np
import tensorflow as tf
NUM_CLASSES = 12
INPUT_SHAPE = (49, 10, 1) # time frames x MFCC coefficients x 1 channel
def build_model():
inp = tf.keras.Input(shape=INPUT_SHAPE)
x = tf.keras.layers.Conv2D(64, (10, 4), strides=(2, 2),
padding="same", use_bias=True)(inp)
x = tf.keras.layers.BatchNormalization()(x)
x = tf.keras.layers.ReLU()(x)
for _ in range(3):
x = tf.keras.layers.DepthwiseConv2D((3, 3), padding="same")(x)
x = tf.keras.layers.BatchNormalization()(x)
x = tf.keras.layers.ReLU()(x)
x = tf.keras.layers.Conv2D(64, (1, 1))(x)
x = tf.keras.layers.BatchNormalization()(x)
x = tf.keras.layers.ReLU()(x)
x = tf.keras.layers.AveragePooling2D(pool_size=x.shape[1:3])(x)
x = tf.keras.layers.Flatten()(x)
out = tf.keras.layers.Dense(NUM_CLASSES)(x) # logits, softmax in firmware
return tf.keras.Model(inp, out)
model = build_model()
model.compile(optimizer="adam",
loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),
metrics=["accuracy"])
# model.fit(train_x, train_y, validation_data=(val_x, val_y), epochs=30)
def representative_dataset():
# A few hundred real, in-distribution feature windows.
for sample in calibration_x[:300]:
yield [sample[np.newaxis, ...].astype(np.float32)]
converter = tf.lite.TFLiteConverter.from_keras_model(model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_dataset
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.int8
converter.inference_output_type = tf.int8
tflite_model = converter.convert()
with open("kws_int8.tflite", "wb") as f:
f.write(tflite_model)
print("model bytes:", len(tflite_model))
# Verify: run the quantized model with the TFLite interpreter on the host.
interp = tf.lite.Interpreter(model_content=tflite_model)
interp.allocate_tensors()
inp = interp.get_input_details()[0]
out = interp.get_output_details()[0]
scale, zero_point = inp["quantization"]
def quantize_input(x):
q = np.round(x / scale + zero_point)
return np.clip(q, -128, 127).astype(np.int8)
correct = 0
for x, y in zip(val_x, val_y):
interp.set_tensor(inp["index"], quantize_input(x[np.newaxis, ...]))
interp.invoke()
correct += int(np.argmax(interp.get_tensor(out["index"])) == y)
print("int8 accuracy:", correct / len(val_y))
Setting supported_ops to the INT8 built-ins makes the converter fail loudly if any operator cannot be quantized, instead of silently leaving a float island that your firmware would then have to support. Setting the input and output types to INT8 removes the quantize and dequantize operators at the model boundary, so the firmware feeds integers straight in; you take responsibility for applying the input scale and zero point, which the script above demonstrates. Compare the printed INT8 accuracy to the float model’s accuracy on the same validation set before you go further. Then turn the file into a C array with xxd -i kws_int8.tflite > model_data.cc and mark the array const and alignas(16) so it stays in flash and is aligned.
Running it on the device
The firmware skeleton below follows the structure used in the TFLM examples. Check the header paths and the MicroInterpreter constructor against the TFLM revision you vendor, because the project evolves and signatures have changed over time.
#include "tensorflow/lite/micro/micro_interpreter.h"
#include "tensorflow/lite/micro/micro_mutable_op_resolver.h"
#include "tensorflow/lite/schema/schema_generated.h"
#include "model_data.h" // g_kws_int8_model[] generated by xxd
namespace {
constexpr int kArenaSize = 40 * 1024;
alignas(16) uint8_t tensor_arena[kArenaSize];
tflite::MicroMutableOpResolver<6> resolver;
tflite::MicroInterpreter* interpreter = nullptr;
TfLiteTensor* input = nullptr;
TfLiteTensor* output = nullptr;
} // namespace
bool setup_model() {
const tflite::Model* model = tflite::GetModel(g_kws_int8_model);
if (model->version() != TFLITE_SCHEMA_VERSION) return false;
resolver.AddConv2D();
resolver.AddDepthwiseConv2D();
resolver.AddAveragePool2D();
resolver.AddFullyConnected();
resolver.AddReshape();
resolver.AddSoftmax();
static tflite::MicroInterpreter static_interpreter(
model, resolver, tensor_arena, kArenaSize);
interpreter = &static_interpreter;
if (interpreter->AllocateTensors() != kTfLiteOk) return false;
input = interpreter->input(0);
output = interpreter->output(0);
// interpreter->arena_used_bytes() reports the real requirement.
return true;
}
int classify(const int8_t* features, int n) {
for (int i = 0; i < n; ++i) input->data.int8[i] = features[i];
if (interpreter->Invoke() != kTfLiteOk) return -1;
int best = 0;
for (int i = 1; i < output->dims->data[1]; ++i) {
if (output->data.int8[i] > output->data.int8[best]) best = i;
}
return best; // argmax on INT8 logits needs no dequantization
}
Notice that the final argmax runs on INT8 logits directly. Because the output scale is positive, ordering is preserved, so you only need to dequantize when you want a calibrated probability or a confidence threshold. When you do, apply real = scale * (q - zero_point) using the output tensor’s params.
Kernel acceleration: CMSIS-NN, Ethos-U and ESP-NN
The reference kernels in TFLM are portable C++ written for correctness. They run everywhere and are slow compared with what the silicon can do. The performance story on MCUs is therefore about swapping in optimised kernels that exploit the instruction set or an attached accelerator, while keeping the same numerical results.
CMSIS-NN is Arm’s library of neural network kernels for Cortex-M processors. Its repository describes the goal as maximising performance and minimising memory footprint, and states that it follows the int8 and int16 quantization specification of TFLM, with int4 weights supported for some operators and floating point described as experimental. It has three implementation tiers: pure C for every operator, which is what runs on Cortex-M0 and M3; SIMD kernels using the DSP extension on cores such as Cortex-M4 or an M33 with DSP; and kernels using the M-profile Vector Extension, branded Helium, on cores such as Cortex-M55 and M85. The library is described as bit-exact with the TensorFlow Lite reference kernels, and where TFL and TFLM reference kernels differ it follows TFLM. That bit-exactness is operationally important: you can validate accuracy on the host with the reference kernels and trust that the optimised kernels on the device produce the same integers.
With TFLM, CMSIS-NN is selected at build time through the optimized-kernel option for the target, so your application code does not change. The practical effect is that the same model runs substantially faster on a DSP-capable or Helium core than on the reference path. How much faster depends on the model, the compiler, the clock and the memory system, so I will not quote a multiplier here; measure on your board.
Arm Ethos-U is a different category. The Ethos-U family are small neural processing units, called microNPUs, designed to sit next to a Cortex-M core and offload the heavy operators. The model is compiled ahead of time with Arm’s Vela compiler, which rewrites supported operators into a command stream for the NPU and leaves unsupported ones on the CPU through CMSIS-NN. The consequence for model design is that INT8 (or INT16 activations) is required, and operator support on the NPU is narrower than on the CPU, so the same compatibility check applies twice. Confirm the exact supported operators and the NPU variants in Arm’s current Vela documentation before committing to a design.
ESP-NN is Espressif’s optimised neural network function library. Its repository lists assembly kernels that use the vector instructions of the ESP32-S3, SIMD optimisations for the ESP32-P4, and generic optimisations for the ESP32 and ESP32-C3, and reports interpreter timings with and without ESP-NN on INT8 models such as person detection. Like CMSIS-NN, it plugs under TFLM so the application is unchanged. If your product is built on an ESP32-S3, it is the first thing to enable.
The common thread is that acceleration happens below the interpreter, keeps INT8 semantics, and is chosen per target. Your job is to pick a chip whose accelerator matches the operators in your model, not the other way round.
MLPerf Tiny: what the benchmark does and does not tell you
MLPerf Tiny, published at NeurIPS Datasets and Benchmarks in 2021, is presented as the first industry-standard benchmark suite for ultra-low-power machine learning systems. Per its abstract, it measures accuracy, latency and energy, and its modular design lets submitters showcase products at any layer of the deployment stack. It defines four tasks: keyword spotting, visual wake words, image classification and anomaly detection.
The v1.3 round, announced by MLCommons on 17 September 2025, reported 70 results across five benchmark tests from four participants (Kai Jiang, Qualcomm, STMicroelectronics and Syntiant), 27 of them power results, along with a new open-source test harness and a new streaming wake-word test built on a 1D depthwise-separable CNN. Five hardware platforms were benchmarked for the first time. I did not verify per-device scores, so none are quoted here.
Use MLPerf Tiny for what it is good at: comparing hardware and software stacks on fixed, public reference models under fixed rules, with energy measured. Do not use it to predict your product’s battery life. Your model, your sensor front end, your sleep schedule and your radio will dominate, and the benchmark measures only the inference window. It also pins the model and the accuracy target, so it tells you little about how a platform behaves on a custom architecture or an operator the reference models do not use.
Architectures that fit
Three families dominate MCU deployments, and all three exist because they are cheap in activations as well as weights.
Depthwise-separable CNNs (DS-CNN) split a convolution into a per-channel spatial filter and a 1×1 pointwise mix, cutting multiply-accumulates and parameters by roughly the kernel area for the same channel counts. They are the standard for keyword spotting, popularised by work on small-footprint keyword spotting on Cortex-M, and for visual wake words with MobileNet-style variants. Their weakness is that depthwise layers are memory-bound and quantize poorly under per-tensor schemes, which is another reason to insist on per-channel weights.
Small dense autoencoders compress a feature vector through a narrow bottleneck and reconstruct it; the reconstruction error is the anomaly score. They are trained only on normal data, which is exactly what you have for healthy machinery. They are tiny in activations and moderate in weights, and map to Fully Connected operators that every accelerator supports.
Tiny 1D CNNs and streaming models operate on time series such as accelerometer and audio streams, and keep a rolling state so the arena stays constant regardless of window length.
Use Case: Vibration Anomaly Detection on a Sensor Node

Figure 4: A vibration anomaly detector for TinyML microcontrollers. The MCU computes FFT features, scores them with an INT8 autoencoder, compares reconstruction error to a threshold, and wakes the radio only on an anomaly.
Predictive maintenance is the canonical industrial TinyML workload, and it connects directly to the condition-monitoring layer of a digital twin: the node supplies a cheap, continuous health signal, and the twin decides what to do with it. Here is a concrete design.
A three-axis MEMS accelerometer is bolted to a motor housing and sampled at a few kilohertz. The firmware takes a window, applies a Hann window function, computes a fast Fourier transform (CMSIS-DSP provides optimised FFT routines on Cortex-M), and keeps the log magnitude of a fixed set of frequency bins as the feature vector. The feature vector is quantized with the model’s input scale and zero point and fed to a dense autoencoder: input layer, a few shrinking dense layers to a small bottleneck, and mirrored layers back out. The autoencoder is trained offline on recordings from healthy operation only. At run time, the firmware computes the mean squared error between the input and the reconstruction, and compares it to a threshold chosen from the distribution of reconstruction errors on held-out healthy data, for example a high percentile plus margin.
Several design points make or break this system.
First, the threshold lives outside the model. Because the model’s output is a reconstruction, small quantization errors shift the baseline error, and a threshold calibrated on the float model will be wrong for the INT8 model. Always calibrate the threshold by running the quantized model, ideally on the device or the host with identical kernels.
Second, put the feature extraction under the same scrutiny as the network. Windowing, scaling and log operations in fixed-point or float32 have to match training exactly. A mismatch between Python feature code and C feature code is the most common reason a model that scores well offline is useless on hardware.
Third, duty cycle the whole pipeline. The accelerometer can run in a low-power mode and wake the MCU on a data-ready interrupt; the MCU wakes, computes features, runs inference in a few milliseconds to tens of milliseconds depending on the model and clock, and goes back to sleep. Radio use is rare and event-driven. This is where the energy advantage over streaming materialises, and it is the number to measure on your own prototype with a power analyser.
Fourth, a single-window anomaly score is noisy. Smooth it, for example by requiring several consecutive windows above threshold, or by tracking an exponential moving average, before alerting. This trades detection delay for fewer false alarms and is usually the right trade for slow-developing mechanical faults.
If you later need to move more computation to a gateway, the sibling article on Jetson Thor versus Hailo-10H versus Coral covers the next tier up, and MLPerf Edge agentic inference v6.1 on Jetson Thor shows how the benchmark story changes at that scale.
Alternatives to TFLM
TFLM is the default, not the only option, and the right choice depends on your team and silicon.
Edge Impulse is a hosted platform that wraps data collection, feature extraction (its DSP blocks), training, quantization and an exportable C++ inference library, which itself uses TFLM and vendor kernels under the hood for many targets. It is the fastest route from raw sensor data to a working prototype and a poor fit if you need full control of the training code or cannot use a hosted service. Check current licensing and deployment terms on their site.
ExecuTorch is PyTorch’s on-device runtime. It targets mobile and embedded devices and includes backends for Arm hardware including the Ethos-U NPUs; I did not verify its current microcontroller maturity or operator coverage in this run, so evaluate it directly if your team lives in PyTorch.
microTVM, part of Apache TVM, compiles models to C code for microcontrollers, and offers ahead-of-time code generation with a memory planner. I could not confirm its current maintenance status in this run, so check the TVM repository’s recent activity before betting a product on it.
Vendor toolchains such as STMicroelectronics’ X-CUBE-AI and similar offerings from other silicon vendors convert models for their own parts and often produce the best result on that silicon at the price of portability.
Trade-offs, Gotchas, and What Goes Wrong
Quantization drift. The most insidious failure is a model that passes validation in float, passes in INT8 on the host, and degrades slowly in the field. Activation ranges were learned from calibration data; when real inputs drift, such as a different microphone, a sensor aged by temperature, or a new operating mode, values clip against the learned INT8 range and accuracy falls without any error being raised. Monitor input statistics on the device: track the fraction of input values that saturate at -128 or 127 and the running mean and variance of the features, and report them with each alert.
Sensor domain shift. Your training data came from one set of units in one place. The deployed sensors have different mounting, different gains, different noise floors. A vibration model trained on one motor does not necessarily transfer to the next of the same model, let alone a different frame size. Collect data from the real installation, and plan for per-site calibration of at least the threshold.
Feature mismatch between training and firmware. Python and C implementations of an FFT, a mel filterbank or a log differ in window definition, scaling, rounding and precision. Test the firmware’s feature output against the Python reference on the same raw input and require they match within a tight tolerance. Do this before you debug the network.
Operator gaps discovered late. A model with an unsupported operator fails at AllocateTensors(), not at conversion. Add the check to continuous integration: convert, then run the model through a host build of TFLM with the exact operator resolver you ship.
Arena fragmentation by neighbours. Arena size assumes the rest of the application leaves enough SRAM. A Bluetooth stack, a display buffer or a logging queue can crowd it out. Budget SRAM for the whole firmware image, and re-measure after every feature added.
Accuracy metrics that hide the failure. Accuracy is a poor metric for rare events. For anomaly detection, report precision and recall at the chosen threshold, plus false alarms per day under healthy operation, which is the number operators actually experience.
Over-trusting the benchmark. MLPerf Tiny numbers describe reference models under fixed rules. They are a guide to relative hardware capability, not a prediction of your product’s performance.
Security and updates. A model in flash is an asset that can be extracted and an attack surface if updated over the air. Sign model updates, and store the quantization parameters with the model so a mismatched pair cannot be flashed.
Practical Recommendations
Start from the constraint, not the model. Write down your flash, SRAM, clock, active power and latency targets, and the memory the non-ML firmware needs. What remains is your budget. Then choose a chip: if you need real throughput, favour a core with DSP or Helium extensions or an attached microNPU, and check that the vendor supports TFLM kernels for it.
Design the network for the budget. Prefer depthwise-separable convolutions or small dense networks, use ReLU or ReLU6 rather than exotic activations, keep the first layer’s output small, and check operator coverage on day one. Train in float, convert with full-integer per-channel quantization, and compare float and INT8 accuracy on the same held-out set. Escalate to QAT only if the gap is unacceptable.
Treat the surrounding pipeline as part of the model. Pin the feature extraction code, test it against the Python reference, and calibrate thresholds on the quantized model. Measure arena use with arena_used_bytes(), then measure energy with a real power analyser, over a full duty cycle that includes the sensor and radio.
A short checklist to run before every release:
- Every operator in the
.tflitefile is registered in the resolver and tested on a host build. - INT8 accuracy is within tolerance of float on a held-out set that matches deployment conditions.
- Per-channel weight quantization is on, and the calibration set spans real operating conditions.
- Arena size equals the measured use plus margin; total SRAM is budgeted for the whole firmware.
- Firmware features match the Python reference within tolerance.
- Thresholds were calibrated on the quantized model.
- Input saturation and drift counters are reported in telemetry.
- Model updates are signed and versioned together with their quantization parameters.
Decision Matrix: Choosing a TinyML Stack
| Option | Best for | Strengths | Limits |
|---|---|---|---|
| TFLM with CMSIS-NN | Cortex-M products, production teams | Reference runtime, bit-exact optimised kernels, broad vendor support, no heap | Limited operator set, manual memory management, no on-device training |
| TFLM with Ethos-U | Cortex-M plus microNPU designs | Offloads convolutions, ahead-of-time compilation | Narrower NPU operator support, INT8 or INT16 only, toolchain complexity |
| TFLM with ESP-NN | ESP32-S3 and related Espressif parts | Vendor-tuned INT8 kernels, large ecosystem | Tied to Espressif silicon |
| Edge Impulse | Fast prototyping, mixed-skill teams | End-to-end data to firmware workflow | Hosted dependency, less control of training |
| ExecuTorch | PyTorch-first teams, Arm NPU targets | PyTorch export path | MCU maturity not verified here, check before use |
| microTVM | Teams wanting compiled C and custom schedules | Ahead-of-time code generation | Maintenance status not verified here |
| Vendor tools such as X-CUBE-AI | Single-vendor products | Often best on that silicon | Lock-in to one vendor |
Frequently Asked Questions
What is TinyML and how is it different from edge AI?
TinyML is machine learning inference on microcontrollers with kilobytes of memory and milliwatt power budgets, usually without an operating system. Edge AI is the broader category that also includes Linux-class devices such as Jetson modules, which have gigabytes of RAM and run full frameworks. The practical difference is the runtime: TinyML needs static memory allocation, a minimal operator set and INT8 models, while edge AI can use general frameworks and larger models.
How much memory does TensorFlow Lite Micro need?
Google’s documentation says the core runtime fits in 16 KB on an Arm Cortex-M3, but that excludes your model weights, the kernels you link, and the tensor arena. Real totals depend on the model: weights go in flash, activations and scratch go in the arena in SRAM. Call AllocateTensors() and read arena_used_bytes() to get the true arena requirement, then add the rest of your firmware.
Should I use post-training quantization or quantization-aware training?
Start with post-training INT8 quantization with per-channel weights and a representative calibration set; it takes minutes and often costs little accuracy on small convolutional networks. Move to quantization-aware training when the measured accuracy gap is too large, which tends to happen with very narrow networks, outlier-heavy activations, or autoencoders whose reconstruction error you threshold. Always measure on a held-out set rather than assuming.
What is the tensor arena and how do I size it?
The tensor arena is a statically allocated, aligned byte array from which TFLM carves activation tensors, kernel scratch buffers and interpreter state. Size it by starting generous, running AllocateTensors(), reading the used bytes from the interpreter, then shrinking the array to that value plus a small margin. A too-small arena fails at start-up, not at run time, so the failure is easy to catch.
What does CMSIS-NN do for TinyML microcontrollers?
CMSIS-NN is Arm’s library of optimised neural network kernels for Cortex-M cores. It provides pure C kernels for every operator, SIMD kernels for cores with the DSP extension, and vector kernels for cores with Helium. It follows the TFLM int8 and int16 quantization specification and is bit-exact with the reference kernels, so enabling it speeds inference without changing results.
Can I train a model on a microcontroller?
Not with TFLM: Google’s documentation lists on-device training as unsupported. The standard pattern is to train on a workstation or in the cloud, convert to INT8, and deploy a frozen model. Adaptation on the device is limited to light techniques such as updating a threshold or a small calibration layer in your own code, and anything heavier belongs on a gateway or in the cloud.
Further Reading
- Edge AI inference on NVIDIA Jetson, Intel Movidius and Arm NPUs: the next tier up from the microcontroller.
- INT4 vs INT8 vs FP8 quantization on edge NPUs: precision trade-offs beyond INT8.
- Jetson Thor vs Hailo-10H vs Coral for edge inference: choosing a gateway-class accelerator.
- MLPerf Edge agentic inference v6.1 on Jetson Thor: how edge benchmarks are read.
- LiteRT for Microcontrollers overview: Google’s official documentation for the runtime.
- CMSIS-NN on GitHub: Arm’s kernel library and operator support tables.
- MLPerf Tiny Benchmark paper: the benchmark’s tasks and methodology.
By Riju — about
