Jetson Orin Nano 2 vs Orin Nano Super: 2x Inference, Same Socket
Last Updated: September 23, 2026
NVIDIA says its new entry-level module doubles inference performance. The peak compute number, though, rises only 16 percent, from 67 to 78 TOPS. Both statements are true, and the gap between them is the whole story of Jetson Orin Nano 2 vs Orin Nano Super. Doubling comes from a new Orin-family chip with better Tensor Cores, faster memory and, according to third-party reporting, a new 40-watt power mode. Whether you actually see 2x depends on whether your workload is limited by compute or by memory bandwidth.
That distinction matters right now because Orin Nano 2 is announced but not shipping. NVIDIA expects the module and developer kit in the first half of 2027, with no price disclosed. Teams planning 2027 robots, drones and smart cameras must decide today whether to design around the Super or wait.
What this covers: a confirmed-versus-reported spec breakdown, the mechanics of where the 2x comes from, worked tokens-per-second estimates for edge LLMs, power and thermal implications, a migration path, and a buy-now-or-wait decision framework.
Context and Background
The Orin Nano line has had an unusual history. The original Jetson Orin Nano 8GB launched in 2023 at 40 sparse INT8 TOPS, with a 1,024-core Ampere GPU running at 635 MHz and 68 GB/s of LPDDR5 bandwidth. In December 2024, NVIDIA turned it into the “Orin Nano Super” with a software update rather than new silicon. The NVIDIA technical blog announcing the Super boost lists the result: GPU clock to 1,020 MHz, memory to 102 GB/s, and 67 sparse TOPS, at a developer kit price cut from $499 to $249.
That software unlock made the Super the default entry point for edge generative AI. It is the board most hobbyists, researchers and many startups use for small language models, vision language models and classic computer vision. It is also what Wing, Alphabet’s drone delivery subsidiary, flies today, according to NVIDIA’s announcement. The Super competes with dedicated accelerators; our Hailo-10H vs Jetson Orin Nano comparison for edge computer vision covers that rivalry for camera pipelines.
On 25 August 2026, NVIDIA announced Jetson Orin Nano 2, calling it a “robotics computer” for entry-level edge AI. The official headline numbers are 78 TOPS, 8 GB of memory, and an 8-core Arm CPU. NVIDIA claims 2x the inference performance of the Super “through improved Tensor Cores and higher memory bandwidth,” in the same compact form factor. It also claims that in 15-watt mode, Orin Nano 2 uses 40 percent less power to match its predecessor’s performance.
Everything beyond those sentences comes from secondary sources. ServeTheHome reported details from an NVIDIA press briefing, including that this is a new Orin-architecture SoC rather than a re-binned original die. JetsonHacks published a spec table. Partner Antmicro confirmed LPDDR5X memory, an 8-core Arm Cortex-A78 cluster and the “same compact SO-DIMM outline.” This post keeps those tiers separate, because hardware decisions made on unconfirmed specs are how product schedules slip.
What Orin Nano 2 Actually Changes: Confirmed vs Reported
Jetson Orin Nano 2 is a new Ampere-based Orin SoC in the Orin Nano’s 260-pin SO-DIMM form factor. NVIDIA confirms 78 TOPS, 8 GB of memory, an 8-core Arm CPU, and 2x inference over Orin Nano Super. Third parties report LPDDR5X at 120 GB/s, a 15–40 W range, 1,536 CUDA cores, PCIe Gen4 and hardware video encode.
The confirmed-versus-reported ledger
Before any spec comparison, sort the facts by who stated them. The table below is the most useful artefact in this post if you are writing a design review.
| Attribute | Orin Nano 2 value | Source tier |
|---|---|---|
| AI compute | 78 TOPS | NVIDIA press release (confirmed) |
| Memory capacity | 8 GB | NVIDIA press release (confirmed) |
| CPU | 8-core Arm | NVIDIA press release (confirmed) |
| CPU core type | Cortex-A78 | Antmicro partner blog (partner-confirmed) |
| Memory type | LPDDR5X | Antmicro partner blog (partner-confirmed) |
| Form factor | Same compact SO-DIMM outline | NVIDIA and Antmicro (confirmed) |
| Performance claim | 2x inference vs Orin Nano Super | NVIDIA press release (confirmed) |
| Efficiency claim | 40% less power at 15 W for same performance | NVIDIA press release (confirmed) |
| Availability | Module and dev kit, 1H 2027 | NVIDIA press release (confirmed) |
| Memory speed and bandwidth | LPDDR5X-7500, 120 GB/s | ServeTheHome, JetsonHacks (reported) |
| Power range | 15 W to 40 W | ServeTheHome, JetsonHacks (reported) |
| CUDA cores | 1,536 | JetsonHacks (reported); ServeTheHome prints “1536?” |
| PCIe | Gen4 | JetsonHacks (reported) |
| Hardware video encode | Present | JetsonHacks (reported) |
| New SoC, not binned Orin | Yes | ServeTheHome, citing NVIDIA briefing (reported) |
| Price | Not disclosed | — |
| JetPack version | Not announced | — |
| DLA presence | Not stated | — |
| FP8 or FP4 support | No — Ampere GPU; INT8 is the lowest precision NVIDIA quotes | Architecture fact, noted by ServeTheHome |
Two confirmed details deserve attention. First, NVIDIA’s release mentions a “significant leap in AI and video-processing performance.” That wording is consistent with the reported addition of a hardware encoder, which the original Orin Nano notably lacked, but it is not an explicit confirmation. Second, the CPU is described by Antmicro as Cortex-A78, whereas the Super uses Cortex-A78AE. The “AE” variant supports split-lock operation for functional safety. Whether Orin Nano 2 keeps it is unconfirmed, and anyone building to IEC 61508 or ISO 26262 should ask NVIDIA directly.
Side-by-side spec table
The table below compares the three Orin Nano generations plus the Orin NX 16GB, which is the natural “step up” alternative. Orin Nano 2 figures marked with an asterisk are reported, not NVIDIA-confirmed.
| Spec | Orin Nano 8GB (2023) | Orin Nano Super (2024) | Orin Nano 2 (1H 2027) | Orin NX 16GB (Super mode) |
|---|---|---|---|---|
| Sparse INT8 TOPS | 40 | 67 | 78 | 157 (GPU plus DLA) |
| GPU | Ampere, 1,024 CUDA, 32 Tensor | Ampere, 1,024 CUDA, 32 Tensor | Ampere, 1,536 CUDA* | Ampere, 1,024 CUDA, 32 Tensor |
| GPU clock | 635 MHz | 1,020 MHz | Not disclosed | Higher in Super mode |
| CPU | 6x Cortex-A78AE, 1.5 GHz | 6x Cortex-A78AE, 1.7 GHz | 8x Cortex-A78 | 8x Cortex-A78AE |
| Memory | 8 GB LPDDR5 | 8 GB LPDDR5 | 8 GB LPDDR5X | 16 GB LPDDR5 |
| Bandwidth | 68 GB/s | 102 GB/s | 120 GB/s* | 102 GB/s |
| Power modes | 7–15 W | 7, 15, 25 W | 15–40 W* | 10–40 W |
| PCIe | Gen3 | Gen3 | Gen4* | Gen4 |
| Video encode | None (software only) | None (software only) | Hardware encode* | Hardware encode |
| DLA | Disabled | Disabled | Not stated | 2x NVDLA |
| Availability | Shipping | Shipping | 1H 2027 | Shipping |
One number in that table is easy to misread. Orin NX TOPS include the two Deep Learning Accelerator (DLA) engines, whereas the Orin Nano figures are GPU-only. The NX’s GPU is the same 1,024-core Ampere block as the Super’s, so for GPU-bound LLM work, NX’s advantage is mostly its 16 GB of memory rather than raw speed.
Why a new die matters more than the TOPS headline
The Super was a clock-and-voltage unlock of existing silicon. Orin Nano 2, per ServeTheHome’s account of NVIDIA’s briefing, is a new chip with LPDDR5X memory controllers and “architecture and microarchitecture” changes to the Ampere Tensor Cores. That has three consequences that do not show up in a TOPS figure.
First, a new memory controller is the only way to move from LPDDR5-6400 to LPDDR5X-7500. The 128-bit bus stays the same width, so bandwidth rises by exactly the ratio of transfer rates: 7,500 / 6,400 = 1.17, or 102.4 GB/s to 120 GB/s. Second, Tensor Core changes can raise delivered throughput without raising peak. If the old Tensor Cores stalled waiting for operands at batch size one, fixing the data path lifts utilisation, which is where a 2x delivered claim can hide behind a 16 percent peak increase. Third, new silicon means TensorRT engines built for the Super will not load on Orin Nano 2. TensorRT engines are tied to the specific GPU they were built on, so every deployed model needs a rebuild and re-validation.
What the reported core count implies
If the 1,536 CUDA core figure holds, there is an interesting inference to draw. Peak TOPS scale with the number of Tensor Cores multiplied by clock and per-core throughput. Assuming Tensor Cores scale with CUDA cores (48 instead of 32) and per-core throughput is unchanged, 78 TOPS would imply a GPU clock of about 1,020 MHz × (78 / 67) / 1.5 ≈ 790 MHz. That is the author’s estimate, not a disclosed figure.
The assumption is shaky. NVIDIA explicitly says the Tensor Cores changed, and ServeTheHome’s write-up refers to the new chip’s “32 tensor cores”, which would contradict a 48-Tensor-Core layout. With 32 Tensor Cores and unchanged per-core throughput, the same arithmetic gives about 1,020 × (78 / 67) ≈ 1,190 MHz instead. The two reported figures do not sit neatly together, which is one more reason to wait for NVIDIA’s datasheet. But the arithmetic shows why the core count matters. A wider, slower GPU is typically more power-efficient than a narrow, fast one, which fits NVIDIA’s 15-watt efficiency claim. It also means CUDA-core work that bypasses Tensor Cores, such as image preprocessing, non-maximum suppression and custom kernels, could gain more than the TOPS figure suggests. At 1,536 cores × 2 FLOPs × 0.79 GHz, FP32 throughput would be about 2.4 TFLOPS, versus about 2.1 TFLOPS on the Super (1,024 × 2 × 1.02 GHz). Again, treat this as a scenario, not a spec.
Where the 2x Actually Comes From
NVIDIA’s claim is 2x delivered inference. ServeTheHome reports a crucial qualifier from the briefing: the 2x compares Orin Nano 2 in its 40 W mode against the Super at its 25 W peak. The 40 percent power saving is a different comparison, Orin Nano 2 at 15 W against the Super at 25 W, at equal performance. Keep those two claims apart, because they describe opposite ends of the power curve.

Figure 1: Where NVIDIA’s 2x claim comes from, split into compute-bound prefill and bandwidth-bound decode.
The diagram separates the four levers. Peak TOPS, Tensor Core efficiency and extra power budget all feed compute-bound work: the prefill phase of an LLM, vision encoders and convolutional detectors. Memory bandwidth, and to a lesser degree power mode, feeds the decode phase, which dominates tokens-per-second in chat and agent workloads. The 2x headline is most achievable on the compute-bound side.
The two phases of LLM inference
A transformer language model does two very different kinds of work. During prefill, it processes the entire prompt in one batched pass. Every weight is read from memory once and multiplied against hundreds or thousands of token vectors, so the arithmetic intensity is high and the Tensor Cores are the bottleneck. Prefill determines time-to-first-token.
During decode, the model produces one token at a time. Every weight must be read from DRAM again for each new token, but it is multiplied against only one vector. Arithmetic intensity collapses to roughly one or two operations per byte, far below what the GPU can sustain. The Tensor Cores sit idle waiting for memory, so tokens per second is capped by bandwidth divided by the bytes read per token.

Figure 2: The token-generation loop on an 8 GB unified-memory Jetson. Each decode step re-reads the weights and the KV cache from LPDDR5X.
The sequence shows why unified memory matters. The CPU, GPU and every accelerator share the same LPDDR5X pool. A camera pipeline, a ROS 2 stack and the LLM all compete for the same 120 GB/s (reported), so the bandwidth available to decode is always less than the datasheet number.
Worked example: decode ceiling for an 8B model
The Super gives us calibration data. NVIDIA’s own measurements show Llama 3.1 8B at INT4 with MLC rising from 14 tokens per second on the original Orin Nano to 19.14 on the Super. Memory bandwidth rose 1.5x (68 to 102 GB/s); tokens per second rose 1.37x. That tracking is exactly what a bandwidth-bound workload looks like.
Now the author’s estimate. An 8-billion-parameter model at 4 bits per weight is about 4.0 GB of weights. Quantisation scales and the higher-precision output projection add overhead, so assume roughly 4.5 GB read per token. The theoretical ceiling on the Super is 102 / 4.5 ≈ 22.7 tokens per second. The measured 19.14 is about 84 percent of that ceiling, which is a realistic efficiency for a well-tuned runtime.
Apply the same 84 percent to Orin Nano 2’s reported 120 GB/s: 120 / 4.5 × 0.84 ≈ 22.4 tokens per second. That is a 1.17x improvement, not 2x. Even at a physically implausible 100 percent efficiency, the ceiling is 26.7 tokens per second, or 1.39x the Super’s measured figure. On single-stream 8B decode, bandwidth alone cannot deliver 2x.
The same arithmetic for Llama 3.2 3B: about 2.0 GB per token gives a Super ceiling of 51 tokens per second against a measured 43.07, again about 84 percent. On Orin Nano 2, 120 / 2.0 × 0.84 ≈ 50 tokens per second, up from 43. The ratio is identical because both models are bandwidth-bound. For measured numbers on today’s boards, see our edge LLM benchmark of Llama, Phi and Gemma on Jetson Orin.
Worked example: where 2x is plausible
Prefill is different. A forward pass costs about 2 × parameters FLOPs per token. A 1,000-token prompt through an 8B model is therefore 2 × 8×10⁹ × 1,000 = 16 TFLOP of work. With INT4 weight-only quantisation, the matrix maths usually runs in FP16 after dequantisation, and NVIDIA lists the Super at 17 FP16 TFLOPS.
At an assumed 40 percent utilisation, the Super sustains about 6.8 TFLOPS, so the prompt takes 16 / 6.8 ≈ 2.4 seconds to prefill. If Orin Nano 2 doubles delivered compute, as NVIDIA claims, that drops to about 1.2 seconds. These are the author’s estimates built on an assumed utilisation; NVIDIA has not published FP16 TFLOPS for Orin Nano 2.
For a voice assistant, a robot receiving a long system prompt, or a vision language model encoding an image into several hundred tokens, time-to-first-token is the latency users feel. Halving it is a real product improvement even if steady-state tokens per second rises only 17 percent. Retrieval-augmented agents that stuff documents into context are dominated by prefill, and so are agent loops that re-send tool output on every step.
Why KV cache makes bandwidth even more important
Decode reads more than weights. It also reads the key-value (KV) cache, which stores attention state for every previous token. For Llama 3.1 8B, with 32 layers, 8 KV heads of dimension 128 and FP16 storage, each token costs 2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes, or 128 KiB.
At 4,096 tokens of context the KV cache is 512 MiB. Each decode step now reads about 5.0 GB rather than 4.5 GB, dropping the Super’s ceiling to 102 / 5.0 ≈ 20.4 tokens per second and Orin Nano 2’s to 120 / 5.0 ≈ 24. At 8,192 tokens, the cache reaches 1 GiB. That is one-eighth of total system memory on either module, before the operating system, the CUDA context and your perception stack have taken their share.
This is why the 8 GB capacity, unchanged between the Super and Orin Nano 2, is arguably the more binding constraint for 2027 workloads. NVIDIA names Gemma 4, Qwen 3, Nemotron and Cosmos as target models. Fitting a 7–8B model plus a useful context window plus a vision encoder into 8 GB of shared memory takes aggressive quantisation and careful budgeting on both modules.
Power Modes, Efficiency and Thermal Design
The two NVIDIA claims describe a performance-per-watt curve with two measured points. Reading them correctly changes how you size a product’s power and thermal budget.
Reading the 40 percent claim
“40 percent less power for the same performance in 15-watt mode” implies the comparison point is the Super at 25 W, because 15 / 25 = 0.60. So Orin Nano 2 at 15 W matches the Super at its maximum. That is a performance-per-watt improvement of 25 / 15 ≈ 1.67x at the low end of the curve.
At the high end, 2x the performance at 40 W versus 25 W gives 2.0 / (40 / 25) = 1.25x performance per watt. Efficiency gains shrink as you push the clock, which is exactly what voltage-frequency physics predicts: dynamic power scales roughly with voltage squared times frequency, and higher frequencies need higher voltages. The practical implication is that battery-powered systems gain the most from Orin Nano 2, not the ones that run flat out.
The 7 W gap
The Super supports a 7 W mode. The reported minimum for Orin Nano 2 is 15 W. If that holds, a solar-powered camera, a small inspection drone or a sensor node designed around 7 W will not transfer directly. Whether NVIDIA adds lower modes via software later is unknown. For power-starved designs, the Super, or a dedicated low-power accelerator, remains the realistic choice.
Thermal sizing: an illustrative calculation
Moving from 25 W to 40 W is a 60 percent increase in heat. For a sealed enclosure, that matters more than any TOPS figure. As an illustrative calculation, suppose ambient is 50 °C and your design keeps the module’s thermal transfer plate at or below 85 °C. That leaves a 35 °C budget.
At 25 W, the cooling path must achieve 35 / 25 = 1.4 °C/W from plate to ambient. At 40 W, it must achieve 35 / 40 ≈ 0.88 °C/W. That is roughly a 37 percent reduction in thermal resistance, which usually means a larger heatsink, a fan where you had none, or a redesigned enclosure. The 85 °C limit here is an assumption for illustration; use NVIDIA’s thermal design guide for Orin Nano 2 once it is published.
Power delivery changes too. A carrier board regulator and input supply sized for 25 W of module draw, plus peripherals, may brown out at 40 W transients. The “same socket” framing is about mechanical and electrical pin compatibility, not about whether your power tree has headroom. Treat 40 W mode as a hardware revision, even if the module drops in.
Should You Buy Orin Nano Super Now or Wait for Orin Nano 2?
This is the question most readers arrive with. The answer depends on ship date, workload shape, power budget and how much qualification risk you can absorb.

Figure 3: Module selection tree for 2027 edge AI products, covering Orin Nano Super, Orin Nano 2, Orin NX 16GB and Jetson T-series Thor.
Read the tree top to bottom. The first gate is precision and memory: if you need FP8 or FP4, or models larger than 8 GB allows, no Orin module fits. The second gate is schedule: if you must ship before mid-2027, Orin Nano 2 is not an option at all. Only then do workload shape and power budget pick between the Super and Orin Nano 2.
Decision matrix
| Your situation | Recommendation | Why |
|---|---|---|
| Shipping before Q3 2027 | Orin Nano Super now | Orin Nano 2 availability is 1H 2027; production qualification adds months |
| Prototyping today, production late 2027 or 2028 | Develop on Super, plan for Orin Nano 2 | Same form factor lets you design the carrier once, pending pinout confirmation |
| Classic CV at 7–15 W | Orin Nano Super | Reported Orin Nano 2 minimum is 15 W; CV rarely needs more |
| SLM up to 4B, latency-sensitive | Orin Nano 2 | Faster prefill and ~17% faster decode (author’s estimate) |
| VLM 7–8B INT4 with images | Orin Nano 2 at 25–40 W | Vision encoding and prefill are compute-bound, where the 2x lives |
| Needs 16 GB or a DLA today | Orin NX 16GB | Same GPU block, double memory, two DLAs, shipping now |
| Needs FP8, FP4 or transformer engine | Jetson T-series Thor | Ampere has no FP8 or FP4; Blackwell does |
| Hardware video encode on an entry module | Wait for Orin Nano 2, or use Orin NX | Super has no NVENC; Orin Nano 2 encoder is reported |
| Functional safety certification | Confirm core type before committing | Super is A78AE; Orin Nano 2 reported as A78 |
When buying the Super now is right
The Super is shipping, well documented, and supported by a mature JetPack 6 stack. NVIDIA extended the Orin family lifecycle through 2032 when it launched the Super. If your product ships in 2026 or early 2027, there is no decision to make.
It is also the right choice for workloads that the 2x does not touch much. A single-stream chatbot on a 3B model gains around 17 percent in decode on the author’s estimate. A classic detector at 15 W may see little difference if it already hits frame-rate targets. Paying later for a module you have never benchmarked, on a JetPack release that has not been announced, to gain 17 percent, is a poor trade for most teams.
When waiting for Orin Nano 2 is right
Wait if your product launches in late 2027 or later and at least one of these holds: latency-sensitive prefill, vision language models, hardware video encoding, PCIe Gen4 NVMe or camera throughput, or a battery budget where the 15 W efficiency point matters. In those cases, designing around the Super now and swapping modules later is sensible, provided you design your carrier and power tree for 40 W from the start.
Waiting also makes sense for teams whose current Super deployment is thermally or power constrained at 25 W. Getting Super-equivalent performance at 15 W opens headroom for sensors, radios or longer battery life, without changing the model.
Orin Nano 2 vs Orin NX: the uncomfortable comparison
The Orin NX 16GB in Super mode offers 157 TOPS (including DLAs), 16 GB of memory, 102 GB/s and a 10–40 W range, and it ships today in the same SO-DIMM family. For LLM decode, NX 16GB is slightly slower than Orin Nano 2 on bandwidth (102 versus a reported 120 GB/s). But it can hold a 7–8B model with a long context window and a perception stack without heroic memory management.
For many 2027 designs, the real choice is not Super versus Orin Nano 2. It is Orin Nano 2’s faster bandwidth and improved Tensor Cores versus NX’s doubled memory and DLAs. If your model plus context plus pipeline exceeds roughly 6 GB of working set, NX is the safer bet regardless of which Nano is faster. NVIDIA has not announced an “NX 2.”
When neither is right
If you need FP8 or FP4 inference, transformer engine support, or large-model memory, move up to Thor. ServeTheHome notes that Orin Nano 2 remains Ampere, so INT8 is, in practice, its lowest supported tensor precision (Ampere’s little-used INT4 tensor mode aside). That is the key differentiator from Blackwell-based modules. Our guide to Jetson T3000 vs T4000 vs T5000 Thor module selection covers that tier.
Migrating an Orin Nano Super Product to Orin Nano 2
“Same form factor” is the most reassuring phrase in the announcement and the most dangerous one. It says the module fits the connector. It does not say your product works unchanged.

Figure 4: A staged migration path from an Orin Nano Super product to Orin Nano 2, with a carrier respin loop if pinouts differ.
The flow front-loads the cheapest checks. Carrier and pinout verification come first, because a mismatch means a board respin and resets everything downstream. Software, model re-validation and thermal work follow, and the fleet rollout comes last.
Carrier board compatibility
NVIDIA and Antmicro both confirm the same compact SO-DIMM outline. JetsonHacks reports a PCIe configuration of one x4 plus two x1 lanes, versus one x4 plus three x1 on Orin Nano. If that is accurate, a carrier that routes a device to the fourth x1 lane would lose it. This is reported, not confirmed. Wait for NVIDIA’s product design guide and pinout spreadsheet before committing a board layout.
PCIe Gen4, if confirmed, also has signal integrity implications. A carrier board routed and validated for Gen3 may not meet Gen4 margins without changes to trace length, via design or redrivers. The link will usually fall back to Gen3 gracefully, but you should not assume Gen4 speeds on an existing carrier.
Software stack
NVIDIA has not announced which JetPack release will support Orin Nano 2. ServeTheHome reports that it runs the existing Jetson Orin software stack, but a new SoC will very likely need a newer board support package than the one your Super fleet runs today, but the exact version, Ubuntu base and CUDA major version are unconfirmed. If your Super fleet runs JetPack 6, budget for a full OS, driver and CUDA major-version migration as part of the switch, not just a module swap.
TensorRT engines are not portable across GPUs. Every model must be rebuilt on Orin Nano 2 hardware, and INT8 calibration caches should be regenerated and re-checked. New Tensor Core microarchitecture also means kernel selection may change, so accuracy regressions of a fraction of a percent are possible even with identical weights and calibration data.
Fleet operations
If you run a mixed fleet, your over-the-air update system needs per-SKU images. The module identifies itself through its board ID, so the update agent must select the right root filesystem, device tree and model engines. Mixing them up bricks devices or, worse, runs models built for the wrong GPU.
Trade-offs, Gotchas, and What Goes Wrong
The 2x will not show up in your chat benchmark. On the author’s arithmetic, single-stream decode on an 8B or 3B model improves about 17 percent from bandwidth, plus whatever runtime efficiency improves. Teams that benchmark only tokens per second will conclude NVIDIA overstated the claim. Measure time-to-first-token, vision-encoder latency and multi-stream throughput too.
Eight gigabytes is the real ceiling. Both modules share 8 GB between the CPU, GPU and every peripheral. A 7B model at INT4, a 4k-token KV cache, a vision encoder, the ROS 2 stack and the OS will not all fit comfortably. Plan the memory budget first and measure it with tegrastats on real hardware.
No FP8, no FP4. Some 2027 model releases may ship quantised checkpoints optimised for FP8 or FP4 on Blackwell. On Ampere, you will need INT8 or INT4 weight-only paths, which means waiting for community or NVIDIA quantisations, or doing them yourself.
40 W is not free. The top performance mode needs a bigger power supply, more cooling and possibly a new enclosure. Products that only qualify 15 W or 25 W modes will never see the 2x at all.
The DLA question is open. The original Orin Nano had its DLA disabled. NVIDIA has not said whether Orin Nano 2 includes one. Do not plan a DLA-offloaded pipeline until the datasheet confirms it.
Early silicon, early software. A new SoC means new errata, new power-management firmware and new driver paths. The first JetPack release for a new module is rarely the one you want in production. Budget at least one maintenance release of soak time.
Price is unknown. The Super’s $249 developer kit set a price anchor. NVIDIA has not disclosed pricing for Orin Nano 2 in either module or developer kit form. Do not build a BOM on an assumed price.
Benchmarks from others are not your benchmark. Unified memory contention, camera pipelines and thermal throttling make edge performance workload-specific. The same model on the same module can differ by 20 percent or more between a clean benchmark and a production image with sensors attached.
Practical Recommendations
For most teams, the right move is not binary. Develop and ship on the Orin Nano Super today, and design your next carrier revision so Orin Nano 2 can drop in later. That means a power tree and thermal path rated for 40 W, clean Gen4-capable PCIe routing, and an OTA system that already handles multiple module SKUs.
Invest in the measurement harness now. Split your benchmarks into time-to-first-token, decode tokens per second, vision-encoder latency and end-to-end pipeline latency, each under realistic memory contention. When Orin Nano 2 developer kits arrive, you will know within a day whether the 2x applies to you. Our on-device SLM inference benchmark on Jetson shows the kind of methodology that separates prefill and decode cleanly.
Checklist:
- Classify your workload as compute-bound (prefill, vision, CNN) or bandwidth-bound (single-stream decode) before estimating gains.
- Size power delivery and cooling for 40 W on any carrier you design from now on.
- Budget 8 GB explicitly: OS, CUDA context, perception, model weights, KV cache and headroom.
- Keep models in INT8 or INT4 paths; do not depend on FP8 or FP4.
- Wait for NVIDIA’s design guide before committing PCIe lane assignments.
- Plan a JetPack major-version migration and TensorRT engine rebuilds.
- Pick Orin NX 16GB if memory, not bandwidth, is the binding constraint.
- Pick Thor if you need Blackwell precisions or more than 8 GB of model memory.
Frequently Asked Questions
What is the Jetson Orin Nano 2 release date?
NVIDIA announced Jetson Orin Nano 2 on 25 August 2026 and says the module and developer kit are expected to be available in the first half of 2027. No exact date has been given. Production modules typically need extra time for carrier board partners to validate designs, so realistic volume shipments in finished products are more likely in the second half of 2027. Partners including Seeed Studio, Connect Tech, ADLINK, Advantech and Antmicro have announced support, which suggests carrier boards will be ready close to module availability.
Is Jetson Orin Nano 2 really twice as fast as Orin Nano Super?
For some workloads, probably yes; for others, no. NVIDIA claims 2x delivered inference, and ServeTheHome reports that this compares Orin Nano 2 at 40 W with the Super at 25 W. Compute-bound work such as LLM prefill, vision encoders and CNN detectors is where 2x is plausible. Single-stream LLM token generation is limited by memory bandwidth, which rises only about 17 percent (reported 102 to 120 GB/s). The author’s estimate is about 1.17x faster decode on an 8B INT4 model.
What are the Jetson Orin Nano 2 specs?
NVIDIA confirms 78 TOPS of AI compute, 8 GB of memory and an 8-core Arm CPU. Partner Antmicro confirms LPDDR5X memory and Cortex-A78 cores in the same SO-DIMM outline. Third-party sources, mainly ServeTheHome and JetsonHacks, report 120 GB/s of bandwidth, a 15–40 W power range, 1,536 CUDA cores, PCIe Gen4 and hardware video encoding. Those reported figures are not yet in an NVIDIA datasheet. Price, JetPack version and DLA presence have not been disclosed.
Should I choose Orin Nano 2 or Orin NX 16GB?
Choose based on memory versus bandwidth. Orin NX 16GB ships today with 16 GB of memory, 102 GB/s, two DLA engines and up to 40 W in Super mode. Orin Nano 2 has 8 GB but reportedly 120 GB/s and improved Tensor Cores. If your model, context window and perception stack need more than about 6 GB of working memory, NX is the safer choice. If your models fit in 8 GB and latency matters, Orin Nano 2 should be faster once available.
Does Jetson Orin Nano 2 support FP8 or FP4?
No. Orin Nano 2 uses an Ampere-architecture GPU, and Ampere Tensor Cores do not support FP8 or FP4. INT8 is the lowest precision in NVIDIA’s performance figures and mainstream toolchains, and INT4 weight-only quantisation runs through dequantisation in software kernels. If your roadmap depends on FP8 or FP4 models, or on a transformer engine, you need a Blackwell-based Jetson T-series Thor module instead. This is the clearest dividing line between Orin Nano 2 and the Thor tier.
Is Orin Nano 2 the best Jetson for edge LLMs in 2027?
It will likely be the best entry-level Jetson for small language models and compact vision language models, especially where time-to-first-token matters. It will not be the best overall. Its 8 GB memory limits model size and context length, and its Ampere GPU lacks the low-precision formats Thor supports. For 3–4B models in a tight power budget, Orin Nano 2 looks strong. For 8B models with long context, Orin NX 16GB or Thor are more comfortable fits.
Further Reading
- Hailo-10H vs Jetson Orin Nano for edge computer vision — when a dedicated accelerator beats a Jetson for camera pipelines.
- Edge LLM benchmark: Llama, Phi and Gemma on Jetson Orin — measured tokens per second on today’s Orin boards.
- Jetson T3000 vs T4000 vs T5000: Thor module selection — the Blackwell tier above Orin, with FP8 and FP4.
- On-device SLM inference benchmark on Jetson — methodology for separating prefill and decode.
References
- NVIDIA Newsroom, “NVIDIA Announces Jetson Orin Nano 2 Robotics Computer to Redefine Entry-Level Edge AI,” 25 August 2026. https://nvidianews.nvidia.com/news/nvidia-announces-jetson-orin-nano-2-robotics-computer-to-redefine-entry-level-edge-ai
- NVIDIA Technical Blog, “NVIDIA Jetson Orin Nano Developer Kit Gets a ‘Super’ Boost,” 17 December 2024. https://developer.nvidia.com/blog/nvidia-jetson-orin-nano-developer-kit-gets-a-super-boost/
- NVIDIA Technical Blog, “NVIDIA JetPack 6.2 Brings Super Mode to NVIDIA Jetson Orin Nano and Jetson Orin NX Modules.” https://developer.nvidia.com/blog/nvidia-jetpack-6-2-brings-super-mode-to-nvidia-jetson-orin-nano-and-jetson-orin-nx-modules
- ServeTheHome, “NVIDIA Announces Jetson Orin Nano 2: Entry-Level Edge Board Gets New Ampere Silicon,” 30 August 2026. https://www.servethehome.com/nvidia-announces-jetson-orin-nano-2-entry-level-edge-board-gets-new-ampere-silicon/
- JetsonHacks, “Jetson Orin Nano 2,” 4 September 2026. https://jetsonhacks.com/2026/09/04/jetson-orin-nano-2/
- Antmicro, “Entry-level edge AI with NVIDIA Jetson Nano 2 SoM,” 25 August 2026. https://antmicro.com/blog/2026/08/entry-level-edge-ai-with-nvidia-jetson-nano-2-som
By Riju — about
