Ternary Bonsai 2 27B: 1.76-Bit LLM Inference on Edge Hardware
A 27-billion-parameter model normally needs a workstation GPU or a server. Ternary Bonsai 2 ships the same model class in a 5.95 GB file, small enough to sit in the memory of a 16 GB laptop or an edge module with room left for a context cache. PrismML released it on 17 September 2026 under Apache 2.0, built from Qwen3.8 27B, with every weight reduced to one of three values: minus one, zero, or plus one.
The headline claim is that this keeps 98.2% of the FP16 baseline’s benchmark average at roughly one ninth of the size. That claim is PrismML’s own, and the details behind it matter more than the headline. The format needs a custom runtime, the accuracy loss is uneven, and “1.76 bits” is a storage figure, not a speed guarantee.
This post separates what is published from what is inferred, shows the storage arithmetic from first principles, and works through what a ~6 GB dense model does to memory bandwidth on Jetson, Apple Silicon, Raspberry Pi and NPU-class hardware.
What this covers: lineage and method, the exact bit budget, runtimes, bandwidth math for edge devices, the accuracy caveats, comparison against INT4, INT8, FP8 and BitNet b1.58, and a decision matrix.
Context and Background
Edge LLM deployment has been constrained by memory, not arithmetic. Decoding one token from a dense model reads every weight once, so throughput is bounded by memory bandwidth divided by model bytes. Quantization is therefore the main lever: fewer bytes per weight means more tokens per second and a model that fits at all. Our comparison of INT4, INT8 and FP8 on edge NPUs covers the mainstream formats, where 4-bit is the usual floor for reasoning-heavy workloads.
Below 4 bits, conventional post-training methods degrade quickly. PrismML’s own comparison, reported on the Hugging Face model card, puts a conventional IQ2_XXS build at 72.59 on its 14-benchmark thinking suite, or 84.1% of FP16, while a 4-bit UD-Q4_K_XL build at 17.6 GB scores 85.18. Two-bit builds often survive knowledge tests while failing multi-step reasoning, which is why casual spot checks can mislead.
Ternary weights have a research lineage. Microsoft Research’s “The Era of 1-bit LLMs” (arXiv 2402.17764, February 2024) proposed BitNet b1.58, where every weight is in {-1, 0, 1} and the model is trained that way from the start. The later BitNet b1.58 2B4T release (arXiv 2504.12285) trained a 2-billion-parameter model on 4 trillion tokens. The catch: native ternary training requires pre-training a new model, which few teams can afford at 27B scale.
PrismML’s bet is different. It converts an existing strong model after the fact. Its first Ternary Bonsai family arrived on 16 April 2026 in 8B, 4B and 1.7B sizes, with the 8B at 1.75 GB and about 9 times smaller than a 16-bit equivalent, according to PrismML’s announcement. That announcement did not name the base model or disclose the training method. The 2 series is the first at 27B, and it is unusually explicit about its lineage.
For readers tracking the small-model side of the story, our on-device SLM benchmark on Jetson shows what 1B to 8B models do on the same hardware. Ternary Bonsai 2 27B asks whether a much larger model can now live where only small models did.
Ternary Bonsai 2 27B: What It Is and How It Is Built
Direct answer: Ternary Bonsai 2 27B is a post-training ternary conversion of Alibaba’s Qwen3.8 27B by PrismML. Weights are {-1, 0, +1} with one FP16 scale per 128 weights, rotated with a Hadamard transform before rounding. It ships at 1.75 bits per weight (5.95 GB) and needs PrismML’s llama.cpp fork or a bundled MLX loader.
The published specification
The following facts come from PrismML’s model documentation, its announcement and the Hugging Face GGUF card.
| Attribute | Published value |
|---|---|
| Base model | Qwen3.8 27B, hybrid attention (about 75% linear, 25% full), SwiGLU MLP, RoPE, RMSNorm |
| Parameters | 27.36B total: 24.35B language, 2.54B embedding and LM head, 0.47B vision tower |
| Weight format | Ternary g128: {-1, 0, +1} with FP16 group-wise scale, blockwise Hadamard rotation |
| Effective bits | 1.72 bpw true ternary; 1.75 to 1.76 bpw as shipped in PTQ1_0 |
| Context | 262,144 tokens; text and image in, text out |
| License | Apache 2.0 |
| Release | 17 September 2026 per PrismML; several outlets date coverage 18 September |
| Sizes | PTQ1_0 5.95 GB; PQ2_0 7.21 GB; MLX 2-bit 8.49 GB with vision tower; FP16 reference 53.8 GB |
Two details deserve attention. First, the model is dense, not Mixture-of-Experts, so every generated token touches the whole language model. Second, roughly 75% of layers use linear attention, which keeps the recurrent state small and the KV cache far smaller than in a fully quadratic 27B model. That is a large part of why a 262K context is plausible on modest memory.

Figure 1: The Ternary Bonsai 2 27B pipeline. Rotation happens offline; the activation-side transform happens at every inference step, which is why stock runtimes cannot load the files.
Figure 1 shows the flow. The FP16 checkpoint is rotated in 1024-element blocks, rounded to three levels, and stored with one FP16 scale per 128 weights. The same ternary values are then packed three ways for different hardware, and each packing has a matching runtime requirement.
Post-training conversion, not native ternary training
Is this quantization-aware training or post-training quantization? PrismML’s announcement page states neither; when I asked it directly, the page said the method was “not stated” and pointed to a whitepaper. Third-party summaries of that whitepaper describe post-training ternarization of the existing Qwen3.8 27B, and a community issue quoting the release describes weights produced “from the FP16 safetensors (not from an existing quant)” with imatrix calibration. I could not open the whitepaper itself, so treat “post-training” as well supported but secondhand.
The distinction matters. BitNet b1.58 trains ternary from initialization, so the network learns to live with three levels. A post-training conversion has to preserve behavior the base model already learned, including tool-calling and reasoning tuning. It avoids the multi-trillion-token pre-training bill and can be repeated whenever a better base appears. It also has a ceiling: the numbers below show where retention breaks down.
Why the Hadamard rotation is the core trick
Ternary rounding fails when a few outlier weights dominate a group’s scale. If one weight is ten times larger than its neighbors, the shared scale is set by that outlier, and most other weights round to zero.
A Hadamard transform is an orthogonal rotation built from plus and minus one signs. Applied to a block, it spreads the energy of an outlier across all coordinates, producing a flatter distribution that three levels can represent far better. A community reading of the code describes a QuaRot/SpinQuant-style system: weights are rotated offline before ternary rounding, and activations go through a matching fast Walsh-Hadamard transform at inference.
Because rotation is orthogonal, applying it to both weights and activations leaves the mathematical function unchanged before quantization. The cost is a runtime contract. Every projection needs an activation-side transform, and the MLX card warns that ordinary loaders “skip the activation transform and the inverse embedding lookup, so they return wrong output rather than an error.” New tensor type codes in the GGUF files exist so that stock llama.cpp refuses to load them instead of producing silent garbage.
The Bit Budget: Where 1.76 Bits Per Weight Comes From
“1.76 bits per weight” is easy to misread as a 1.76-bit number system. It is an average storage cost that includes the scales, the packing overhead and a small set of higher-precision tensors. The arithmetic can be reproduced from the published figures.
From log2(3) to 1.75
A value with three equally likely states carries log2(3), about 1.585 bits of information. That is the floor for any ternary weight. On top of it, each group of 128 weights stores one FP16 scale: 16 bits divided by 128 weights adds 0.125 bits per weight. The sum, about 1.71 bits, is the ideal ternary rate, and PrismML rounds its “true ternary” figure to 1.72 once the higher-precision tensors are counted.
Real storage cannot use fractional bits, so trits are packed in base 3. Five trits fit in one byte because 3^5 is 243, which is less than 256. That gives 1.6 bits per trit, close to the 1.585 floor. A community teardown of the GGUF format describes PTQ1_0 as packing 120 of every 128 weights at five trits per byte and the remaining 8 at four trits per byte, plus one FP16 scale.

Figure 2: One PTQ1_0 group. Trit payload plus scale is 28 bytes per 128 weights, which reproduces the published 1.75 bpw.
Working through the layout in Figure 2 (my arithmetic from that community description, not a PrismML specification):
- 120 weights at 5 trits per byte: 24 bytes
- 8 weights at 4 trits per byte: 2 bytes
- 1 FP16 scale: 2 bytes
- Total: 28 bytes, or 224 bits, for 128 weights
- 224 divided by 128 equals exactly 1.75 bits per weight
That matches the 1.75 bpw PrismML lists for PTQ1_0. The 1.76 headline is the same layout with a little extra overhead from tensors that are not ternary. PQ2_0 is simpler: 2 bits per weight for 128 weights is 32 bytes, plus a 2-byte scale, giving 34 bytes or 2.125 bpw. The published 2.13 bpw matches. The trade is 21% more bytes for cheaper unpacking, because a 2-bit field needs only shifts and masks rather than base-3 division.
Reproducing the 5.95 GB file size
Now the total. The 2.54B embedding and LM head parameters appear to be ternarized along with the 24.35B language parameters, because that is the only reading that makes the published FP16 size work: (24.35B + 2.54B) is 26.89B parameters, and at 2 bytes each that is 53.78 GB, matching the published 53.8 GB F16 reference. The 0.47B vision tower ships separately.
At 1.75 bits per weight, the same 26.89B parameters occupy:
26.89 x 10^9 x 1.75 / 8 = 5.88 x 10^9 bytes, or about 5.88 GB
PrismML also reports 26.2M parameters (0.0976%) held in higher precision for recurrent state and normalization. At 2 bytes each that is about 0.05 GB, so the file lands near 5.93 GB, which matches the 5.93 GB figure in press coverage of the whitepaper and sits close to the 5.95 GB on the model card. For PQ2_0 the same calculation gives 26.89B x 2.125 / 8 = 7.14 GB, plus 0.05 GB, or about 7.19 GB against a published 7.21 GB. The small residuals are file metadata and rounding. This is a reconstruction; PrismML does not publish the per-tensor table in the pages I could read.
The compression ratio follows directly: 53.8 divided by 5.93 is about 9.1 times.
How the ratio compares with other formats
Illustrative sizes for the same 26.89B language-plus-embedding parameters, computed from bits per weight (real files vary by tensor-level choices):
| Format | Approx. bits per weight | Approx. size |
|---|---|---|
| FP16 or BF16 | 16 | 53.8 GB (published) |
| FP8 or INT8 with scales | about 8.1 | about 27 GB |
| INT4, group 128, FP16 scale | about 4.125 | about 13.9 GB |
| UD-Q4_K_XL (measured) | mixed | 17.6 GB (published) |
| PQ2_0 | 2.125 | 7.2 GB |
| PTQ1_0 | 1.75 | 5.9 GB |
The measured 4-bit build is larger than a naive 4-bit estimate because dynamic 4-bit schemes keep sensitive tensors at higher precision. The ternary file is about 3 times smaller than that 4-bit build, and about 4.5 times smaller than an FP8 checkpoint.
Runtimes and Kernels: What Actually Runs It
The format is only useful if something can execute it. This is the biggest practical constraint on Ternary Bonsai 2 today, and it is stated bluntly on the model card: stock llama.cpp cannot load these files, and PrismML’s fork is required.
The PrismML llama.cpp fork
GGUF users need the PrismML llama.cpp fork at build prism-b10658 or newer. The card lists CUDA, Metal and CPU paths with custom ternary hybrid-attention kernels. The fork’s README also advertises other backends, including HIP, Vulkan and Snapdragon, but whether the 27B ternary types are implemented on each is not confirmed on the pages I retrieved, so verify per backend before planning around one.
Two packings are available because two hardware regimes exist. PTQ1_0 moves fewer bytes and is reported faster on Ada-generation GPUs and the L4, where memory bandwidth is the limit. PQ2_0 costs more bytes but less unpacking arithmetic and is reported faster on H100, A100 and Blackwell, where compute headroom is abundant and bandwidth is plentiful. The fork’s README prefers PQ2_0 across Metal, CUDA, HIP and CPU for the Bonsai line generally.
MLX on Apple Silicon
Apple users get an MLX pack that runs on stock MLX libraries but needs the loader bundled in the repository’s runtime/ folder. Its trick is clever: ternary levels {-s, 0, +s} are reproduced exactly by setting the MLX quantization scale to s and the bias to -s, so 2-bit codes {0, 1, 2} decode to -s, 0, +s. The price is 2.25 bpw, since MLX stores a scale and a bias per 128-weight group. That is why the MLX pack is larger on disk (about 8.49 GB with the vision tower) than the GGUF files, despite carrying identical ternary values.
Where support is missing
I found no evidence of support in TensorRT-LLM, vLLM or ONNX Runtime, and independent porting efforts describe the work as substantial. One issue on a Rust inference project concluded that the architecture is supported but the quantization “needs a bespoke rotation-quant reimplementation,” and a community MLX patch adds the PTQ1_0 and PQ2_0 codecs with run-time rotation. For Jetson teams, this matters: our TensorRT-LLM versus llama.cpp comparison on Jetson explains why llama.cpp is the more flexible path for exotic formats, and Ternary Bonsai 2 fits that pattern. A TensorRT-LLM engine would need a custom plugin for the rotation and the trit unpacking.
There is also a hosted route. OpenRouter lists the model with a single provider at $0.075 per million input tokens and $0.50 per million output tokens, a 262,144-token context and a 32,768-token output cap, at an observed 18 tokens per second at the time I checked. Hosted access says little about edge behavior, but it is a cheap way to test quality before committing hardware.
Edge Hardware: Bandwidth, Memory and Realistic Throughput
For a dense model, the decode ceiling is simple: tokens per second cannot exceed memory bandwidth divided by the bytes read per token. Ternary Bonsai 2 reads about 5.9 GB per token from the language model, so bandwidth sets the ceiling and everything else (Hadamard transform, unpacking, attention state, scheduling) pulls real numbers below it.

Figure 3: One projection in a ternary Bonsai layer. The activation is rotated, then the kernel unpacks trits, adds or subtracts activations, skips zeros and applies the group scale.
What the published measurements imply
PrismML and press coverage report decode throughput on desktop and laptop hardware only. Reported figures include 142.5 to 143 tokens per second on an RTX 5090 (PQ2_0, with a lower 129.9 tokens per second measured in an earlier round), 96.7 on an RTX 4090, 32.1 on a 72 W L4, and 46.8 to 47.0 on an Apple M5 Max, with the M5 Pro at 27.7 to 28.7. Energy is reported at 0.714 mWh per token on an RTX 4090, and 0.582 mWh per token on the 5090 in one secondary write-up.
Multiplying throughput by file size gives the effective bytes streamed per second, a useful sanity check:
- RTX 5090: 142.5 tok/s x 5.93 GB is about 845 GB/s
- M5 Max: 46.8 tok/s x 5.93 GB is about 278 GB/s
- M5 Pro: 27.7 tok/s x 5.93 GB is about 164 GB/s
- L4: 32.1 tok/s x 5.93 GB is about 190 GB/s
The implied figures are effective numbers, not spec-sheet peaks, and I have not verified each device’s published memory bandwidth. They do show the model is being fed at a healthy fraction of what these devices can move, which is consistent with a bandwidth-bound workload. Energy converts easily too: 0.714 mWh is 2.57 joules per token, so at 96.7 tokens per second the RTX 4090 draws roughly 250 W in that measurement. That is a board-level figure for a desktop card, not an edge power budget.
The ceiling on Jetson, Pi and other edge parts
PrismML publishes no Jetson, Raspberry Pi, iPhone or NPU numbers for the 27B model (the first-generation Ternary Bonsai 8B did report 27 tokens per second on an iPhone 17 Pro Max, but that is a different, smaller model). The table below is therefore my ceiling calculation, using module memory bandwidth from a public Jetson comparison and PTQ1_0 at 5.93 GB per token. It is illustrative, not a benchmark.
| Device | Memory | Peak bandwidth | Ceiling, ternary 5.93 GB | Ceiling, 4-bit 17.6 GB | Ceiling, FP8 about 27 GB |
|---|---|---|---|---|---|
| Raspberry Pi 5, 16 GB | 16 GB LPDDR4X-4267 | about 17 GB/s (my calc, 4267 MT/s x 4 B) | about 2.9 tok/s | does not fit comfortably | no |
| Jetson Orin Nano 8 GB | 8 GB LPDDR5 | 102 GB/s | about 17 tok/s but only 2 GB headroom | no | no |
| Jetson Orin NX 16 GB | 16 GB LPDDR5 | 102.4 GB/s | about 17 tok/s | no, 17.6 GB exceeds 16 GB | no |
| Jetson AGX Orin 64 GB | 64 GB LPDDR5 | 204.8 GB/s | about 34 tok/s | about 11.6 tok/s | about 7.5 tok/s |
| Jetson AGX Thor 128 GB | 128 GB LPDDR5X | 273 GB/s | about 46 tok/s | about 15.5 tok/s | about 10 tok/s |
Three caveats keep this honest. Theoretical bandwidth is rarely achieved: an NVIDIA developer-forum thread on an AGX Orin 32GB reported STREAM host-copy near 55 GB/s and device-to-device near 152 GB/s against the 204.8 GB/s spec, and NVIDIA called the spec figure theoretical. Applying the 152 GB/s measurement to the AGX Orin row would put the ternary ceiling nearer 26 tokens per second than 34. The Orin Nano’s 8 GB cannot hold the 5.93 GB weights plus operating system, runtime and cache comfortably, and PrismML lists a 24 GB GPU or a 16 GB laptop as the practical minimum. And nothing above accounts for the fork’s CUDA kernels: the published notes favor Ada, Hopper, Ampere-class A100 and Blackwell, and whether the fork’s CUDA build runs well on Orin’s Ampere or Thor’s Blackwell generation is unverified.
The ratio is the durable insight. A ternary file cuts the bytes per token by about 3 times against a measured 4-bit build and about 4.5 times against FP8, so the bandwidth ceiling rises by the same factor on any device. Whether that converts to real speed depends on kernel quality, which is why our Jetson SLM benchmark methodology (measure tokens per second at fixed context, warm cache, fixed power mode) applies here unchanged.
Memory beyond the weights
The weights are only one line of the memory budget. Linear-attention layers hold a fixed-size recurrent state rather than a growing KV cache, which is why 75% linear attention matters for long contexts. The full-attention quarter still grows with sequence length. A third-party deep dive estimates about 9.4 GB peak at the full 262K context with an FP16 KV cache; treat that as an outside estimate and measure your own build. The vision tower adds 0.63 GB for the Q8_0 mmproj file in GGUF (0.92 GB in MLX), loaded only when images are used.
NPUs and fixed-function accelerators
Ternary weights are attractive for hardware in principle, because multiplication by minus one, zero or one becomes add, subtract or skip. In practice, current NPUs are built around INT8 and INT4 datapaths, and mobile NPU toolchains do not accept custom ternary types with a Hadamard activation transform. PrismML’s shipped kernels target GPUs, Apple Metal and CPUs. Expect Ternary Bonsai 2 to run on the GPU or CPU of an edge SoC, not its NPU, until vendors add ternary or 2-bit operators. The edge NPU quantization comparison explains what those NPUs natively support today.
Accuracy: What the 98.2% Actually Hides
The retention figure is real but reads differently once you split it. PrismML reports two suites: an average of 83.9 against 85.4 for FP16 across 20 benchmarks, and 84.78 against 86.32 across 14 thinking-mode benchmarks. Both give 98.2%. The Hugging Face card is the source for the second suite and the other page for the first.
Category breakdown
On the 20-benchmark suite, per PrismML’s announcement:
| Category | Ternary Bonsai 2 27B | Qwen3.8 27B FP16 | Retention |
|---|---|---|---|
| Agentic and tool use | 77.57 | 79.74 | 97.3% |
| Coding | 81.58 | 82.17 | 99.3% |
| Instruction following | 82.66 | 81.25 | 101.7% |
| Knowledge and reasoning | 83.95 | 86.66 | 96.9% |
| Math | 96.57 | 97.06 | 99.5% |
| Vision | 78.59 | 81.64 | 96.3% |
Math and coding are essentially preserved; knowledge and vision give up three to four points. Instruction following slightly exceeds the baseline, which is within the range of benchmark noise and should not be read as a real improvement.
The long-horizon gap
The aggregate hides the weakest area. Press coverage of the whitepaper reports Terminal-Bench 2.1 at 52.8 versus 69.7 for FP16 and SWE-bench at 60.8 versus 80.6. Those are retentions of about 76% and 75%, and they sit far below the 98% headline. Long-horizon agentic tasks compound small per-step errors: if each step is a little less reliable, the chance of completing fifty steps falls sharply. For a model that is marketed with tool-calling and agentic features, this is the number that matters for robotics and industrial automation.
Sources of doubt
- Vendor-reported. PrismML’s results are its own, and I found no independent replication of the benchmark averages at the time of writing.
- Suite mismatch. The two suites give different absolute numbers and even different category values (coding is 81.58 on one and 89.42 on the other), so compare within a suite only.
- Reasoning effort. The card notes that “low” reasoning effort is not supported. OpenRouter lists the model as thinking by default at “xhigh” effort, so a comparison against a non-thinking baseline is unfair to both sides.
- Sampling. The card recommends temperature 1.0, top-p 0.95, top-k 20 and min-p 0.05 for thinking mode, and temperature 0.7 with top-p 0.80 for instruct mode. Changing them changes scores.
- Vision. The GGUF card reports 66.19 against 71.36 for FP16 on its vision set, a bigger relative drop than the announcement’s category table suggests.
The honest summary is that ternary conversion preserved short-form reasoning, math and code very well, and lost more on knowledge recall, vision and above all sustained multi-step agent work.
Ternary vs INT4, INT8, FP8 and BitNet b1.58
The useful comparison is not “which is best” but “which constraint binds.” Each format trades bytes, accuracy risk, tooling maturity and hardware fit differently.
Against BitNet b1.58
BitNet b1.58 and Ternary Bonsai 2 share the same value set and the same absmean-style idea of a per-tensor or per-group scale, but they differ in almost everything else. BitNet is trained ternary from scratch, so the network’s optimization landscape accounts for the constraint. Bonsai starts from a finished FP16 model and rounds it, using rotation to make that survivable. The publicly documented BitNet b1.58 2B4T model is 2 billion parameters trained on 4 trillion tokens; Bonsai 2 is 27 billion parameters and inherits its training from Qwen3.8. Native ternary models are proven at smaller scale; post-training ternary at 27B is the newer and less-tested claim.
The runtime story also differs. Microsoft’s BitNet models run on a dedicated CPU and GPU inference framework, while Bonsai runs on a llama.cpp fork. Neither is available in stock mainstream runtimes. On the format side, BitNet-style quantization typically pairs ternary weights with 8-bit activations, whereas Bonsai’s activations stay in higher precision and are rotated. I did not find PrismML publishing head-to-head numbers against BitNet, so any claim that one is more accurate than the other at equal size is unsupported.
Against INT4, INT8 and FP8
INT8 and FP8 are the safe defaults. They are close to lossless for most models, well supported by TensorRT-LLM, ONNX Runtime and NPU toolchains, and give a 2 times size reduction. INT4 with group scales and outlier handling is the mainstream sweet spot: 4 times smaller, usually within a point or two of FP16, and supported nearly everywhere. Our FP8 vs INT8 vs INT4 LLM quantization benchmark and the edge NPU quantization comparison cover these trade-offs in detail.
In PrismML’s own comparison, the 4-bit UD-Q4_K_XL scores 85.18 at 17.6 GB against 84.78 for ternary at about 5.9 GB. That is 0.4 points for a 3 times smaller file. The interpretation depends on what the extra 11.7 GB costs you. On a 64 GB module it costs nothing that matters. On a 16 GB module it is the difference between fitting and not fitting.

Figure 4: Deployment decision flow. Memory headroom and runtime freedom come first; bandwidth then sets the throughput class.
Decision matrix
| Scenario | Best fit | Why | Watch out for |
|---|---|---|---|
| 16 GB laptop or mini PC, want a 27B-class model | Ternary Bonsai 2 27B | Only 27B option that leaves KV headroom | Custom fork; verify agent tasks |
| 64 GB Jetson AGX Orin or Thor, production stack | INT4 or FP8 on stock runtime | Memory is free, tooling is mature | Slower decode than ternary by size ratio |
| Same device, need max tokens per second | Ternary PTQ1_0 if kernels perform | 3 to 4.5 times fewer bytes per token | Unpacking and rotation overhead |
| Long-horizon coding or tool agents | FP8 or INT4 27B, or smaller FP16 model | Reported 75% retention on SWE-bench and Terminal-Bench | Ternary gap widens with step count |
| Raspberry Pi 5 16 GB | Smaller model | Ternary ceiling near 3 tok/s | 27B is a demo, not a product |
| NPU-only target | INT8 or INT4 model | NPUs lack ternary operators | Kernel support absent |
| Safety-critical or regulated | Standard formats with validated toolchain | Auditable, common runtimes | Custom forks complicate certification |
Trade-offs, Gotchas, and What Goes Wrong
Silent-failure risk. The most dangerous failure is loading a rotated model in a loader that ignores the rotation. PrismML deliberately uses new tensor type codes and a bundled MLX loader to force an error, but derivative tools that patch support incompletely could produce fluent nonsense. Pin the runtime build, hash the model files and add a smoke test that checks a known prompt’s output in CI.
Fork dependency. Production systems built on prism-b10658 inherit the fork’s release cadence. The fork’s own README warns against a stale prism-v6 branch and against mixing its libraries with stock llama.cpp builds. Budget for tracking upstream drift, especially if you also serve other GGUF models from the same host.
Bandwidth is not the only bottleneck. The Hadamard activation transform runs on every projection. A third-party summary of the whitepaper quotes it as “one of the larger non-matrix-multiply costs” at batch size one. At larger batch sizes, or on devices with weak compute, unpacking and rotation can erase part of the bandwidth win. That is the logic behind shipping two packings.
Batching and prefill change the picture. Ternary compression helps decode most, because decode is bandwidth-bound. Prefill and large-batch serving are compute-bound, where fewer bytes matter less. Reported prefill rates (4,121 tokens per second on the 5090 and 397 on an M5 Pro in one write-up) are fast but not proportional to the decode gains.
Quality cliffs on agents. A 98% average retention does not mean 98% of tasks succeed. The reported 75% retention on long-horizon benchmarks means a ternary agent may need more retries or human checks than the FP16 original. Evaluate on your own task distribution, not on aggregate scores.
Fine-tuning is unclear. I found no documentation on fine-tuning the ternary weights. LoRA adapters on top of a ternary base are technically plausible but unvalidated, and merging adapters back into ternary weights would re-quantize and could lose the adaptation. Plan on prompt engineering and retrieval rather than adaptation until PrismML documents a path.
Thermals and power. Edge modules run in constrained power modes. The reported energy figures come from desktop GPUs, so do not extrapolate joules per token to a 15 W or 25 W module. Measure under your production power mode.
Licensing and provenance. The Ternary Bonsai 2 weights are Apache 2.0. Qwen3.8’s own license terms determine what a derived model can do, and PrismML’s licensing statement covers its release. Have legal review the chain before shipping in a product.
Practical Recommendations
Treat Ternary Bonsai 2 27B as a memory-capacity tool first and a speed tool second. It earns its place when a 27B-class model would not otherwise fit, or when the freed memory buys a long context or a second concurrent model. It is a weaker choice where a mature INT4 or FP8 stack already fits comfortably.
Start with a proof of value on your own workload. Pull the PTQ1_0 GGUF, build the fork at the pinned tag, and run the same prompts against a 4-bit build of Qwen3.8 27B if one fits, otherwise against a hosted endpoint. Measure quality on your task set, then decode tokens per second at your real context length, then energy at your real power mode.
Use the flow in Figure 4. If you cannot free 8 GB or more, pick a smaller model. If you cannot run a custom fork, pick INT4 or FP8. If both pass, ternary is viable and bandwidth sets the speed class.
- Confirm at least 8 GB of free memory for weights, runtime and cache; 16 GB is safer, matching PrismML’s guidance.
- Pin the fork tag
prism-b10658or newer and record the commit hash. - Test both PTQ1_0 and PQ2_0; the faster one depends on your GPU generation.
- Use PrismML’s recommended sampling parameters for thinking and instruct modes.
- Do not select “low” reasoning effort; it is unsupported.
- Run a long-horizon agent test before trusting it with multi-step tool use.
- Measure real bandwidth with a STREAM-style test rather than trusting the spec sheet.
- Keep a fallback path to a stock-runtime model.
- Re-check for independent benchmarks and upstream llama.cpp support before you lock in a design.
Frequently Asked Questions
What is Ternary Bonsai 2 27B?
Ternary Bonsai 2 27B is PrismML’s ternary-weight version of Qwen3.8 27B, released on 17 September 2026 under Apache 2.0. Each weight is minus one, zero or plus one, with an FP16 scale per 128 weights and a Hadamard rotation applied before rounding. It ships as a 5.95 GB GGUF file with a 262K-token context, a text-and-image input, and requires PrismML’s llama.cpp fork.
How can it be 1.76 bits per weight if ternary is 1.585 bits?
The 1.585 figure is log2(3), the information in one three-state value. Storage adds an FP16 scale per 128 weights (0.125 bits per weight) and packing overhead. In the PTQ1_0 layout each 128-weight group takes about 28 bytes, or 224 bits, which is exactly 1.75 bits per weight. The extra hundredth up to 1.76 comes from small higher-precision tensors that are not ternarized.
Does it really keep 98.2% of the original accuracy?
PrismML reports 98.2% retention on two different suites: 83.9 against 85.4 on 20 benchmarks, and 84.78 against 86.32 on 14 thinking benchmarks. Math and coding are close to lossless, but long-horizon agent tasks retain only about 75% (SWE-bench 60.8 versus 80.6). The figures are vendor-reported and, as of writing, I found no independent replication, so validate on your own tasks.
Can I run Ternary Bonsai 2 27B on a Jetson or Raspberry Pi?
PrismML has published no Jetson or Raspberry Pi results for this model. Arithmetic suggests a bandwidth ceiling of roughly 17 tokens per second on a 102 GB/s Orin NX 16 GB, about 34 on an AGX Orin 64 GB, and around 3 on a Pi 5 16 GB, before real-world losses. Kernel support for those GPU generations in the fork is unverified. Treat those numbers as ceilings, not results.
How is it different from BitNet b1.58?
BitNet b1.58 trains a model from scratch with ternary weights, so the network adapts to the constraint; the public 2B4T release has 2 billion parameters trained on 4 trillion tokens. Ternary Bonsai 2 converts an existing 27B model after training, relying on Hadamard rotation to limit accuracy loss. Bonsai reaches a larger scale cheaply, while BitNet-style native training has stronger evidence at small sizes. Both need custom runtimes.
Should I use ternary or 4-bit for a 27B model?
Use 4-bit when memory is not binding and you want mature tooling, since its reported score is about 0.4 points higher (85.18 versus 84.78) at 17.6 GB. Use ternary when the 4-bit file does not fit, when you need more room for context, or when you want the lowest bytes per token. Avoid ternary for long-horizon agent work until you have tested it on your own tasks.
Further Reading
- INT4 vs INT8 vs FP8 on edge NPUs: the mainstream format trade-offs that ternary is measured against.
- On-device SLM inference benchmark on Jetson: methodology and results for small models on the same hardware.
- TensorRT-LLM vs llama.cpp on Jetson: why a llama.cpp fork is the natural home for exotic formats.
- Hailo-10H vs Jetson Orin Nano for edge vision: accelerator choices for the surrounding edge system.
- PrismML Ternary Bonsai 2 27B documentation and the Hugging Face GGUF model card: primary sources for every figure attributed to PrismML.
- The Era of 1-bit LLMs, arXiv 2402.17764: the original BitNet b1.58 paper.
By Riju — about
