Tencent Hy4 Explained: 770B MoE, Sherry Quantization and Open Weights
A 770-billion-parameter model that you can actually download, under Apache 2.0, with a one-million-token context window, sounds like a hosting problem rather than a model release. The weights in BF16 occupy roughly 1.5 TB, which is more memory than any single server holds. Tencent Hy4 preview, announced on 28 August 2026, is interesting less for the raw parameter count than for what surrounds it: a sparse design that touches only 49 billion parameters per token, an FP8 checkpoint on day one, and a follow-up claim that a 1.25-bit-class quantization called Sherry compresses the checkpoint to about 214 GB.
That last number is the one that changes who can experiment with a frontier-scale open model. It moves the question from “which cluster” to “which pooled set of workstations.” This post separates what Tencent has published from what third parties have reported, explains the architecture and the quantization mechanism from first principles, and works through hardware sizing with arithmetic that is clearly labelled as illustrative. One correction up front: the release date circulating in some weekly roundups is early September, but the model card and several news reports date the launch to 28 August.
What this covers: the Hy4 lineage and design, the 78-layer MoE architecture, how Sherry packs weights into 1.25 bits, what the reported 214 GB number does and does not mean, benchmarks with caveats, deployment and pricing, failure modes, and a comparison matrix against GLM-5.3 and Kimi K3.
Context and Background
Hy4 is the latest generation in Tencent’s Hunyuan-derived model line, which the company now brands under the shorter “Hy” name. Reports compare it against its predecessor Hy3 and against peers such as GLM-5.3, Kimi K3, DeepSeek V4 Pro and Qwen3.8 Max. The “preview” label matters: Tencent’s own materials describe this as an early version, and we should read every benchmark in that light. For a sense of how we have treated comparable releases, see our analysis of GLM-5.2 as an open-weight LLM, which sets the baseline for what open weights at this tier now deliver.
The strategic backdrop is a shift in how frontier labs ship. Two years ago, “open weights” at the top end usually meant a model you could technically download but not practically serve. The current pattern pairs the release with an inference story: sparse activation so per-token compute is modest, a low-precision checkpoint so memory is tractable, and a serving recipe for vLLM or SGLang so the deployment path is not guesswork. Hy4 follows that pattern, and the quantization thread is the novel part.
Quantization itself is not new, and readers who want the general trade-offs should start with our FP8 vs INT8 vs INT4 LLM quantization benchmark. What is new here is the extreme end of the curve. Ternary and 1-bit research has mostly been demonstrated on models up to a few billion parameters; our piece on ternary 1-bit LLMs for edge inference covers that small-model regime. Applying a related idea to a 770B MoE is a different bet, because the weights are dominated by expert feed-forward matrices that are individually less sensitive than attention layers.
The primary sources for this post are the Hy4-preview model card on Hugging Face for architecture and deployment, the Sherry paper on arXiv for the quantization method, and trade-press coverage such as TechNode and GIGAZINE for launch details. Where I rely on a third-party summary rather than a document I could read directly, I say so.
What Hy4 Is: A Sparse 770B Model With a Small Active Footprint
Tencent Hy4 preview is a sparse Mixture-of-Experts transformer with 770 billion total parameters and 49 billion activated per token, a 1M-token context window, and an Apache 2.0 license. Each of 77 MoE layers holds 256 routed experts plus one shared expert, and each token uses eight routed experts plus the shared one. A 10B-parameter multi-token-prediction layer supports speculative decoding.

Figure 1: Hy4 preview block structure as described on the model card: one dense layer, 77 MoE layers, top-8 routing plus a shared expert, and a separate MTP head.
The diagram shows the data path for a single token. After embedding, the first layer uses a standard dense feed-forward network; layers two through 78 replace that feed-forward block with a routed expert bank. Attention in every layer uses what the model card calls Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache, which reuses sparse attention indices across layers so the cost of deciding which past tokens to attend to is paid less often. The multi-token-prediction (MTP) layer sits outside the backbone and drafts several candidate tokens per step for the verifier to accept or reject.
Why the active-parameter ratio is the number that matters
The ratio of active to total parameters is 49 over 770, about 6.4 percent. That ratio governs two different costs in opposite directions. Compute per token scales with the 49B active parameters, so the floating-point work per generated token is comparable to a dense model of roughly that size. Memory, however, scales with the full 770B, because any expert might be selected on any token and so every expert must be resident or quickly reachable.
This is the core tension of every large MoE. You buy cheap per-token arithmetic and pay for it in capacity. At batch size one, decode is memory-bandwidth-bound: each step must read the weights of whichever experts fire. At large batch sizes across many users, nearly all experts get touched every step, so the bandwidth advantage of sparsity erodes and the system behaves more like it is reading the whole model repeatedly. Serving architecture, not the model alone, determines whether the 6.4 percent figure turns into real savings, a theme we explore in our guide to expert-parallel MoE inference serving.
The reported dimensions
Beyond the headline counts, secondary coverage that reproduces the model card lists a hidden size of 6,144, 64 attention heads, a vocabulary of 120,832 tokens, and query and key-value compression dimensions of 2,048 and 512 respectively. The compression terms indicate a latent-attention style design in which keys and values are projected into a small shared representation before caching, which directly shrinks the key-value (KV) cache. IntuitionLabs additionally reports identity Hyper-Connections in the residual path and a context length of 1,048,576 tokens. Tencent does not, in the pages I could read, disclose a full training-token count or compute budget, so I will not offer one.
Shared expert plus top-8 routing
Routing eight of 256 experts, with one always-on shared expert, is a deliberate choice about specialization versus stability. The shared expert absorbs common knowledge every token needs, such as syntax and frequent facts, so the routed experts can specialize more narrowly instead of each redundantly relearning the basics. Top-8 out of 256 is relatively fine-grained: many small experts and many active at once give the router more combinatorial freedom than a classic top-2 of 8 layout, at the cost of more all-to-all communication when experts are sharded across devices.
Fine-grained routing also interacts with quantization, which becomes relevant later. When each expert is small and many contribute to every token, the error introduced by aggressively compressing one expert is averaged with seven others plus the shared path. That averaging is one plausible reason expert-heavy models tolerate low-bit weights better than dense models of the same total size, though Tencent has not published an ablation that isolates the effect.
Training and the Self-Improvement Claim
Tencent has disclosed little about pre-training data scale, token counts or compute, so this section is mostly about what is reported rather than what is known. TechNode quotes the company as saying the training data was informed by the workflows of software engineers, game developers, finance analysts and security specialists, which signals a deliberate tilt toward professional, agentic and coding work rather than general chat. The post-training recipe is not described in the sources I could access, so I cannot say whether it used supervised fine-tuning, preference optimization, reinforcement learning with verifiable rewards, or a mix.
A more unusual claim appears in secondary coverage: that the model participated in optimizing its own training methods and in tuning its own inference-serving stack, with a measured 31.8 percent throughput improvement on serving. I could only find this in a third-party explainer, so treat it as reported by that outlet rather than confirmed by an independent measurement. Even taken at face value it describes engineering automation, a model proposing and testing configuration changes, rather than anything like a model rewriting its own weights. The useful reading is that Tencent is using agentic coding to accelerate infrastructure work, which is plausible and increasingly common.
The MTP head and speculative decoding
The separate 10B-parameter MTP layer, with roughly 0.7B parameters activated per step according to the model card, is trained to predict multiple future tokens. At inference it acts as a built-in draft model, so there is no need to find and align a separate small model. Because the draft head shares the backbone’s representations, its acceptance rate tends to be higher than a generic small model’s, which is why integrated MTP heads have become a standard feature of recent large releases. The model card documents speculative decoding through this layer for both vLLM and SGLang.
Sherry Quantization: How 1.25 Bits Per Weight Works
Sherry is a ternary quantization method from researchers at City University of Hong Kong, Tencent and McGill University. It constrains every block of four weights to contain exactly three non-zero values (each plus or minus one) and one zero, then packs each block into five bits. Five bits over four weights is 1.25 bits per weight, a layout that fits SIMD hardware cleanly.

Figure 2: From BF16 to the reported 214 GB checkpoint. The 1.25-bit figure is the ideal Sherry rate; the whole-model average is higher once other tensors are kept at more bits.
The arithmetic of the packing is the elegant part. In each group of four weights, one position is forced to zero, so there are four choices for where the zero sits. The other three positions each take a sign, giving 2 cubed, or 8, sign patterns. Four times eight is 32 distinct block patterns, and 32 values need exactly 5 bits. The paper describes this as one sign-and-index scheme that maps neatly onto 128-bit SIMD registers without bit-shuffling.
Why 3:4 sparsity instead of plain ternary
Standard ternary quantization allows each weight to be -1, 0 or +1, which carries log2(3), about 1.58 bits of information. Packing three-valued symbols efficiently is awkward, because 3 is not a power of two. The common workaround packs three weights into 5 bits, about 1.67 bits per weight, but that layout misaligns with how CPUs and accelerators load data. Sherry’s fine-grained 3:4 structured sparsity gives up some freedom, since exactly one of four must be zero, in exchange for a layout that is both smaller and aligned. The paper reports roughly 25 percent bit savings and about a 10 percent speedup against a prior ternary baseline on an Intel i7-14700HX CPU.
The cost of a hard structural constraint is that the optimizer has less room. The authors identify a failure they call weight trapping: under the 3:4 constraint, weights collapse toward a binary-like distribution and gradients homogenize, which hurts accuracy. Their remedy is a training-time module named Arenas that injects an annealing residual path, decaying to zero by the end of training, so the final network pays no inference overhead for it. This is quantization-aware training (QAT), not post-training rounding, which is a crucial distinction for what follows.
What the paper actually demonstrates
The Sherry paper evaluates LLaMA-3.2 at 1B and 3B scale on PIQA, ARC-Easy, ARC-Challenge, HellaSwag and WinoGrande. At 1B it reports an average accuracy of 0.519, matching the state-of-the-art ternary baseline with the bit and speed gains above; at 3B it reports 0.567 against 0.576 for the strongest baseline. The authors themselves list the limitation that validation stops at models up to 3B parameters, that only weights are quantized while activations stay in BF16, and that QAT adds training cost. Nothing in the paper, as I read it, addresses a 770B MoE.
That gap is the right place for skepticism. The jump from a 3B dense model to a 770B sparse model is two orders of magnitude, and the paper’s evidence does not cover it. What does exist is Tencent’s own announcement for Hy4, which I treat in the next section as a vendor claim with partial third-party echo.
The 214 GB Claim: What Is Reported and What Is Not
Tencent AI’s official account reportedly stated that Hy4 preview went from 1.5 TB to 214 GB using Sherry, calling it seven times smaller with barely a dent in quality. A community post by Ivan Fioravanti, as quoted in search results, lists BF16 versus Sherry scores of MCP Atlas 83.7 to 83.2, SWE-Bench Multilingual 82.9 to 81.3, and MRCR 81.3 to 81.1. I could not open either post directly, so those figures are second-hand quotations from search snippets and should be verified against the originals.
Two details deserve care. First, 1.5 TB divided by 214 GB is about seven, consistent with the “seven times smaller” phrasing. Second, 214 GB across 770 billion weights works out to roughly 2.2 bits per weight on average (illustrative arithmetic: 214e9 bytes times 8 bits, divided by 770e9 weights). That is well above 1.25. The gap implies that Sherry applies to most of the bulk of the model, while some tensors, such as embeddings, attention projections, routers, norms, the shared expert and possibly the MTP layer, stay at higher precision.
GIGAZINE’s account adds a wrinkle. It describes a 1-bit variant at 213 to 214 GiB that alternates 1.31-bit STQ1_0 and 2.06-bit IQ2_XXS quantization types, alongside a Q4_K_M build at 435 GiB. Those names resemble llama.cpp-ecosystem formats, which suggests community tooling may be involved in at least some of the shipped files. The sources I read do not state clearly whether the 214 GB artifact is a pure Sherry output, a llama.cpp conversion of it, or a hybrid, so I would not assume any particular file format or runtime until you read the actual release notes.
Why accuracy can hold up better than intuition suggests
Low-bit survivability at this scale rests on three mechanisms, none of which Tencent has formally published an ablation for in the material I saw. First, MoE redundancy: with 256 experts per layer and many co-active, individual weight noise averages out across experts. Second, QAT: the network is trained or fine-tuned while experiencing the quantization constraint, so it learns to compensate, unlike post-training rounding. Third, mixed precision: sensitive tensors stay wide, so the structural bottlenecks of the network retain fidelity while the bulky expert matrices absorb the compression.
None of this guarantees that every task survives. Benchmark deltas of a few tenths of a point on aggregate scores can hide larger drops on specific long-tail behaviors such as rare-language knowledge, precise numerical recall, or long-context retrieval at the far end of the window. The reported MRCR near-parity is encouraging for long-context retrieval, but a single benchmark is not a guarantee, and you should run your own evaluation before committing a workload.
Capabilities and Benchmarks, With Caveats
The scores on the model card, as reproduced by secondary sources, are GPQA Diamond 92.3, SWE-bench Pro 65.7, DeepSWE 64.3, SWE-bench Multilingual 82.9 and SkillsBench v1.1 62.9. These are self-reported by Tencent. GPQA Diamond is a graduate-level science question set where the top of the leaderboard has compressed into a narrow band, so a 92.3 tells you the model is in the frontier cluster but not much about ordering within it. The software-engineering numbers are more discriminating because they measure end-to-end repair and implementation on real repositories.
| Benchmark | Reported score | Source type |
|---|---|---|
| GPQA Diamond | 92.3 | Vendor model card |
| SWE-bench Pro | 65.7 | Vendor model card |
| DeepSWE | 64.3 | Vendor model card |
| SWE-bench Multilingual | 82.9 (BF16) | Vendor model card |
| SkillsBench v1.1 | 62.9 | Vendor model card |
The most quoted comparison is not a standard benchmark at all. Tencent ran a blind evaluation in which 163 internal experts compared model outputs on 203 engineering tasks. Against GLM-5.3, Hy4 preview won 46.8 percent, tied 12.8 percent and lost 40.4 percent. Against Kimi K3 it won 51.2 percent, tied 7.9 percent and lost 40.9 percent. TechNode summarizes this as a mean rating of 2.99 out of 4 versus 2.92 for GLM-5.3 and 2.94 for Kimi K3. Tencent states the evaluation has not been independently verified.
Reading the blind evaluation honestly
A win rate of 46.8 against a loss rate of 40.4 is a lead of 6.4 points over 203 tasks. As an illustrative statistical check, with about 203 paired comparisons the standard error on a win-minus-loss difference is on the order of five to six points, so the Hy4 versus GLM-5.3 gap sits roughly within one standard error of zero. The Kimi K3 gap, about ten points, is somewhat stronger evidence but still modest. The raters were Tencent employees, which introduces plausible in-group bias even if outputs were anonymized, and the task mix was chosen by Tencent.
None of that makes the result meaningless. It is a decent signal that Hy4 preview is competitive with its two named peers on engineering work. It is not evidence of a clear win, and it should not be cited as “beats Kimi K3.” Community testers reportedly found weak spots on SWE-Marathon, ProgramBench, HorizonMath and BrokenArXiv, according to IntuitionLabs, which fits the preview framing and the vendor’s own admission that the model sometimes reasons longer than necessary and over-verifies answers.
Long context: a feature to test, not assume
A one-million-token window is a specification, not a guarantee of usable recall. Effective context, the length at which retrieval and reasoning stay reliable, is usually shorter than the advertised maximum, as we examined in our piece on long-context LLM benchmarks and RULER effective context. Sparse attention helps economics, because attending to a selected subset of past tokens cuts the quadratic cost, but selection can miss the needle. The quantized checkpoint’s reported MRCR score of 81.1 against 81.3 in BF16 is one data point that long-range retrieval survived compression; it does not tell you how either performs at 900K tokens on your documents.
Access and Deployment
Hy4 preview is distributed as open weights, in full precision and FP8, through Hugging Face, ModelScope, GitCode and CNB under the Apache License 2.0. Apache 2.0 is permissive: commercial use, modification and redistribution are allowed with attribution and a patent grant, with no field-of-use restrictions of the kind that appear in some “open” model licenses. Check the repository’s LICENSE and any model-specific notices before shipping, since I read the license summarized rather than the file.
For hosted access, the model is reported available through Tencent Cloud TokenHub and through OpenRouter under the identifier tencent/hy4-preview. Pricing as reported by DataNorth and others is $0.834 per million input tokens, $2.501 per million output tokens and $0.042 per million cached-input tokens. TechNode reported an initial two-week free period from launch, which has presumably ended by the time you read this, so confirm current rates before budgeting.

Figure 3: Reported checkpoint sizes and the kind of hardware each class implies. Device groupings are illustrative sizing, not Tencent recommendations.
Serving stack
The model card provides a Docker image and recipes for vLLM and SGLang, with IntuitionLabs noting vLLM 0.29.0 or later. Both engines support tensor and expert parallelism, which you will need: no practical single-GPU deployment exists at this size. Our comparison of vLLM, SGLang and TensorRT-LLM is a good starting point for choosing between them, and the choice matters more for a model with fine-grained routing, because all-to-all communication efficiency dominates throughput.
Illustrative hardware sizing
The following is illustrative arithmetic, not a benchmark and not vendor guidance. The reported sizes are: BF16 about 1.5 TB, Q4_K_M about 435 GiB, and the Sherry-class build about 214 GiB, plus an FP8 build that, at one byte per weight, should land near 770 GB plus the roughly 10B-parameter MTP layer and scale metadata.
| Precision | Approx. weight memory | Example pooling (illustrative) |
|---|---|---|
| BF16 | ~1.5 TB | Two 8-GPU nodes of 141 GB cards (2.25 TB) |
| FP8 | ~0.8 TB | One 8-GPU node of 141 GB cards (1.13 TB) |
| Q4 class | ~435 GiB | Four 128 GB unified-memory machines (512 GB) |
| Sherry class | ~214 GiB | Two 128 GB machines (256 GB), tight; three safer |
Weights are not the whole budget. You also need KV cache, activations and runtime overhead. With a latent-attention design the KV cache per token is small, but at hundreds of thousands of tokens it still runs to tens of gigabytes per sequence, so plan headroom of 15 to 30 percent over raw weights for modest context, and more if you intend to use the long window.
Throughput ceilings from memory bandwidth
Decode speed at batch size one is bounded by how many bytes must be read per token. As an illustrative upper bound, 49B active parameters at an effective 2.2 bits per weight is about 13.5 GB per token (49e9 times 2.2 divided by 8). On a pooled system whose memory delivers an aggregate 250 GB/s, that bounds decode near 18 tokens per second before any communication overhead, expert cache misses or dequantization cost. At FP8, the same calculation gives 49 GB per token, so the same system would cap near 5 tokens per second. These ceilings are optimistic, and real numbers will be lower, but they show why low-bit weights help decode speed as well as capacity.
This is also where machines with large unified memory become interesting. A workstation class like the one in our look at the AMD Ryzen AI Max Pro 400 for local LLM inference offers a lot of addressable memory per dollar, though with far less bandwidth than datacenter GPUs. Tencent’s announcement reportedly describes pooling GPUs you already own across machines. I could not read that post, so I do not know the networking requirements or the software path. Expert-parallel inference over ordinary Ethernet is hard, because all-to-all traffic every layer punishes latency, so expect a throughput-oriented, not interactive, experience until someone publishes measurements.
How a Request Flows Through Hy4 at Inference Time
Understanding the request path explains where the cost lands and where the failure modes live. A prompt arrives at the serving engine, which runs prefill: the backbone processes all prompt tokens in parallel, building the compressed KV cache with sparse attention. Prefill is compute-bound and benefits from the 49B active footprint. For a very long prompt, prefill dominates time to first token, and the cost grows with prompt length even when attention is sparse.

Figure 4: Decode loop with the integrated MTP head drafting tokens and the backbone verifying them in a single pass.
Decode then proceeds in a loop. The MTP head drafts several tokens, the backbone scores them in one forward pass, and the engine accepts the longest correct prefix. When the draft is mostly right, you generate multiple tokens per backbone pass, which matters because each pass must stream the selected experts’ weights from memory. When the draft is wrong, you fall back to one token per pass plus the drafting overhead. Acceptance rates vary by task: code and structured output tend to be highly predictable, free-form creative text less so.
Where quantization changes the picture
Weight-only quantization, which is what the Sherry paper addresses, reduces memory traffic but does not by itself reduce arithmetic: weights are dequantized to a higher-precision type before multiplication unless the kernel is specially built for ternary arithmetic. The Sherry paper’s design goal is precisely that special kernel, with 5-bit blocks that unpack without bit shuffling on SIMD hardware. The paper’s speed numbers come from a CPU, not from a GPU cluster, so I would not extrapolate its 10 percent speedup to a multi-node Hy4 deployment. What transfers is the bandwidth saving: fewer bytes per weight mean fewer bytes moved.
Activations remain in BF16 in the paper’s setting, which keeps accuracy higher but means the compute path still involves wide types. A deployment that wants the full benefit needs inference kernels that fuse dequantization with the matrix multiply. If your engine lacks them, you may see a smaller-than-expected speed gain even though the memory footprint shrinks as advertised. This is a software maturity question, and it is the main reason I would wait for runtime-specific benchmarks before planning production traffic around the 214 GB build.
Cost per token versus cost per node
For an operator choosing between the API and self-hosting, the comparison is not about price per million tokens alone. The reported API rate of $2.501 per million output tokens means a workload generating a billion output tokens a month costs on the order of $2,500 in output charges (illustrative, ignoring input and caching). A multi-GPU node costs far more than that per month to rent, so self-hosting only wins at sustained high utilization, or when data residency, latency control or customization outweighs raw cost. Our analysis of AI inference cost optimization walks through the utilization break-even logic that applies here.
Trade-offs, Gotchas, and What Goes Wrong
The first gotcha is the word “preview.” Tencent describes this as an early version, with acknowledged tendencies to reason longer than needed on complex tasks and to over-verify. In an agent loop, over-thinking is not just a quality quirk, it is a latency and cost multiplier, because each tool-use step may consume a long reasoning trace. Test with realistic agent traces and measure tokens per completed task, not only accuracy.
The second is evaluation provenance. Nearly every number in circulation is self-reported or comes from the vendor’s blind study with employee raters. Independent leaderboards will take weeks to settle, and the community findings of weakness on several harder benchmarks suggest the picture will be more mixed than launch headlines. If your decision hinges on a benchmark delta of a few points, run your own evaluation on your own data.
Quantization-specific risks
The 214 GB build is not the same model as the BF16 checkpoint. Aggregate benchmark deltas look small in the reported figures, with the largest quoted drop being about 1.6 points on SWE-Bench Multilingual, but behavior differences are likely to concentrate in places aggregate scores miss: calibration of uncertainty, rare tokens, exact copy tasks and tool-call formatting. A one-point drop on a benchmark can correspond to a much larger failure rate on a narrow production task such as emitting strictly valid JSON for a particular schema.
There is also a reproducibility risk. If your numbers were produced with a specific kernel, a different runtime may yield different outputs, because low-bit formats depend on implementation details such as block layout and scale handling. Pin the checkpoint hash, the engine version and the kernel build, and re-run your regression suite on any upgrade. Our guidance on constrained decoding for structured output applies directly if you need schema-valid responses from a heavily quantized model.
Operational failure modes of large MoEs
Expert load imbalance is the classic one. If routing sends a disproportionate share of tokens to a few experts, the devices holding those experts become stragglers, and aggregate throughput falls to the speed of the slowest shard. Real traffic is rarely uniform: a coding-heavy workload will concentrate on a different expert subset than a translation workload. Monitoring per-expert token counts and device utilization is a prerequisite for tuning, not an optional nicety.
Communication is the second. Expert parallelism requires an all-to-all exchange in each MoE layer, 77 times per forward pass in Hy4. Over NVLink or fast InfiniBand this is manageable; across commodity networks it can dominate. Finally, memory fragmentation and KV cache growth at long context can cause out-of-memory failures well after a deployment looks healthy, so set admission limits on context length per request rather than relying on the headline 1M window.
How It Compares
Because the benchmarks are largely vendor-reported, I would compare these models on deployment characteristics and on the evidence quality for each use case, rather than declare a winner. The matrix below reflects what the cited sources support and flags where I have no verified data. GLM-5.3 and Kimi K3 details are limited to what the Hy4 coverage says about them, since I did not research those models independently for this post.
| Use case | Tencent Hy4 preview | GLM-5.3 | Kimi K3 |
|---|---|---|---|
| Agentic coding evaluation | Competitive in vendor blind study, 46.8% win vs GLM-5.3 | Lost slightly in that study, not independently verified | Lost slightly in that study, not independently verified |
| Self-hosting an open-weight frontier model | Apache 2.0, FP8 and reported 214 GB build | Not assessed here | Not assessed here |
| Hosted API cost | Reported 25-36% cheaper than GLM-5.3 and 70-82% cheaper than Kimi K3 (aggregator figures) | Reference price | Reference price |
| Long-context retrieval | 1M window, reported MRCR 81.3 BF16 | Not assessed here | Not assessed here |
The cost row comes from a single secondary source, so treat the percentages as indicative. If cost is your deciding factor, pull current price pages for all three on the day you decide, because preview pricing and promotional periods change quickly.
Practical Recommendations
Start with the hosted API, not the weights. It lets you test task quality, token consumption and the over-thinking behavior against your real prompts for very little money, and it separates the question “is this model good for my workload” from “can I afford to run it.” Only once the first answer is yes should you invest in hardware.
If you do self-host, begin with the FP8 checkpoint on a node whose aggregate memory covers weights plus at least 25 percent headroom, using the vLLM or SGLang recipe from the model card. Treat the Sherry-class build as a second step, justified when capacity or cost, not accuracy, is your binding constraint, and validated with your own regression suite before it carries traffic. Compare the two builds on the same prompts and the same sampling settings so that differences are attributable to quantization.
A short checklist for a pilot:
- Read the model card and LICENSE directly, and record the checkpoint hash you deploy.
- Run 100 to 200 of your own tasks on the API first and log tokens per completed task.
- Cap per-request context length to what your memory budget can hold, well below 1M.
- Enable speculative decoding through the MTP layer and measure acceptance rate on your traffic.
- Monitor per-expert load and all-to-all latency, not only GPU utilization.
- Evaluate FP8 versus the 214 GB build on schema validity, tool-call accuracy and long-context retrieval.
- Re-verify third-party claims, including those in this post, against primary sources before committing budget.
Frequently Asked Questions
What is Tencent Hy4?
Tencent Hy4 preview is an open-weight Mixture-of-Experts language model announced on 28 August 2026. It has 770 billion total parameters with 49 billion active per token, a one-million-token context window and an Apache 2.0 license. It uses 78 layers, 256 routed experts plus one shared expert per MoE layer, and a built-in multi-token-prediction head for speculative decoding. Tencent describes it as an early preview and its benchmark claims are self-reported.
What is Sherry quantization?
Sherry is a ternary quantization method from researchers at City University of Hong Kong, Tencent and McGill University. It forces three of every four weights to be plus or minus one and one to be zero, giving 32 block patterns that fit in five bits, or 1.25 bits per weight. A training-time module called Arenas prevents weight collapse. The paper validates it on LLaMA-3.2 models up to 3B parameters, so large-model results come from Tencent’s own announcement.
How big is Hy4 after quantization?
The BF16 checkpoint is about 1.5 TB. Tencent AI reportedly states the Sherry-compressed version is 214 GB, roughly seven times smaller, and GIGAZINE reports a 214 GiB 1-bit variant plus a 435 GiB Q4_K_M build. An FP8 checkpoint is also published. Because 214 GB over 770 billion weights averages about 2.2 bits, not 1.25, some tensors evidently stay at higher precision. Verify exact sizes against the release files.
Can I run Tencent Hy4 locally?
Not on a single consumer machine. Even at 214 GB you need roughly 256 GB or more of pooled fast memory plus room for KV cache, which means a multi-device setup, such as several large unified-memory workstations or multiple GPUs. Expect modest decode speeds on bandwidth-limited hardware, and treat networked pooling as experimental until measurements are published. For most teams, trying the hosted API first is far more practical.
Is Tencent Hy4 really open source?
It is open weights under the Apache License 2.0, which permits commercial use, modification and redistribution. That is more permissive than many restricted “open” licenses. It is not necessarily open training: Tencent has not disclosed the training data or full recipe in the sources I reviewed. So you can run, fine-tune and ship it, but you cannot fully reproduce how it was made.
How does Hy4 compare with GLM-5.3 and Kimi K3?
In Tencent’s own blind study of 203 engineering tasks judged by 163 internal experts, Hy4 preview won 46.8 percent against GLM-5.3 and 51.2 percent against Kimi K3, with losses of 40.4 and 40.9 percent. The study has not been independently verified, and the margin against GLM-5.3 is within plausible noise. Community tests reportedly show weaker results on some harder benchmarks, so evaluate on your own workload.
Further Reading
- Expert-parallel MoE inference serving architecture for how to shard and schedule models with hundreds of experts.
- AI inference cost optimization for utilization break-even logic between APIs and self-hosting.
- AMD Ryzen AI Max Pro 400 for local LLM inference for large unified-memory workstation sizing.
- FP8 vs INT8 vs INT4 LLM quantization benchmark for the wider quantization trade-off landscape.
- Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization (arXiv 2601.07892) and the Hy4-preview model card as primary sources.
By Riju — about
