Index-Translate-35B-A3B Explained: Bilibili’s MoE Translation Model for 150 Languages
A video platform released a translation model that beats a field of specialist systems on a headline benchmark, and the most interesting thing about it is how little of it runs per token. Index-Translate-35B-A3B is a Mixture-of-Experts (MoE) model from Bilibili’s Index LLM team: 35 billion parameters in total, roughly 3 billion active for any one token, Apache-2.0 licensed, built on Alibaba’s Qwen3.5 foundation, and trained to translate text across 150 languages. It is explicitly labelled a preview, and the family around it covers subtitles, dubbing and book-length documents.
It matters now because translation is one of the few LLM workloads where the quality bar is measurable, the latency budget is tight, and the data often cannot leave your network. A small-active-parameter model with a permissive license changes the build-versus-buy arithmetic for localization pipelines, subtitle tooling and multilingual support systems.
This post reads the primary sources only: the Hugging Face model card, the GitHub repository and the vendor demo page. You leave with the architecture, the training recipe as disclosed, the benchmark numbers with their caveats, a sizing calculation you can reuse, and an honest list of what is not yet public.
What this covers: the model family and lineage, the Qwen3.5-derived architecture, the three-phase training pipeline, every published score and what it does and does not prove, serving and hardware sizing, failure modes, and a decision matrix against alternatives.
Context and Background
Machine translation split into two camps over the last few years. Dedicated neural machine translation (NMT) systems such as encoder-decoder transformers are small, fast and cheap, but they translate sentence by sentence and follow instructions poorly. General-purpose large language models handle context, tone and formatting constraints, but they are expensive per token and uneven in low-resource languages. The practical question for a localization team has been which of the two failure profiles it can live with.
Index-Translate is Bilibili’s attempt to collapse that choice. Bilibili is a large Chinese video platform, so its translation problem is unusually specific: subtitles that must fit on screen, dubbing where the target sentence must fit the spoken duration, internet slang and memes, and long serialized fiction. The Index LLM team built a family rather than a single model, and the text model is the foundation the rest are derived from. If you want the serving mechanics behind models of this shape, our explainer on expert-parallel MoE inference serving architecture covers how sparse experts map onto GPUs.
The release dates need care. The GitHub README states a release date of September 30, 2026, and the model card citation points to September 2026 with a technical report on arXiv (identifier 2609.40181 as given on the model card). Press coverage, including Pandaily’s report, describes an open-source family built on Qwen3.5 covering 150 languages under Apache-2.0. I could not open the arXiv abstract during this run because the fetch proxy rate-limited the request, so everything below that goes beyond the model card and README is flagged as such.
Two facts frame everything that follows. First, the 35B-A3B model is a preview: the README’s TODO list names an “official” Index-Translate-35B-A3B release as future work. Second, the 2B and 9B siblings are not previews, and on several measures the 9B model is within noise of the 35B-A3B. Reading the benchmark tables with that in mind is the difference between a useful evaluation and a press-release summary.
The model family at a glance
The model card and README list the following members. All are described as built on Qwen3.5.
| Model line | Sizes | Task |
|---|---|---|
| Index-Translate | 2B, 9B, 35B-A3B (preview) | Text translation, 150 languages, instruction following |
| Index-Echo S2TT | 2B, 9B | Speech to translated subtitles |
| Index-Echo S2ST | 2B, 9B | Speech-to-speech dubbing with timbre preservation |
| Index-Homura | 2B, 9B | Syllable-count controlled translation for dubbing |
| Index-NaiLong (README also writes NativeLong) | 2B, 9B | Long-document translation |
Note the naming inconsistency: the README prose says NativeLong while the model IDs read Index-Nailong. The vendor demo page spells it NaiLong. Use the Hugging Face model IDs when scripting, not the prose names.
Index-Translate-35B-A3B Architecture: A Hybrid-Attention MoE Inherited from Qwen3.5

Figure 1: Layer pattern of the Qwen3.5 MoE backbone that Index-Translate-35B-A3B builds on, with three linear-attention blocks for every full-attention block and a router selecting eight of 256 experts per token.
Direct answer: Index-Translate-35B-A3B is a sparse Mixture-of-Experts decoder with 35 billion total and about 3 billion active parameters. It inherits Qwen3.5’s hybrid layout of Gated DeltaNet linear attention interleaved with Gated Attention, supports a 262,144-token position range, and is fine-tuned for translation across 150 languages under the Apache-2.0 license.
The model card states the headline numbers directly: base model Qwen3.5, MoE architecture, 35B total parameters, 3B active, max_position_embeddings of 262,144, and a tensor type of BF16 with the Hub reporting 36B parameters. The 35B versus 36B difference is a counting convention, most plausibly the multi-token-prediction head and embeddings, but the card does not say, so treat it as unexplained.
The finer architectural detail is not on the Index card at all. It lives on the Qwen3.5-35B-A3B card, and the Index checkpoint’s configuration declares the same Qwen3_5MoeForConditionalGeneration class. I am therefore reporting the following as the base model’s published architecture, and assuming, without Index-side confirmation, that the fine-tune did not change the topology. That assumption is supported by the matching parameter counts and context range, but it is an inference.
What the backbone looks like
Per the Qwen3.5-35B-A3B card: 40 layers, hidden dimension 2,048, laid out as ten repetitions of three Gated DeltaNet blocks followed by one Gated Attention block, each paired with an MoE feed-forward layer. The MoE layer has 256 experts, with 8 routed experts plus 1 shared expert active per token and an expert intermediate dimension of 512. The vocabulary is 248,320 entries (padded). The base model supports 262,144 tokens natively and is documented as extensible to roughly 1,010,000.
The config file adds the attention geometry. The Gated DeltaNet layers use 16 query/key heads and 32 value heads at head dimension 128 with a convolution kernel of 4. The full-attention layers use 16 query heads and only 2 key-value heads at head dimension 256, with partial rotary embeddings. A full_attention_interval of 4 confirms the 3:1 ratio. The config also lists one multi-token-prediction layer and a router auxiliary loss coefficient of 0.001. The backbone additionally carries a vision encoder, which the translation fine-tune does not advertise.
Why this layout suits translation serving
Two properties matter for a translation workload and both are consequences of the design rather than of Bilibili’s training. The first is sparse activation: with 8 routed plus 1 shared of 257 experts firing, each token touches a small fraction of the feed-forward weights, so per-token compute resembles a 3B dense model while the knowledge capacity resembles a 35B one. The second is that only 10 of the 40 layers hold a growing key-value cache. The other 30 carry a fixed-size recurrent state, so memory growth with sequence length is roughly a quarter of what an all-full-attention stack of the same depth would show. Section three quantifies this with an illustrative calculation.
The catch is that sparse activation reduces compute, not memory. Every expert must be resident, because any token can route to any expert. That is why the weights occupy the memory of a 36B model while the arithmetic is that of a 3B one, and why the serving topic ties into AI inference cost optimization: memory capacity and bandwidth, not FLOPs, set the floor.
Context length: three different numbers
Reading the card closely, there are three context figures that are easy to conflate. The architectural limit is 262,144 positions. The recommended serving context is 32,768 tokens, which is what the documented vLLM command uses via --max-model-len 32768. The evaluation context was 16,384 tokens for text benchmarks and 3,584 for WMT26. The README’s quick-start for the small model even uses 4,096.
So “262K context” is true of the position range and silent about translation quality at that length. Nothing published shows 35B-A3B quality curves beyond the evaluation lengths. If your use case is whole-document translation, the family’s answer is the separate NaiLong line, which was trained on sequences up to 128K tokens and evaluated at 64K, not this model.
Tokenizer and 150 languages
The vocabulary of 248,320 entries is inherited, and a large multilingual vocabulary is a quiet contributor to translation economics: scripts that fragment into many tokens in smaller vocabularies cost more to process and generate. The model card claims text translation across 150 languages and points to the technical report for the documented coverage. I could not retrieve that coverage list, so I cannot say which languages are strong. The card’s own caveat is the right default: scores vary across languages and tasks, and low-resource directions remain harder than core-language translation.
Deeper Analysis: Training, Benchmarks, Serving and Sizing
The disclosed training recipe

Figure 2: The three disclosed phases, multilingual mid-training, specialist SFT and RL, then expert integration by parameter interpolation and multi-teacher on-policy distillation.
The model card gives a compact but unusually specific description of three phases.
Phase one, multilingual mid-training. The model saw 167.77 billion tokens of a mix of general, parallel and monolingual data. During the constant-learning-rate stage the ratio is 1:1:1 across those three categories, and during the decay stage it shifts to 1:4:2, which upweights parallel data four-fold and monolingual twofold relative to general text. That shift is a sensible curriculum: end the schedule with the data that most directly resembles the target task. Compared with the trillions of tokens in a from-scratch pre-training run, 167.77B is small, which fits the framing of a continued-training adaptation of an existing Qwen3.5 checkpoint rather than a new foundation model.
Phase two, specialist SFT and RL. Task-specific supervised fine-tuning and reinforcement learning follow, with three named reward ingredients: XCOMET-XXL, Rubric-as-Reward and RIVAL. XCOMET-XXL is a learned translation-quality metric. Using a neural metric as a reward is powerful and also a known hazard, which I return to in the caveats. Rubric-as-Reward suggests rubric-scored judgments for instruction compliance, which would explain the strong instruction-following numbers. I could not verify what RIVAL stands for or its exact formulation from the model card alone, so I will not guess.
Phase three, expert integration. Rather than shipping one model trained on everything, the team trained specialists and then merged them using parameter interpolation and multi-teacher on-policy distillation. Parameter interpolation, often called weight averaging or model soups, blends checkpoints in weight space. On-policy distillation has the student generate its own outputs and learn from teacher feedback on those outputs, with several teachers here. This pattern reduces the task-interference that plagues multi-objective RL, but it also means the final behavior is partly an average, which can blunt the peaks of individual specialists.
What is not disclosed in the sources I could read: the parallel-data sources, the language-by-language data distribution, the compute used, the number of GPU hours, the RL hyperparameters, and the exact teacher set. The README itself says no detailed pipeline specification appears in its documentation, deferring to the technical report. Anything beyond the three phases above would be invention.
The published scores

Figure 3: How the three text models relate on the published suite. The 35B-A3B leads on aggregate scores but the 9B is close, and wins on one low-resource column.
The README publishes two tables for the three text models. These are the vendor’s own evaluations, not independent ones.
| Model | FLORES COMET-22 | WMT26 Judge | instTrans Quality | instTrans IFscore | MEME |
|---|---|---|---|---|---|
| Index-Translate-35B-A3B | 0.8794 | 76.76 | 0.6901 | 0.8336 | 0.7405 |
| Index-Translate-9B | 0.8789 | 75.35 | 0.6771 | 0.8209 | 0.7387 |
| Index-Translate-2B | 0.8655 | 60.26 | 0.5391 | 0.7569 | 0.6443 |
| Model | FLORES minor COMET-22 | FLORES minor XCOMET-XXL | Off-target rate | instTrans minor Quality | instTrans minor IFscore |
|---|---|---|---|---|---|
| Index-Translate-35B-A3B | 0.8168 | 0.7164 | 2.4% | 0.5151 | 0.7715 |
| Index-Translate-9B | 0.7992 | 0.6805 | 4.0% | 0.5222 | 0.7725 |
| Index-Translate-2B | 0.7377 | 0.4817 | 4.2% | 0.3050 | 0.6586 |
The model card also reports 77.6% on C-Eval and 49.1% on GPQA-Diamond, which indicate that general knowledge and reasoning were not wiped out by translation specialization. The card does not publish the base model’s scores in the same table in the text I retrieved, so the delta against Qwen3.5 is unknown here.
The vendor’s demo page claims first place on FLORES-200 (0.8794) and on instruction-following IFscore (0.8336), and second on MEME (0.7405). The card says the model achieves the highest FLORES COMET-22 score among the comparison systems. I could not retrieve the list of comparison systems, so “first” has no stated denominator in this post. That is the single most important caveat: a ranking is only as meaningful as the field it was drawn from.
Reading the tables critically
Look at the gap between the two larger models. On FLORES COMET-22 the difference is 0.0005, which is far inside any plausible metric noise. On WMT26 the 35B-A3B leads by 1.41 points, and on instTrans quality by 0.013. In the low-resource table the 35B-A3B wins clearly on COMET-22 (0.8168 versus 0.7992), XCOMET-XXL, and off-target rate, but the 9B is marginally ahead on instTrans minor quality and instruction-following. The 2B collapses on WMT26 (60.26) and on low-resource XCOMET.
My reading, labelled as interpretation: for core-language translation, parameter count beyond 9B buys little on these metrics, while for low-resource directions the larger model’s extra capacity shows up in fewer off-target outputs. If you serve high-resource pairs, the 9B dense model is a serious competitor on simplicity alone. If you serve long-tail languages, the 35B-A3B earns its memory footprint. The preview tag also means these numbers may move.
The off-target rate deserves its own note. A 2.4% off-target figure on the minor-language pair set means roughly one output in forty came back in the wrong language, which is a failure no quality metric computed on the intended language will describe well. For a pipeline translating thousands of strings, that rate demands a language-identification check on the output.
Metric caveats every translation reader should apply
COMET-22 and XCOMET-XXL are learned metrics: neural networks trained to predict human quality judgments. They correlate with human ratings better than BLEU or chrF, but they inherit the biases of their training data and can be gamed. When XCOMET-XXL is both a training reward and an evaluation metric, as the model card implies by listing it in the RL phase and in the low-resource table, the evaluation is partly measuring what the model was optimized against. The README does not say whether the evaluation sets were held out from reward modelling, though FLORES is a standard public benchmark and WMT26 is a new evaluation.
The WMT26 Judge score uses an LLM judge, the instTrans suite is the vendor’s own benchmark (5,793 tasks: 3,000 Chinese-to-20-language plus 2,793 low-resource, per the README) and the MEME set holds 3,638 Chinese-to-English examples. The README lists open-sourcing instTrans, SandGlass-V2, nailong-bench and meme-bench as future work, so independent replication of those columns is currently impossible. Only FLORES and WMT can be re-run by third parties today.
Finally, the README notes that the benchmark runs used temperature 0.3 with thinking disabled and that outputs vary between runs, while the model card’s client defaults to temperature 0 with greedy decoding. Those are two different decoding settings. Expect your own reproduction to land near, not on, the published digits, and treat differences in the third decimal as noise.
Serving and Sizing: What It Takes to Run It

Figure 4: A reference serving path for translation workloads, with constraint-bearing prompts going through an OpenAI-compatible vLLM endpoint and a post-hoc language check catching off-target outputs.
What the vendor documents
The model card documents vLLM serving with an OpenAI-compatible endpoint:
pip install -U vllm
vllm serve IndexTeam/Index-Translate-35B-A3B-preview \
--served-model-name IndexTeam/Index-Translate-35B-A3B-preview \
--host 127.0.0.1 --port 8000 --max-model-len 32768
The client examples set temperature=0, max_tokens=1024, and pass chat_template_kwargs with enable_thinking set to false. The card states a CUDA environment is required, a runtime that supports Qwen3.5 MoE and the complete checkpoint is needed, and multi-GPU deployments use tensor-parallel arguments. It publishes no VRAM figure, no throughput figure, and no price. The README says GGUF (for llama.cpp) and FP8 (for vLLM) quantized variants exist for the family, per the vendor page; I did not verify which sizes have them, and in particular whether the preview 35B-A3B has any.
The prompt format is a Chinese-language instruction template even for non-Chinese pairs: “Please translate the following {source-language} text into {target-language}, output the translation directly without any explanation,” followed by the source text. A second template supports constraints: it takes a genre, a source text and a numbered constraints block that distinguishes hard requirements from soft notes, and asks for the translation only. Hard constraints documented include strict glossary enforcement, preservation of structured data such as JSON, CSV and markdown, code blocks and variable placeholders. Soft constraints cover tone, from formal to meme style, and domain disambiguation.
Illustrative sizing arithmetic
None of the following is a vendor figure. It is arithmetic from the published parameter count, and real deployments add activation memory, CUDA context, fragmentation and the runtime’s own buffers.
Weight memory. The Hub lists 36B parameters in BF16, 2 bytes each: about 72 GB for weights alone. At 8-bit (FP8) that halves to roughly 36 GB, and at 4-bit it is about 18 GB, plus quantization scales. So BF16 does not fit a single 48 GB card, FP8 plausibly fits a single 48 GB or 80 GB card with headroom depending on cache needs, and a 4-bit build approaches the territory of a 24 GB card with little room for context. The Ryzen AI Max local inference workstation discussion is relevant here because large unified-memory machines are exactly the class of hardware where a 36B-resident, 3B-active model is attractive.
KV cache. Using the config geometry for the full-attention layers: 10 layers, 2 KV heads, head dimension 256. Per token and per layer, keys plus values are 2 x 2 x 256 = 1,024 elements, or 2,048 bytes at 16-bit. Across 10 layers that is 20,480 bytes, about 20 KB per token. At the recommended 32,768 context that is roughly 0.67 GB per sequence. At the full 262,144 positions it is roughly 5.4 GB per sequence. This excludes the fixed recurrent state of the 30 DeltaNet layers, which I have not sized because I did not verify the state layout. Compare that with a hypothetical stack where all 40 layers were full attention at the same KV geometry: four times the cache. The hybrid design is why a long position range is cheap on paper.
Decode bandwidth bound. At batch size one, decoding is memory-bandwidth bound. If each token reads only the active parameters, roughly 3B at 2 bytes is 6 GB per token in BF16, or 3 GB in FP8. On a card with 1 TB/s of memory bandwidth, the idealized ceiling is about 166 tokens per second in BF16 and 333 in FP8. Real numbers will be lower because routing makes the read pattern irregular and shared components and attention add traffic. At larger batch sizes, different tokens activate different experts, and the union of touched experts grows quickly, so batched throughput approaches the full 72 GB read per step. That is the reason MoE batch economics differ from dense models, a point developed in vLLM vs SGLang vs TensorRT-LLM.
Output length. Translation output is roughly proportional to input length, and the model card’s default max_tokens of 1,024 will silently truncate long paragraphs. Set it from input length, not from a constant.
Using the constraint interface
The instruction-constrained mode is the feature most relevant to production localization. A glossary requirement states that a source term must be rendered in a specific target term; a format requirement says to leave placeholders such as {user_name} untouched. The model card claims hard-constraint compliance is a focus, and the IFscore of 0.8336 is the vendor’s measurement of it. Two cautions follow. The IFscore definition is in the technical report I could not read, so I cannot say whether it is a strict pass rate or a graded mean. And the low-resource IFscore is 0.7715, so combined constraints in rare languages are exactly where the card says performance drops. Always validate placeholder integrity programmatically: a regex over input and output placeholders is cheap and catches the failures a metric averages away.
Trade-offs, Gotchas, and What Goes Wrong
Preview status is a real constraint, not a label. The README lists an official 35B-A3B version as a TODO. Weights, benchmark numbers and even the repository may change. There is also a -backup repository on Hugging Face under the same organization whose purpose is undocumented in what I read. Pin a revision hash, record it with every translation batch, and plan for re-evaluation when the official release lands.
Memory, not compute, is the bill. The “3B active” headline invites the mistake of budgeting for a 3B model. You provision for 36B of resident weights. For low-traffic internal tools a dense 9B may be cheaper to host and nearly as accurate on core pairs, per the table above. The MoE advantage appears when you have enough concurrent traffic to keep GPU compute busy and amortize the large resident footprint.
Quality is uneven by design, and the vendor says so. The card’s own limitation is that low-resource directions and combined constraints remain harder. The 2.4% off-target rate on minor languages is a pipeline concern, not just a score. Add language identification on the output, retry or route to a fallback model when it fails, and keep a human-review sample for any language you cannot read.
Learned-metric inflation. If XCOMET-XXL was optimized in training and reappears in evaluation, expect some of the score to reflect metric fit. Before committing, run a small human or independent-judge evaluation on your own domain text. Legal, medical and technical manuals look nothing like subtitles and memes, and the model was shaped by a platform whose corpus is video-centric.
Prompt language and template drift. The documented templates are in Chinese. English instructions may work, but the vendor’s numbers came from the documented templates, so deviating from them is a variable you must test. The same applies to the enable_thinking flag: the card’s examples disable it, and enabling reasoning is untested in the published results.
Greedy decoding hazards. Temperature 0 gives determinism but can produce repetition loops on long outputs, and the README’s evaluation used 0.3. Neither card documents repetition penalties. Monitor for outputs that stretch to the token limit, which is the typical symptom.
Context claims. A 262,144-position range does not imply tested quality at that length for translation. For documents, chunk at paragraph boundaries within the 32K serving context, carry a glossary in every chunk, and consider the NaiLong model for book-length work. Our analysis of long-context benchmarks and effective context explains why advertised and effective context diverge.
Safety and content. I found no published red-team results or safety evaluation for the translation model. Translation models can faithfully translate harmful content, since fidelity is the objective, and a source-language filter is a separate component you own. Chinese-origin models also embed the content-governance norms of their training pipeline, which can surface as refusals or softened renderings on sensitive topics. Test your own sensitive-content cases rather than assuming neutrality.
How It Compares: A Decision Matrix
No head-to-head table against named competitors can honestly be built from the sources I read, because the comparison field in the vendor’s ranking was not retrievable and no third-party evaluation of the preview exists yet. What I can do is compare on structural properties that are verifiable, and be explicit that quality columns are judgments to confirm with your own test set.
| Use case | Index-Translate-35B-A3B preview | Index-Translate-9B | Dedicated NMT service or small NMT model | Frontier general LLM via API |
|---|---|---|---|---|
| High-volume core-language pairs, tight cost | Good, but 36B resident memory is hard to justify | Strong default: near-identical FLORES score, simpler hosting | Cheapest and fastest, weak on instructions | Highest unit cost |
| Long-tail and low-resource languages | Best of the family on published low-resource COMET and off-target rate | Acceptable, 4.0% off-target versus 2.4% | Often limited language list | Variable, test per language |
| Glossary and placeholder constrained localization | Supported by design, validate programmatically | Comparable IFscore on published sets | Usually glossary features, no free-form constraints | Strong but no weights to audit |
| Data cannot leave the network | Yes, Apache-2.0 weights | Yes | Depends on vendor | No, unless private deployment |
| Production stability needed today | Preview, pin and re-test | Non-preview release | Mature | Mature, but model versions rotate |
The honest summary of the matrix: the 35B-A3B is a specialist for the long tail and for teams that already operate MoE-capable GPU serving. The 9B is the safer default. A related build-versus-API tradeoff appears in the open-weight analysis of GLM-5.2 benchmarks, and for another sparse model with a very different scale see Tencent HY4 770B MoE and Sherry quantization.
A reasonable evaluation protocol
Before adopting any translation model, assemble 300 to 500 segments from your real content, including your hardest examples: placeholders, glossary terms, ambiguous pronouns, numerals and units, and mixed-language strings. Translate with each candidate at fixed decoding settings and at least two seeds if you use nonzero temperature. Score with a reference-free quality metric for a quick ranking, then have bilingual reviewers grade a stratified sample on adequacy, fluency and terminology, with a separate tally for catastrophic errors such as dropped sentences or wrong language. Track the catastrophic-error rate separately from the mean score, because a single mistranslated dosage or price outweighs hundreds of small stylistic gains.
Then measure the operational side: tokens per second at your expected concurrency, p95 latency for a typical segment, and cost per million characters. The published sources give none of those numbers, so this measurement is entirely yours, and it is where the MoE choice stands or falls.
Practical Recommendations
Treat Index-Translate-35B-A3B as a promising preview to evaluate, not yet a component to build a critical path around. Begin with the 9B model for core-language work, and bring in the 35B-A3B only for the language pairs where your own test set shows the 9B falling short, most plausibly the low-resource ones the published tables point to.
Keep the architecture of your pipeline model-agnostic: an OpenAI-compatible endpoint, a thin translation service that owns the prompt template, glossary injection and placeholder validation, and an output language check. That lets you swap the preview for the official release, or for a competitor, without touching callers. For serving, start at the documented 32,768 maximum length and tune from there, and size memory for 36B resident parameters at your chosen precision before you price hardware.
Do not skip the unglamorous checks. The off-target and placeholder failure modes are cheap to detect and costly to ship. And budget a re-evaluation when the official release or an independent benchmark appears, because every number in this post is the vendor’s.
Adoption checklist
- [ ] Pin the exact Hugging Face revision of the preview and log it with each batch.
- [ ] Build a 300 to 500 segment in-domain test set with glossary and placeholder cases.
- [ ] Compare the 9B and 35B-A3B per language pair, not on an aggregate score.
- [ ] Compute memory: 36B parameters at your precision, plus KV cache at your real context length.
- [ ] Use the documented prompt templates and the same decoding settings as your baseline.
- [ ] Set
max_tokensfrom input length rather than the 1,024 default. - [ ] Add output language identification and a retry or fallback route.
- [ ] Validate placeholders, numerals and structured data with code, not with the model.
- [ ] Run a human review on a stratified sample for languages your team cannot read.
- [ ] Re-test when the official 35B-A3B release or independent evaluations appear.
Frequently Asked Questions
What is Index-Translate-35B-A3B?
Index-Translate-35B-A3B is a preview Mixture-of-Experts translation model from Bilibili’s Index LLM team. It has 35 billion total parameters with about 3 billion active per token, is built on Alibaba’s Qwen3.5, supports text translation across 150 languages, and is released under the Apache-2.0 license. It sits at the top of a family that also includes 2B and 9B text models plus speech, dubbing and long-document variants.
How many parameters does Index-Translate-35B-A3B actually use per token?
About 3 billion are active per token, out of 35 billion total, though the Hugging Face page lists 36B in BF16. The Qwen3.5 base routes each token to 8 of 256 experts plus 1 shared expert. Compute therefore resembles a 3B dense model, but all weights must stay in memory because any token can reach any expert, so budget for roughly 36B resident parameters.
Is Index-Translate-35B-A3B better than the 9B model?
Not uniformly. On the published tables the 35B-A3B leads on WMT26 (76.76 versus 75.35), instTrans quality, and low-resource COMET-22 and off-target rate (2.4% versus 4.0%). But the FLORES COMET-22 gap is only 0.0005, and the 9B is marginally ahead on low-resource instTrans quality and instruction-following. For core-language pairs the 9B is often the pragmatic choice.
Can I use Index-Translate-35B-A3B commercially?
The model card and README state Apache-2.0, which permits commercial use, modification and redistribution with attribution and license notices. Two cautions: the 35B-A3B is a preview that may change, and the Qwen3.5 base is also listed as Apache-2.0 on its own card. Have counsel read the actual LICENSE files in the repositories you download from, since training-data terms are not covered by the weight license.
What hardware do I need to run Index-Translate-35B-A3B?
The vendor publishes no VRAM figure. From the 36B parameter count, illustrative weights-only memory is about 72 GB in BF16, 36 GB in FP8 and 18 GB at 4-bit, before KV cache and runtime overhead. The card requires CUDA and a vLLM-style runtime that supports Qwen3.5 MoE, with tensor parallelism for multi-GPU. Plan for an 80 GB class GPU at FP8 and measure throughput yourself.
How reliable are the published translation benchmark scores?
They are the vendor’s own results, and several come from internal benchmarks (instTrans, MEME) that are not yet open-sourced. COMET-22 and XCOMET-XXL are learned metrics, and XCOMET-XXL also appears as a training reward, which can inflate scores. The set of comparison systems behind the “first place” claims was not retrievable. Treat the numbers as a starting hypothesis and verify on your own domain text.
Further Reading
- Expert-parallel MoE inference serving architecture for how sparse experts are placed and routed across GPUs.
- AI agent frameworks benchmark: LangGraph, OpenAI and Google ADK if you plan to wrap translation in an agentic localization workflow.
- AI inference cost optimization for the memory-versus-compute cost levers behind MoE deployments.
- AMD Ryzen AI Max Pro 400 local LLM inference workstation for unified-memory hardware suited to large resident, small active models.
- INT4 vs INT8 vs FP8 quantization for edge NPUs for the precision trade-offs behind the sizing above.
- External: the Index-Translate GitHub repository and the Hugging Face model card.
References
- IndexTeam, “Index-Translate-35B-A3B-preview” model card, Hugging Face. https://huggingface.co/IndexTeam/Index-Translate-35B-A3B-preview
- Bilibili Index LLM Team, “Index-Translate: A Multilingual Translation Model Family,” GitHub README. https://github.com/bilibili/Index-Translate
- Qwen Team, “Qwen3.5-35B-A3B” model card and configuration, Hugging Face. https://huggingface.co/Qwen/Qwen3.5-35B-A3B
- Index-Translate project page and online demo. https://index-translate.bilibili.com/
- Pandaily, “Bilibili Open-Sources Index-Translate, a Qwen3.5-Based Translation Model Family for 150 Languages.” https://pandaily.com/bilibili-index-translate-open-source-qwen3-5-150-languages
- Technical report, arXiv 2609.40181 (identifier as cited on the model card; abstract not retrieved in this run).
By Riju – about
