Ling 3.1 Flash Explained: Architecture, Pricing and Benchmarks
Most model write-ups in the first week after a launch repeat the press release. This one cannot, because Ling 3.1 Flash shipped on 30 September 2026 as an announcement, a free trial and a handful of vendor-reported scores, with no weights, no model card and no published price. Ant Group’s inclusionAI team describes a roughly 560 billion parameter mixture-of-experts model that activates about 25 billion parameters per token. Beyond those headline numbers, almost everything an engineer needs to plan a deployment is either undisclosed or contradicted between sources.
That gap matters now because the model is free only until 13 October 2026, and teams are already deciding whether to route traffic to it. This post separates what is documented from what is rumoured, derives what the documented numbers imply for memory and compute, compares the model with its open predecessor Ling 3.0 Flash, and gives you a procedure for evaluating a closed-weight LLM you cannot yet inspect.
What this covers: the Ling lineage, the documented specifications and where sources disagree, what a 560B/25B sparse design implies in practice, the benchmark evidence and its limits, access and pricing, failure modes, and a decision procedure for adopting it.
Context and Background
The Ling family comes from inclusionAI, the open-source arm of Ant Group, and its models have followed a consistent theme: very sparse mixture-of-experts (MoE) designs that trade total parameter count for low per-token compute. Ling-flash-2.0, for example, was published with 100 billion total and 6.1 billion active parameters, a 1/32 expert activation ratio, a 32K native context extendable to 128K with YaRN, and an MIT license, according to its Hugging Face model card. The point of that design was throughput: the card claims performance comparable to roughly 40 billion parameter dense models at over 200 tokens per second on H20 hardware, a claim you should treat as the vendor’s own.
Ling 3.0 Flash, announced on 26 July 2026 according to the Business Wire release, moved to a native hybrid linear attention design and a sparser 1/64 expert ratio. It has 124 billion total parameters, 5.1 billion active, a 256K training context and MIT-licensed weights on Hugging Face. That release is the reference point for everything below, because it is the last Ling model you can download and inspect.
Ling 3.1 Flash is a different kind of release. It is roughly 4.5 times larger in total parameters and about 4.9 times larger in active parameters than Ling 3.0 Flash. It is also described as a hybrid reasoning model, meaning it exposes reasoning effort levels rather than a single fixed mode. In vocabulary terms the “Flash” label now covers a much heavier model than it did two months ago, so do not assume the cost and latency profile of earlier Flash models carries over.
The market context is a crowded, fast-moving tier of cheap agent-oriented models. If you are tracking this tier, our earlier explainers on DeepSeek V4.1 Flash and Qwen3.8 Omni Flash cover the same question for other vendors: how much capability can you buy per active parameter. For the cost side of that question in general, see our guide to AI inference cost optimization.
What Is Actually Documented About Ling 3.1 Flash
Direct answer: Ling 3.1 Flash is a hybrid reasoning mixture-of-experts language model from inclusionAI (Ant Group) with about 560 billion total and 25 billion active parameters, announced on 30 September 2026. It is served with a 262,144 token context window and a 32,768 token output cap during a free trial ending 13 October 2026. Weights, license and final pricing are not published.

Figure 1: The evidence ladder for Ling 3.1 Flash. The upper rungs exist today; the lower rungs, open weights and reproducible evaluation, do not yet.
Figure 1 frames the situation. The announcement and the provider listings are real and consistent on the core specifications. The benchmark claims are first-party. The step that converts claims into evidence, an independent party running the model on its own tests with controlled settings, depends on weights or a stable paid API, and neither exists yet. The rest of this section walks the ladder from the top.
The specifications that sources agree on
Four numbers recur across every source I checked. The model has about 560 billion total parameters. About 25 billion are activated per token. The model is described as a mixture-of-experts system with hybrid reasoning. And the developer is inclusionAI, the Ant Group unit. TechNode’s report, citing IT Home, positions it for agent tasks, search, office software and specialist applications. The OpenRouter listing repeats “25B active parameters out of 560B total”, and the Vercel AI Gateway changelog says the same.
The activation ratio follows directly: 25 divided by 560 is about 4.5 percent, or roughly one parameter in 22 touched per token. Because both figures are stated with a tilde, treat that ratio as approximate. It is notably denser than Ling 3.0 Flash, which activates 5.1 of 124 billion parameters, about 4.1 percent, and which the model card describes as a 1/64 expert ratio.
The context window discrepancy
The announcement says the model supports up to 1 million tokens. The trial is capped lower: TechNode reports a 256,000 token limit, and the OpenRouter, Vercel and llmreference listings show 262,144 tokens, which is 256 times 1,024 and almost certainly the same limit expressed in binary units. The maximum output is listed as 32,768 tokens on those same pages.
For planning, use 262,144 input tokens and 32,768 output tokens, because that is what you can call today. The 1 million token figure is a vendor statement about capability that, per the reporting, will be enabled later. Nobody outside the vendor has published a long-context evaluation at that length, so do not budget for it.
Where the sources disagree
Several secondary details do not reconcile, and a careful reader should know which are noise and which matter. The table below lists them.
| Item | What sources say | How to treat it |
|---|---|---|
| Release date | 29 Sep (Vercel model page, one review blog), 30 Sep (TechNode, Vercel changelog, llmreference), 1 Oct (Dataconomy), 2 Oct (OpenRouter, Kilo) | Announcement was 30 Sep; the others look like listing dates |
| Price | Free on OpenRouter, Vercel and Kilo; one aggregator shows $0.30 input and $0.90 output per million tokens | Free is the only price documented by providers; treat the paid figure as unconfirmed |
| Weights | Not published; vendor says open-sourcing is planned | Closed today |
| License | Not stated anywhere I could find | Unknown; do not assume MIT |
| Architecture | Only “MoE, hybrid reasoning” stated | Attention type, expert count, layers undisclosed |
| Latency | 2.3 to 3.1 s time to first token, 60 to 88 tokens per second across three listings | Provider and load dependent |
None of these discrepancies is surprising for a model that appeared on aggregators within hours of an announcement. They do mean that any single-source article, including this one, can be wrong on a detail, which is why I have named the sources inline.
Architecture: What a 560B Total, 25B Active Design Implies
Inclusion of the word “Flash” suggests a speed-first model, but the numbers say something more nuanced. A sparse MoE decouples two costs that dense models couple together. Memory footprint scales with total parameters, while per-token compute scales with active parameters. Ling 3.1 Flash therefore has the memory profile of a very large model and the compute profile of a mid-sized one. Everything in this section is derived from the two documented numbers plus general MoE mechanics, and I flag where Ling 3.1 specifics are simply not public.

Figure 2: Generic MoE routing. A router scores every expert per token, only the top k run, and every expert still has to be resident in memory. This is the general mechanism, not a disclosed Ling 3.1 layout.
Figure 2 is deliberately generic. inclusionAI has not published the expert count, the number of routed experts per token, the layer count or the attention design for Ling 3.1 Flash. What the figure captures is the property that determines your infrastructure bill: inactive experts cost memory but not compute.
Memory: the cost you pay regardless of sparsity
Weight storage is parameter count times bytes per parameter. At 560 billion parameters, 16-bit weights occupy about 1.12 terabytes, and 8-bit weights occupy about 560 gigabytes. These are my own derivations from the stated parameter count, ignoring embeddings overhead, quantization scale tensors, activations and the key-value cache. They are lower bounds on a real deployment.
As a calibration point, the Ling 3.0 Flash weights are listed at roughly 255 GB in BF16 and 128 GB in FP8 for 124 billion parameters, per the OrcaRouter comparison. That is consistent with two bytes and one byte per parameter, so the same arithmetic applied to 560 billion parameters is reasonable. In practice an FP8 deployment of Ling 3.1 Flash would need at least eight 80 GB accelerators for weights alone, and realistically more once you reserve cache space. That is an illustrative planning estimate, not a vendor requirement.
The consequence is blunt. Even after weights are released, self-hosting is a multi-GPU, likely multi-node proposition for most teams. The model’s economics favour a provider who amortises that fixed memory across many concurrent requests, which is exactly the situation the free trial on shared providers represents.
Compute: where the sparsity pays off
A transformer forward pass costs roughly two floating point operations per active parameter per token, ignoring attention over long contexts. At 25 billion active parameters that is about 50 GFLOPs per generated token, versus about 10 GFLOPs for Ling 3.0 Flash at 5.1 billion active. This is a first-order estimate I derived, not a measured figure.
Against a hypothetical dense 560 billion parameter model, which would need about 1.1 TFLOPs per token, the sparse design is roughly 22 times cheaper in arithmetic. Against its own predecessor it is about five times more expensive. If you are moving a workload from Ling 3.0 Flash, you should expect per-token cost and latency to rise unless the provider’s stack compensates, and you should expect the quality gain to be what justifies it.
The practical bottleneck at decode time is usually memory bandwidth and expert dispatch rather than raw arithmetic. In an expert-parallel deployment, each token is routed across devices, and the all-to-all communication can dominate. Our deep dive on expert-parallel MoE inference serving explains why very sparse models often show good aggregate throughput but uneven tail latency, which is relevant when reading the 60 to 88 tokens per second figures reported on provider pages.
What the Ling 3.0 Flash card tells us, and what it does not
Because Ling 3.1 Flash has no model card, the closest primary-source evidence about inclusionAI’s current design choices is the Ling 3.0 Flash card. It lists a hybrid linear architecture with 35 Kimi Delta Attention layers and 7 gated MLA layers in a 5:1 ratio, 2 dense layers, 512 routed experts with 8 activated plus 1 shared, 32 attention heads, a hidden size of 2,560, an expert intermediate size of 768, and a vocabulary of 157,184 tokens. The context schedule was 8K, then 32K, then 256K tokens.
Kimi Delta Attention is a linear attention variant: instead of attending over the full key-value history, it maintains a fixed-size recurrent state, which keeps long-context memory growth bounded. The gated MLA layers every sixth layer retain exact token-to-token lookups. Secondary summaries disagree on how to expand “MLA” (one expands it as Multi-head Latent Attention, another as something else), and the card itself does not spell it out, so I write MLA without expansion. The design logic is sound regardless: cheap recurrent layers for most depth, exact attention in a minority.
It would be tempting to assume Ling 3.1 Flash inherits this stack scaled up. That assumption is plausible, since the 1M token claim is easier to deliver with linear attention layers, but it is an inference. The Ling 3.1 materials say only “mixture of experts” and “hybrid reasoning”. If you see an article stating that Ling 3.1 Flash uses KDA and MLA in a 5:1 ratio, ask for the source, because I could not find one.
Hybrid reasoning as an interface contract
Vercel’s listing says the model supports reasoning at multiple effort levels, tool calling and long-context text processing, and Kilo’s listing adds structured outputs with JSON schema validation. A hybrid reasoning model lets the caller choose how many thinking tokens to spend. That changes the cost model: the same prompt can cost very different amounts depending on the effort setting, and the 32,768 token output cap may interact with reasoning tokens. I could not confirm how providers count them against that cap, so check before you set budgets.
This interface also changes evaluation. A benchmark score is only meaningful alongside the effort level used, and the vendor-reported figures do not state that setting. When you run your own tests, fix the effort level, record it, and test at least two levels so you know where the quality curve flattens.
Training and Post-Training: What Is Disclosed
For Ling 3.1 Flash, training details are essentially undisclosed. I found no published token count, no compute figure, no description of the pre-training data mix and no post-training recipe. Any article quoting a training-token number for this model is either inferring from an earlier Ling release or inventing it.
What can be said comes from the lineage. Ling-flash-2.0’s card reports over 20 trillion tokens of pre-training data, followed by supervised fine-tuning and multi-stage reinforcement learning. The Ling 3.0 Flash release says the model was trained on more than 10,000 interactive environments for agentic tasks across coding, general knowledge and deep research, and it highlights self-correction and long-horizon planning. The 3.1 positioning, agent tasks, search, office software and specialist applications, is consistent with continued investment in that environment-based reinforcement learning, but that is an interpretation of the marketing language, not a disclosed fact.
The benchmark selection offers a second, indirect hint. The vendor-reported scores emphasise work-style evaluations rather than classic exam benchmarks, which suggests the post-training targets professional task completion. That is a reading of what was chosen for display, and a vendor chooses what to display.
Capabilities and Benchmarks: Three Tiers of Evidence
Benchmark numbers for Ling 3.1 Flash fall into three tiers, and mixing them is the most common mistake in early coverage. Tier one is the vendor’s own announcement. Tier two is provider-page figures that are also vendor-supplied but republished. Tier three is independent measurement, currently limited to Artificial Analysis. I list each with its provenance, and I did not run any of these myself.
Tier one: vendor-reported scores
The announcement, as relayed by OrcaRouter’s analysis, cites three headline figures. The model is reported at 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional, the last evaluated in an undocumented “AQ environment”. The same analysis notes that the comparison baselines in the announcement, Claude Opus 5 and GPT-5.6 Sol, had already been superseded by newer versions at publication time.
I could not retrieve the original announcement text, so I rely on this secondary relay for those three numbers. I also could not find definitions of GDPVal-AA v2.1 scoring in the vendor material, and Elo on that scale is not interchangeable with a percentage score. A reader should therefore hold these figures loosely: they are what the vendor chose to show, against baselines the vendor chose, in environments the vendor controls.
Tier two: provider-page agentic and coding figures
A review on Build Fast with AI lists further provider-reported scores: AutomationBench 52.5 percent, SkillsBench 68.7, CyberGym 87.9, Finance Agent v2 57.9, DRACO 85.5, Terminal-Bench 4 at 40.4, SWE Atlas Codebase QnA 55.9 and HealthBench Professional 65.3. The review itself flags that these are provider-reported only. The small difference on HealthBench Professional, 65.3 versus 65.35, looks like rounding, which at least suggests the same underlying source.
| Benchmark | Reported value | Provenance |
|---|---|---|
| GDPVal-AA v2.1 | 1,673 Elo | Vendor announcement, via secondary relay |
| FrontierSWE | 75.16 | Vendor announcement, via secondary relay |
| HealthBench Professional | 65.35 | Vendor announcement, undocumented environment |
| AutomationBench | 52.5% | Provider page, via review blog |
| SkillsBench | 68.7 | Provider page, via review blog |
| CyberGym | 87.9 | Provider page, via review blog |
| Finance Agent v2 | 57.9 | Provider page, via review blog |
| DRACO | 85.5 | Provider page, via review blog |
| Terminal-Bench 4 | 40.4% | Provider page, via review blog |
| SWE Atlas Codebase QnA | 55.9% | Provider page, via review blog |
Notice what is missing: no MMLU-Pro, no GPQA Diamond, no LiveCodeBench, no AIME for this model. The vendor chose agentic and professional-task evaluations, many of which are new, have limited public leaderboards and are hard to cross-check. That is not evidence of weakness, but it does mean you cannot place the model on the familiar academic scales.
Tier three: independent measurement
The only independent numbers I found come from Artificial Analysis, shown on the OpenRouter listing: an Intelligence Index of 41.1, AA-LCR (a long-context reasoning evaluation) at 83.0 percent, SciCode at 54.1 percent, GDPval-AA at 56.1 percent, Humanity’s Last Exam at 39.4 percent and a non-hallucination rate of 62.1 percent. A Dataconomy summary of the same data ranks it 52nd out of 681 models on the Index, and prints the HLE score as “394%”, which is evidently a typo for 39.4.
These independent numbers are more useful than the vendor’s because they come from a fixed methodology applied across many models. Two cautions apply. First, the index version matters, so compare only against other models on the same version. Second, a 62.1 percent non-hallucination rate means roughly 38 percent of the time the model either answered wrongly or failed to abstain on that test’s questions, which deserves attention if you plan to use it for retrieval or reporting tasks.
Note the apparent tension between the two GDPval figures. The vendor’s 1,673 Elo is on an Elo scale for GDPVal-AA v2.1, while Artificial Analysis reports 56.1 percent on GDPval-AA. These are different scales and possibly different versions, so they cannot be compared directly, and I make no claim that they conflict.
Comparing with Ling 3.0 Flash
The OrcaRouter comparison reports an Artificial Analysis Intelligence Index of 20 for Ling 3.0 Flash, with generation speed of 325 tokens per second. Ling 3.0 Flash’s own card reports SWE-Bench Pro 56.6 percent, SWE-Bench Multilingual 72.4 percent, AIME 2026 93.2 percent and HMMT February 2026 at 87 percent, all vendor-reported.
If the index versions match, 41.1 against 20 would be a large quality step, and 60 to 88 tokens per second against 325 would be a large speed penalty. Both inferences are cautious: the two figures were gathered at different times, on different providers, and I cannot confirm they use the same index revision. What you can safely conclude is the direction of the trade, more capability and more latency, which is exactly what a fivefold increase in active parameters predicts.
Benchmark caveats you should apply to any new model
Contamination is the first concern. Newly published agentic benchmarks have short histories, and there is little public evidence on whether their tasks leaked into training data. Environment definition is the second: scores on “AQ environment” or tool-harness-specific suites depend on scaffolding that the vendor controls. Selection is the third: a vendor lists the evaluations it wins.
None of this means the numbers are wrong. It means they are marketing-grade until a third party reproduces them, which requires either open weights or a stable paid endpoint. The independent Artificial Analysis data is the only partial exception.
Access, Pricing and Deployment
You can use Ling 3.1 Flash today through hosted APIs, but only on promotional terms. The model is served by NovitaAI behind OpenRouter and Vercel AI Gateway, and it is reachable through Ant’s own Ling Studio interface. No provider lists a durable price, so any cost model you build now is provisional.

Figure 3: Release pattern. Ling 3.0 Flash moved from announcement to MIT weights in roughly two weeks according to OrcaRouter; Ling 3.1 Flash has followed the first steps and has a plan, not a delivery, for the last.
Pricing: free now, undefined later
On OpenRouter, Vercel AI Gateway and Kilo the listed price is zero for input and output. The Vercel changelog says the standard model ID, inclusionai/ling-3.1-flash, transitions to paid billing after the promotion ends on 13 October 2026, while the inclusionai/ling-3.1-flash-free variant simply stops serving. As of this writing, the post-promotion rate has not been published by the vendor, and one aggregator displays $0.30 per million input tokens and $0.90 per million output tokens, which I could not trace to a provider and therefore label unconfirmed.
For calibration, Ling 3.0 Flash was listed at $0.07 per million input tokens and $0.22 per million output tokens, per the OrcaRouter comparison. If the 3.1 price lands near the unconfirmed figure, that would be about four times the input price of its predecessor, broadly in line with the five-fold increase in active parameters. This is a plausibility check on an unverified number, not a forecast.
An illustrative workload shows why the uncertainty matters. A batch job consuming 100 million input tokens and producing 20 million output tokens would cost $48 at the unconfirmed $0.30/$0.90 rates, versus $11.40 at Ling 3.0 Flash’s rates, versus nothing during the free window. These are arithmetic illustrations on assumed prices, not quotes. The point is that a free trial tells you nothing about unit economics.
Open weights: promised, not delivered
The announcement says inclusionAI plans to open-source the model. TechNode reports that release would follow the trial, and the OrcaRouter analysis points out the precedent: Ling 3.0 Flash was announced in late July with no weights and its MIT-licensed weights appeared around two weeks later. That pattern is encouraging but not a guarantee, and a license for 3.1 has not been stated. Earlier Ling releases used MIT, so it is a reasonable expectation, but I would not write contracts around it.
Until weights arrive, there is no Hugging Face repository, no quantization ecosystem, no vLLM or SGLang recipe and no way to run the model in your own boundary. If data residency or air-gapped operation is a requirement, Ling 3.1 Flash is not an option today, and Ling 3.0 Flash is the closest usable family member.
Latency and throughput on provider pages
Provider-reported latency varies: about 2.3 seconds (P50) and 60 tokens per second on OpenRouter, about 2.4 seconds and 67 tokens per second on Vercel’s 24-hour window with an 88 percent cache hit rate, and roughly 3.1 seconds and 88 tokens per second in the review’s reading of the Vercel and Novita route. These are three snapshots of one shared free endpoint under unknown load. A free endpoint with a single provider also shows weaker availability, and OpenRouter lists about 94 percent over three days, which is not production-grade.
Treat all of this as indicative. Measure time to first token and tokens per second yourself, at your own prompt lengths, on your own region, at the hour you intend to run.
How Ling 3.1 Flash Compares and How to Evaluate It
The closed state of the model changes how you should compare it. A normal comparison puts benchmark tables side by side. Here the more useful comparison is by decision: which workloads can adopt a model that has no weights, a temporary price and vendor-only scores, and which cannot.
A decision matrix against the alternatives
The table compares Ling 3.1 Flash with its open predecessor and with the generic category of reasoning-first frontier APIs. I avoid naming specific competitor numbers I have not verified this run.
| Use case | Ling 3.1 Flash | Ling 3.0 Flash | Frontier reasoning API |
|---|---|---|---|
| Prototype agent on free budget | Good fit until 13 Oct | Good fit, cheap API or self-host | Costly for a prototype |
| Regulated or air-gapped deployment | Not possible today | Possible, MIT weights | Usually not possible |
| Long documents above 262K tokens | Not served yet | Up to 262K via YaRN | Varies by vendor |
| Stable production SLA | Not yet, single provider, ~94% availability reported | Better, weights allow multiple hosts | Strongest, with contracts |
| Cost-sensitive high volume | Unknown until price is set | Strong at $0.07/$0.22 | Expensive |
Read the matrix as a snapshot. After weights are published and a paid price exists, several cells will change, particularly the second and fourth rows. For the reasoning-versus-cost question at the frontier tier, our explainer on test-time compute scaling in OpenAI’s o3 lays out why spending more tokens at inference is the knob that a hybrid reasoning model like Ling 3.1 Flash also exposes.
A procedure for evaluating a model you cannot inspect

Figure 4: Adoption flow. Residency and context requirements gate the choice; everything else goes through a pilot on your own golden set before the price is known.
Figure 4 compresses the procedure. Start with the hard constraints, because they decide the question without any benchmark. If you need weights in your own boundary, you wait or use Ling 3.0 Flash. If you need contexts beyond 262,144 tokens, you wait for the vendor to enable the larger window.
If neither constraint applies, run a pilot. Build a golden set of 50 to 200 real tasks from your own workload, with expected outcomes you can score automatically or with a rubric. Run the model at two reasoning effort levels, record cost in tokens and latency, and repeat each task a few times because sampling variance on agentic tasks is large. Then run the same set on your incumbent model.
Agent workloads need extra care. Multi-step tool use compounds errors, so measure end-to-end task success rather than single-turn accuracy, and log every tool call. If you are choosing an orchestration layer for that pilot, our benchmark of agent frameworks covers LangGraph, the OpenAI agents SDK and Google ADK and how they behave under load.
Finally, decide before the trial ends. The free window closes on 13 October, and the standard model ID moves to paid billing at an unpublished rate. A pilot that ends with a recommendation to adopt “if the price is right” should include the break-even price at which the quality gain over Ling 3.0 Flash still pays for itself.
Trade-offs, Gotchas, and What Goes Wrong
The first risk is evidence quality. Every performance claim beyond the Artificial Analysis figures is vendor-supplied, and several rely on new benchmarks with undocumented environments. Treat a published Elo or percentage as a hypothesis to test, not a ranking to cite in a procurement document.
The second risk is lock-in by default. A free tier invites integration work, prompt tuning and evaluation harness building around one model’s quirks. If the paid price disappoints, you have sunk cost. Mitigate this by keeping prompts and tool schemas provider-neutral and by retaining a fallback route to a second model through a gateway.
The third risk is behavioural drift. A hosted, closed model can be updated without notice, and the two model IDs on Vercel already show promotional variants being retired. Pin model identifiers, store evaluation results with timestamps and rerun the golden set on a schedule.
The fourth is hallucination and abstention. A 62.1 percent non-hallucination rate on one independent test suggests you should not use the model for ungrounded factual output. Pair it with retrieval, ask for citations and verify them programmatically. Hybrid reasoning can also over-think simple requests, so route trivial queries to a smaller model.
Fifth, operational constraints on the free tier matter more than they appear. A single provider, about 94 percent availability over three days and shared capacity mean retries, timeouts and circuit breakers are mandatory, and you should not put it on a latency-critical path.
Sixth, the self-hosting story, when it arrives, will be expensive. As derived above, an FP8 copy needs about 560 GB for weights alone, so a single eight-GPU node with 80 GB per device is already tight once the key-value cache is included. Unless you have steady high utilisation, a hosted API will almost certainly beat your own cluster on cost, and the trade resembles the one we describe in our inference cost optimization guide.
Seventh, there is a communication risk inside your team. Early articles, including aggregator pages generated automatically, mix dates, prices and benchmark numbers from different sources. A claim such as “open-source 560B model” is wrong today. Write the facts you rely on into your design document with their source and date.
Finally, safety documentation is absent. I found no system card, no red-team results and no description of refusal behaviour for Ling 3.1 Flash. If your use case has safety or compliance requirements, you will have to run your own adversarial testing, including prompt injection tests for any tool-using agent.
Practical Recommendations
For most teams the right move this week is a bounded, time-boxed pilot rather than adoption. Use the free window to learn whether the model’s quality on your tasks justifies a price you do not yet know, and make the decision to continue conditional on that price.
If you are a prototype-stage team with no data residency constraints, use the free endpoint now, but route through a gateway that lets you switch models without code changes. If you are in a regulated or air-gapped environment, stay on Ling 3.0 Flash or another open-weights model and revisit when 3.1 weights and a license appear. If you need contexts above 262K tokens, do not plan on the 1M window until a provider serves it and a long-context evaluation exists.
A short checklist:
- Record the date, provider and model ID of every evaluation run, and pin them.
- Build a golden set from real tasks and score at two reasoning effort levels.
- Measure latency and throughput yourself, at your own prompt lengths.
- Compute a break-even price against Ling 3.0 Flash and against your incumbent.
- Add retries, timeouts and a fallback model, because free-tier availability is not an SLA.
- Do not cite vendor-only benchmarks externally without labelling them as such.
- Recheck this model’s status after 13 October for price, weights, license and a model card.
Revisit the decision when four things are public: a license, weights, a paid price and independent long-context results. Until then, Ling 3.1 Flash is best understood as a promising, expensive-to-host, thinly documented model that is cheap to try and expensive to depend on.
Frequently Asked Questions
What is Ling 3.1 Flash?
Ling 3.1 Flash is a hybrid reasoning mixture-of-experts language model from inclusionAI, the open-source arm of Ant Group, announced on 30 September 2026. It has about 560 billion total parameters and about 25 billion active per token. It is positioned for agent tasks, search, office software and specialist applications, and it is currently accessible through hosted APIs and Ant’s Ling Studio rather than as downloadable weights.
Is Ling 3.1 Flash open source?
Not yet. The vendor says it plans to open-source the model, and reporting indicates the release would follow the free trial that ends on 13 October 2026. As of this writing there are no published weights, no Hugging Face repository, no model card and no stated license. The predecessor, Ling 3.0 Flash, did receive MIT-licensed weights after a short delay, but that is a precedent rather than a commitment.
How much does Ling 3.1 Flash cost?
It is free on OpenRouter, Vercel AI Gateway and Kilo during the promotional period through 13 October 2026. The standard model ID on Vercel moves to paid billing afterwards at a rate that has not been published. One aggregator shows $0.30 per million input and $0.90 per million output tokens, but I could not confirm it with a provider, so treat it as unverified.
What is the context window of Ling 3.1 Flash?
The vendor says the model supports up to 1 million tokens, but the trial deployments serve 262,144 tokens, which TechNode reports as a 256,000 token limit. The maximum output listed on providers is 32,768 tokens. Plan around the served limits, because the larger window is described as something to be enabled later and has no public long-context evaluation yet.
How does Ling 3.1 Flash differ from Ling 3.0 Flash?
It is much larger: about 560 billion total and 25 billion active parameters versus 124 billion and 5.1 billion for Ling 3.0 Flash. Ling 3.0 Flash has MIT-licensed weights, a published hybrid linear attention design and a $0.07/$0.22 per million token price. Ling 3.1 Flash has no published weights, license, architecture details or final price, so a direct like-for-like comparison is not possible yet.
How good is Ling 3.1 Flash on benchmarks?
Independent Artificial Analysis figures show an Intelligence Index of 41.1, ranked 52nd of 681 models in one summary, with 54.1 percent on SciCode and 39.4 percent on Humanity’s Last Exam. The vendor reports 1,673 Elo on GDPVal-AA v2.1 and 75.16 on FrontierSWE, but those are first-party and unreproduced. Run your own evaluation on your workload before relying on either set.
Further Reading
- Expert-parallel MoE inference serving architecture for why very sparse models behave the way they do in production.
- OpenAI o3 and test-time compute scaling for the reasoning-effort trade-off.
- AI agent frameworks benchmark: LangGraph, OpenAI and Google ADK for building the agent harness you will evaluate with.
- AI inference cost optimization for modelling cost per task rather than per token.
- External: the Ling-3.0-flash model card on Hugging Face and the Vercel AI Gateway changelog for Ling 3.1 Flash.
By Riju — about
