MiniMax M3.1 Flash Preview Explained: Architecture, Pricing and Benchmarks
Most model write-ups on launch week are built on a spec sheet. MiniMax M3.1 Flash Preview, which MiniMax confirmed on September 27, 2026, does not have one. There is no public model card, no parameter count, no per-token price, no benchmark table, and no open weights. What exists is a short vendor description, a handful of API documentation pages, and a flood of third-party articles that mostly repeat each other and, in a few cases, assert numbers that the vendor has never published.
That makes this an unusual deep-dive, and arguably a more useful one. Engineers deciding whether to route real work to a preview model need to know exactly where the documented ground ends. This post separates what MiniMax has actually stated from what is inferred and what is simply circulating, explains the mechanisms that are documented (one-million-token context, five-level effort control, always-on thinking, strict reasoning-block handling), and gives you an evaluation protocol that does not depend on numbers nobody has published.
What this covers: the verified facts and the verified gaps, how the preview differs from M3, how access and quotas really work, what the effort control means for cost, the tool-use contract you must respect, honest failure modes, and a decision matrix against M3 and other options.
Context and Background
MiniMax is a Shanghai-based lab that has spent 2026 pushing hard on agentic coding models. In June it released M3, an open-weight model that we analyzed in detail in our MiniMax M3 open-weight LLM benchmark analysis. That post covered a Mixture-of-Experts design reported at roughly 428 billion total parameters with about 23 billion active per token, a sparse attention scheme MiniMax calls MSA, and a headline 59.0 on SWE-Bench Pro as published by MiniMax. Those M3 figures came from the vendor’s launch materials, and we flagged then that they were not independently reproduced. They belong to M3. None of them has been published for M3.1 Flash Preview, and they should not be transplanted.
The preview arrived roughly four months later, with a different positioning. MiniMax describes it in its API documentation as a frontier multimodal coding model with a one-million-token context window and tunable thinking depth, and the documentation says it is built for agentic reasoning, tool use, coding and structured task execution. The “Flash” label signals the intent: a faster, more practical model for everyday development rather than a research flagship. The company has pointed the model at the loop of locating a bug, implementing a fix, running tests and delivering the change.
The industry context matters for how to read this release. Frontier labs now ship coding models as products first and as APIs later, and they increasingly expose a “reasoning effort” dial so the buyer can trade latency and tokens for deliberation. The general mechanics of that trade are covered in our AI inference cost optimization analysis of inference cost levers, which is the right companion read if you are budgeting for any effort-tiered model. M3.1 Flash Preview fits this pattern exactly, and, like several recent launches, it ships to subscribers before it ships to the open market.
For readers working in industrial IoT, digital twin and PLM engineering, the relevance is practical. Coding agents are now used to generate OPC UA client code, refactor simulation harnesses, and maintain data-pipeline glue. A one-million-token window is attractive for repositories with large vendor SDKs. But a preview model with no published price, no SLA and no independent benchmark is not something to wire into a production toolchain without a controlled evaluation, and the rest of this post is organized around making that evaluation cheap and honest.
One more framing point. The primary sources for this post are MiniMax’s own API documentation (the text generation guide and the Anthropic-compatible API reference) and its Token Plan pricing page. Secondary sources, such as news and review sites, are used only where they agree with those pages, and where a secondary source asserts something the vendor pages do not, we say so explicitly. The full list of what was checked appears in the review log that accompanies this post.
What MiniMax M3.1 Flash Preview Actually Is
MiniMax M3.1 Flash Preview is a preview-stage multimodal reasoning model that accepts text, image and video input, offers a one-million-token context window, always performs internal thinking at one of five selectable effort levels, and is currently available only through MiniMax Code and the subscription Token Plan, not through pay-as-you-go API billing.
That sentence is the entire documented specification. Everything else in this section is about reading it carefully, because each clause has consequences.
The documented facts, and the documented gaps
MiniMax’s text generation documentation lists the model identifier as MiniMax-M3.1-Flash-Preview, states a context window of 1,000,000 tokens, and says the model supports multimodal inputs including text, images and video. It also states that the model is available only through the plan product and MiniMax Code “for now”. The Anthropic-compatible API reference adds the thinking contract: effort is tuned through output_config.effort with values low, medium, high, xhigh and max, and the model “always thinks” and returns an HTTP 400 error if thinking is disabled.
Now the gaps. As of this writing, a search of the vendor documentation and of the independent coverage turns up no parameter count, no statement of whether the model is dense or Mixture-of-Experts, no attention mechanism, no tokenizer details, no training data description, no benchmark scores, no latency or throughput figures, no per-token pricing, and no license or weights release. Reviewers who looked closely said the same thing: one reviewer explicitly noted that MiniMax had not published a standalone benchmark table for M3.1 that supports a numerical coding score, and another noted the model cannot currently be independently benchmarked at scale because of the access restrictions.

Figure 1: The evidence map for MiniMax M3.1 Flash Preview. The left branch is what the vendor documents. The right branch is what remains unpublished at the time of writing.
Figure 1 is the mental model to keep. It is tempting to fill the right-hand branch by analogy with M3, but that is exactly the error to avoid. A model released four months after M3 under a different name could share its backbone, be a distillation of it, be a smaller sibling, or be something else entirely. MiniMax has not said. Any article that states an active-parameter count or an expert count for the preview is guessing, and a guess presented as a spec is how bad capacity plans get made.
The unverified numbers in circulation
At least one aggregator-style report we reviewed lists a precise set of figures for the preview: a SWE-bench Verified score in the low seventies, a generation speed of 165 tokens per second, a 145 millisecond time to first token, a very low per-million-token price with batch and cached tiers, a sparse expert count, and a three-billion-parameter speculative drafter. We could not find any of these in MiniMax’s documentation. They contradict the documentation on access (that report claims a pay-as-you-go platform API, whereas the vendor documentation says plan and MiniMax Code only), and a separate API-focused outlet states that no per-token price, cache rate, stable API identifier or benchmark has been published. We treat all of those numbers as unverified and exclude them from this post.
This is worth dwelling on, because search results for a new model name are now dominated by pages that synthesize plausible specs. The practical rule is simple: a spec is real when it appears on the vendor’s documentation, model card or pricing page. When it appears only in a summary article, it is a claim about a claim.
Preview means the contract can move
The word “Preview” is in the model identifier itself. That has three concrete implications. The behavior of the model can change without a version bump you can pin. The identifier may be retired or renamed when a general-availability release arrives. And one API aggregator specifically warns that the label shown in the MiniMax Code interface is not, by itself, a verified API request identifier, so teams should take the identifier from the documentation rather than from a screenshot. If you build against a preview, build the model name and the effort level into configuration, not into code.
Core Architecture: What We Can and Cannot Say
The architecture of M3.1 Flash Preview is not disclosed. What is disclosed is the interface contract: a one-million-token window, three input modalities, five effort levels, mandatory thinking and strict reasoning-block continuity. Those interface facts are enough to design an integration and an evaluation, even without knowing parameter counts.
One million tokens: capacity, not fidelity
A one-million-token window is a capacity statement. It says how many tokens the API will accept in a request. It does not say how reliably the model uses information at position 700,000, and the vendor has published no long-context retrieval curve for the preview. For M3, our earlier analysis noted that the launch materials reported long-context evaluation only at 256K and that a showcased 24-hour optimization run was an existence proof rather than a systematic test. The same caution applies here with more force, because there is no benchmark at all.
To put one million tokens in engineering terms, a rough rule of thumb is that source code runs somewhere around three to four characters per token for typical tokenizers, so a million tokens is on the order of three to four million characters of code. That might hold a mid-sized service repository plus its dependencies’ headers, but not a large monorepo. Whether the preview’s tokenizer matches M3’s is undisclosed, so treat that as an illustrative estimate, not a spec. The useful habit is to measure your own repository with the provider’s token-counting endpoint before assuming it fits.
Cost and latency also scale with how much of the window you fill. Even without published prices, the structural point holds: prefill work grows with input length, and in an agent loop you re-send a growing history on every turn unless the provider caches it. Whether MiniMax offers prompt caching for this preview is not documented, so a multi-turn agent that fills 400K tokens of context may be re-processing that prefix repeatedly. That is the single largest unknown in any cost model for this model, and it is why we recommend measuring it directly.
Three modalities and what video input implies
The documentation says images and video are accepted as input. Output is text. For a coding model, image input is useful for screenshots of failing UIs, architecture diagrams and rendered charts, and video input is useful for screen recordings of a bug reproduction. Video is also the most expensive input type in nearly every multimodal system, because frames expand into many tokens. MiniMax has not published how it tokenizes video, how many frames it samples, or the maximum video length. Before relying on video, run a single representative clip and look at the usage counters the API returns.
In IoT and digital-twin work, a realistic use is feeding a screenshot of a dashboard anomaly plus the relevant telemetry-processing code and asking for a diagnosis. That is a text-plus-image workload and is well within what the documentation describes. A long screen recording of a SCADA session is a different matter, and should be treated as an experiment.
Mandatory thinking and the five effort levels
The most distinctive documented feature is the effort control. The model accepts low, medium, high, xhigh and max, with max as the default when the parameter is omitted. The model always thinks. Sending a disabled thinking configuration returns HTTP 400. Both an Anthropic-compatible shape (output_config.effort) and an OpenAI-compatible shape (reasoning_effort) are documented in third-party coverage of the API, and the Anthropic-compatible page confirms the first.
The default matters more than it looks. Because omitting the parameter selects the maximum depth, an integration that never sets effort gets the slowest, most token-hungry behavior. A team that benchmarks “the model” without setting effort is benchmarking max, and a team that later sets low for latency is running a materially different configuration. Always log the effort value alongside every evaluation result.

Figure 2: Access paths. The preview is reachable through MiniMax Code and a Token Plan key, governed by rolling quota windows, while M3 remains on per-token billing.
Figure 2 shows how the routes divide. The practical consequence is that for the preview, your cost is a subscription and a quota, not a token meter, and that changes how you reason about budgets.
How M3.1 Flash Preview differs from M3
The two models share a family name and a one-million-token context window, and differ in everything else that is documented. M3 is a callable pay-as-you-go model with a published identifier (MiniMax-M3 in the text generation guide), per-token pricing reported by third parties, and open weights from its June release. M3.1 Flash Preview has a preview identifier, plan-only access, an effort parameter, mandatory thinking, and explicitly documented multimodal input. The text generation guide lists an effort parameter only for the preview.
| Dimension | MiniMax M3 | MiniMax M3.1 Flash Preview |
|---|---|---|
| Status | Released June 2026 | Preview, confirmed September 27, 2026 |
| Context window | 1,000,000 tokens | 1,000,000 tokens |
| Effort control | Not documented in the guide | low, medium, high, xhigh, max (default max) |
| Thinking | Not documented as mandatory | Always on, disabling returns HTTP 400 |
| Access | Pay-as-you-go routes | MiniMax Code and Token Plan only, for now |
| Per-token price | Reported by third parties | Not published |
| Weights | Open weights (see our M3 post) | None announced |
| Benchmarks | Vendor-reported, see M3 post | None published |
Notice what the table does not contain: a row for parameters. For M3 we have a reported figure and for the preview we have nothing. That asymmetry is the whole story of this release.
Deeper Analysis: Operating a Preview Model Without a Spec Sheet
If you cannot read capacity off a datasheet, you derive it from the interface and from measurement. Four areas deserve attention: how plan-based access shapes cost, how the effort dial shapes both cost and quality, what the tool-use contract demands of your harness, and how to build an evaluation that survives the model graduating from preview.
Plan-based access changes the economics
MiniMax’s Token Plan pricing page lists three tiers: Plus at $22 per month, Max at $55 per month and Ultra at $132 per month, with usage governed by rolling five-hour and weekly windows and agent-concurrency allowances that grow with tier. The same page also lists credit top-ups, in the range from five dollars for 5,000 credits up to one hundred dollars for 100,000 credits. The page, at the time we read it, did not list the preview by name, which is a reminder to check the live page rather than rely on a blog’s summary. The vendor’s text generation guide is the document that says the preview is reachable through the plan.
For a budget owner, this has a counterintuitive effect. A fixed subscription turns marginal tokens into free tokens until the quota binds, so the incentive is to use max effort and long contexts liberally. The binding constraint is the rolling window. Heavy agent loops that fill the context on every turn can exhaust a five-hour window quickly, and the failure shows up as throttling at an inconvenient moment rather than as a surprising invoice.
The correct metric for comparison is therefore not dollars per million tokens, which does not exist for this model, but dollars per accepted change within your quota. If a Plus subscription lets one developer complete a fixed number of agent tasks per day and a Max subscription lets them complete many more, the relevant comparison is subscription price divided by accepted tasks, set against the same figure for a pay-as-you-go model. Our AI inference cost optimization guide walks through the general approach of converting per-token prices into cost per solved task, and the same arithmetic applies in reverse here.
MiniMax also ran a promotion that doubled daily free quota for the preview and for some other models between September 28 and October 7, 2026, according to one API-aggregator report that we could not confirm on the vendor pages. Even if accurate, a promotional quota tells you nothing about steady state. Do not size a team rollout on promotion-period behavior.
The effort dial is a cost, latency and quality lever, not a quality lever
Five effort levels give you five operating points, but the vendor has published none of the curves. Without a published curve, assume nothing about linearity. In reasoning models generally, the relationship between thinking budget and accuracy is concave: the first increments buy a lot and later increments buy little, and the relationship depends heavily on task type. Simple refactors and boilerplate generation saturate early. Multi-file debugging and algorithmic tasks keep improving longer. These are general patterns in the literature on test-time compute, not measured properties of this preview.
Here is an illustrative, clearly hypothetical worked example of why you should sweep. Suppose that on your private task set, low solves 52 of 100 tasks using an average of 4,000 output tokens, medium solves 61 using 7,000, high solves 67 using 12,000, xhigh solves 69 using 20,000 and max solves 70 using 32,000. Those numbers are invented for illustration. In that hypothetical, going from high to max buys three more solved tasks for 20,000 extra tokens per task, which is a large increase in latency and quota burn for a tiny gain. The shape, not the values, is the lesson, and your real curve will differ.
The default of max means that an unconfigured integration lands at the far right of that curve. A sensible practice is to choose a different default per workload: low or medium for autocomplete-style and boilerplate tasks, high for typical bug-fix work, and max reserved for cases a first pass has already failed. That mirrors the escalation pattern many teams already use with other effort-tiered models, and it is also how the AI agent framework benchmark comparisons are best read: harness behavior and retry policy move results as much as the model does.
The tool-use contract: preserve the whole assistant message
The documentation is explicit about multi-turn tool use. When a response includes thinking blocks, you must preserve them unchanged in later turns, and in multi-turn function-calling conversations the complete assistant message must be appended to the conversation history to maintain continuity of the reasoning chain.

Figure 3: The agent loop. The harness must send back the full assistant message, including thinking blocks, with each tool result.
This is a real integration trap. Many harnesses normalize messages into a simple role-and-text structure and discard content block types they do not recognize. A harness that strips the thinking block before sending the next turn will not necessarily crash. It will more likely produce subtly degraded behavior: the model loses its own prior reasoning and may repeat work, contradict an earlier plan or abandon a multi-step approach. That failure mode looks like model weakness in your evaluation and is actually a harness bug.
A minimal sketch of the correct pattern follows. This is illustrative pseudocode in the Anthropic-compatible shape, not copied from vendor samples, and field names should be checked against the current reference.
messages = [{"role": "user", "content": task}]
while True:
resp = client.messages.create(
model="MiniMax-M3.1-Flash-Preview",
max_tokens=8192,
output_config={"effort": "high"}, # set explicitly, default is max
tools=tools,
messages=messages,
)
# Append the COMPLETE assistant content, thinking blocks included.
messages.append({"role": "assistant", "content": resp.content})
calls = [b for b in resp.content if b.type == "tool_use"]
if not calls:
break
results = [run_tool(c) for c in calls]
messages.append({"role": "user", "content": results})
Three details deserve emphasis. The effort value is explicit, so evaluations are reproducible. The assistant message is appended as a whole, not reconstructed. And the loop has no code path that disables thinking, since that returns an error. If you use a framework rather than a raw client, test this behavior directly by logging the exact payload of the second request, because abstractions sometimes drop unknown block types.
Designing an evaluation that outlives the preview
Because no independent benchmark exists, your private evaluation is the only evidence available. Build it so it survives the model changing underneath you. Pin a set of real tasks from your own repositories, each with an automated pass criterion such as a test suite. Record, for every run, the model identifier, the effort level, the date, the harness version, wall-clock time, tool calls, retries and token counts from the usage fields. Keep the raw transcripts. When the preview changes or reaches general availability, re-run the same set and diff the results.
A good task set for an IoT or PLM team mixes several categories: protocol client generation (for example an MQTT or OPC UA subscriber with reconnection logic), data-model transformations between a PLM export and a digital-twin schema, test-writing for existing telemetry code, and debugging tasks seeded with a known fault. Include some tasks that require reading a large amount of context, so that the one-million-token claim gets exercised rather than assumed. The evaluation should report the fraction of tasks solved, the median and 90th percentile time to completion, and the cost proxy for your access path.

Figure 4: A decision workflow. A published API or SLA requirement rules the preview out; otherwise sweep effort, measure cost per solved task, and adopt behind a flag.
Figure 4 encodes the decision. The first gate is blunt on purpose: if your use case needs a contractual availability guarantee or pay-as-you-go access today, this model is not an option regardless of quality, and M3 or another generally available model is the answer. If it passes that gate, the sweep tells you which effort level to deploy and whether the economics beat your current baseline.
A note on benchmark hygiene. When MiniMax or a third party eventually publishes scores for this model, ask the standard questions: which harness and scaffolding were used, which effort level, how many attempts, whether the benchmark has known contamination, and whether the comparison models were current at the time. Our M3 analysis noted that agentic scores depend heavily on scaffolding, with a comparison model already superseded on launch day. The same questions apply here. A score without an effort level is, for this model, an incomplete score.
How It Compares: A Decision Matrix
Because the preview has no benchmarks, a comparison cannot rank it on quality. It can only compare access, controllability and risk, which is what actually determines whether you can use it this quarter. The matrix below compares the preview to M3 and to a generic “generally available hosted coding model” category rather than to named competitors, since we have no verified head-to-head data to cite.
| Use case | M3.1 Flash Preview | MiniMax M3 | Generic GA hosted model |
|---|---|---|---|
| Developer-seat coding assistant, individual use | Good fit if you accept preview risk and use MiniMax Code | Good fit via API | Good fit |
| Production agent with SLA and per-token billing | Not currently possible | Possible, check terms | Best fit |
| Self-hosting and data residency | Not possible, no weights | Possible, open weights (large cluster) | Varies |
| Latency-sensitive interactive use | Unknown, test low and medium |
Unknown, test | Published figures usually exist |
| Reproducible research or audit trail | Weak, preview can change | Better, versioned and open weights | Better, versioned |
Read the matrix as a risk map, not a score. The preview’s strengths are its documented controllability (effort, thinking) and its one-million-token window. Its weaknesses are all about unpublished information and restricted access.
Trade-offs, Gotchas, and What Goes Wrong
No spec means no capacity planning. With no parameter count and no throughput numbers, you cannot estimate latency under load from first principles. You are limited to measuring it, and a measurement taken in the promotional period or at off-peak hours may not reflect later behavior. If you are going to run a pilot, run it over several days and at different times of day.
Quota exhaustion looks like an outage. Under plan-based access, a fleet of agents sharing a subscription can hit the five-hour or weekly window and begin to fail or throttle. That is a different failure mode than a per-token bill, and it needs monitoring. Track remaining quota where the product exposes it, and make your harness degrade gracefully, for instance by falling back to a lower effort or a different model when throttled.
Default max effort inflates cost and latency. Forgetting to set effort is the most likely self-inflicted wound. It will not error. It will simply be slow and consume quota faster than expected, and it will make the “Flash” label look unearned in your own tests. The label describes positioning, and MiniMax has published no measurement that supports a speed claim.
Dropped thinking blocks degrade multi-step work silently. As discussed, harnesses that rebuild messages from text alone can lose the reasoning chain. The symptom is an agent that seems to forget its plan between tool calls. Diff the outgoing request against the previous response to confirm thinking blocks are present and unmodified.
Long-context claims are untested for this model. A one-million-token window invites stuffing the entire repository into every request. That is expensive, slow, and, for sparse-attention or retrieval-style architectures in general, can reduce accuracy on details buried in the middle. Even if the preview handles it well, retrieval of just the relevant files is usually cheaper and often more accurate. Test both: whole-repo context versus targeted context, on your own tasks.
Third-party numbers contaminate planning. The most dangerous trap here is not a model weakness but an information one. A procurement or architecture document that quotes a specific tokens-per-second or price from an aggregator article is quoting something unverified. Insist that every figure in a design doc carry a vendor-documentation link.
Preview-to-GA drift. When general availability arrives, identifiers, pricing, quotas and even behavior can change. Anything you hard-code now is technical debt. Keep model name, effort and endpoint in configuration, keep your evaluation set, and plan to re-run it.
Data handling is undocumented in the sources we reviewed. For regulated or proprietary code, check MiniMax’s terms for data retention and training use before sending source to any preview. We did not find a statement specific to this preview, so we make no claim either way.
Practical Recommendations
Treat the model as a controlled experiment rather than an infrastructure dependency. The right posture for most teams is a time-boxed pilot with explicit exit criteria, run by one or two engineers, on non-sensitive repositories, with a fallback model already in place.
Start by reading the live vendor documentation rather than any summary, including this one, since a preview’s facts can change. Pull the model identifier from the documentation, put it in configuration, and set effort explicitly in every call. Run the sweep described above on a private task set and compute cost per accepted change under the actual quota. Decide up front what result would make you adopt, wait or abandon, and write it down before you look at the numbers, which prevents rationalizing a disappointing outcome after the fact.
For teams in industrial IoT, digital twin and PLM contexts, keep proprietary schemas, customer telemetry and unreleased product data out of the pilot until the data-handling terms are confirmed. Use public or synthetic fixtures for the evaluation. If you need self-hosting for residency reasons, the preview cannot help today, and the relevant question is whether M3’s open weights meet your needs. Our guide to expert-parallel MoE inference serving covers what serving a large MoE model of that class involves.
A short checklist:
- [ ] Take the model identifier and effort values from the vendor documentation, not from a UI label or article.
- [ ] Set
output_config.effortexplicitly; never rely on themaxdefault. - [ ] Verify your harness sends back complete assistant messages with thinking blocks unchanged.
- [ ] Log model, effort, date, harness version, latency, tool calls, retries and token usage for every run.
- [ ] Measure cost per accepted change under your real quota, not dollars per token.
- [ ] Test whole-context versus retrieved-context prompts on your own repositories.
- [ ] Keep sensitive code out until data-handling terms are confirmed.
- [ ] Keep a generally available fallback model wired in, and re-run the evaluation at GA.
Frequently Asked Questions
What is MiniMax M3.1 Flash Preview?
It is a preview multimodal coding and reasoning model from MiniMax, confirmed on September 27, 2026. The vendor documentation describes a one-million-token context window, text, image and video input, five thinking-effort levels and mandatory thinking. It is available through MiniMax Code and the Token Plan, not pay-as-you-go billing. Architecture, parameter count, benchmarks and per-token pricing have not been published.
Is MiniMax M3.1 Flash open source or open weight?
No weights or license have been announced for the preview, and coverage we reviewed states that none exists so far. This differs from M3, which MiniMax released as an open-weight model in June 2026. If you need to self-host, the preview is not an option today, and you should evaluate M3 instead. Check the vendor pages again later, since a preview’s status can change at general availability.
How much does MiniMax M3.1 Flash cost?
No per-token price has been published. MiniMax’s Token Plan page lists Plus at $22, Max at $55 and Ultra at $132 per month, with usage governed by rolling five-hour and weekly windows, and the text generation guide says the preview is available through the plan product and MiniMax Code. Some third-party sites list far lower per-token prices, but we could not find them on vendor pages, so treat them as unverified.
How do the effort levels work?
You set output_config.effort (Anthropic-compatible) to low, medium, high, xhigh or max, and the default when omitted is max. Higher levels let the model think longer, which generally raises latency and token use. The model always thinks, and sending a disabled thinking setting returns an HTTP 400 error. MiniMax has not published accuracy or latency curves per level, so sweep them on your own tasks.
Does MiniMax M3.1 Flash have published benchmark scores?
No. As of this writing MiniMax has not published a benchmark table for the preview, and reviewers specifically warn against copying M3’s scores, such as its vendor-reported SWE-Bench Pro result, onto it. Some aggregator pages quote precise coding and speed figures, but we found no vendor source for them. Build a private evaluation on your own repositories and record the effort level with every result.
Why does the model need thinking blocks preserved in tool use?
MiniMax’s documentation says the complete assistant message, including thinking blocks, must be appended to the conversation history in multi-turn function-calling so the reasoning chain stays continuous. If your harness strips those blocks, the model can lose its earlier plan and repeat or contradict itself. Log the second request in any agent loop and confirm the thinking content arrives unchanged.
Further Reading
- MiniMax M3 open-weight LLM benchmark analysis — the previous model in this series, with architecture and benchmark caveats.
- Expert-parallel MoE inference serving architecture — what serving a large Mixture-of-Experts model involves.
- AI agent framework benchmark: LangGraph, OpenAI and Google ADK — how harness choices move agent results.
- AI inference cost optimization — turning prices and effort levels into cost per solved task.
- MiniMax API documentation: text generation — the primary source for model identifiers, context and effort values.
- MiniMax API reference: Anthropic-compatible API — thinking-block and effort behavior.
- MiniMax Token Plan pricing — plan tiers and quota windows.
By Riju — about
