AMD Ryzen AI Max+ Pro 400: Local LLM Inference on Workstations Explained

AMD Ryzen AI Max+ Pro 400: Local LLM Inference on Workstations Explained

AMD Ryzen AI Max+ Pro 400: Local LLM Inference on Workstations Explained

A workstation laptop with 192 GB of memory that the GPU can partly address sounds like a spec-sheet stunt, until you remember that the binding constraint on running large language models is rarely arithmetic. It is where the weights live and how fast they can be read. The Ryzen AI Max+ Pro 400 family, AMD’s refresh of its Strix Halo-class workstation processors, is a bet that a very large, moderately fast, shared memory pool beats a small, very fast, discrete one for a growing class of local inference jobs.

That matters now because open-weight mixture-of-experts models have made “big but sparse” the common shape of capable models, and enterprises increasingly want inference on-premises for privacy, latency and predictable cost. This article walks through what AMD has published about the chips, how the CPU, integrated GPU and NPU divide the work, which model sizes genuinely fit, and a first-principles estimate of tokens per second. You leave with a method you can reuse to judge any unified-memory machine, plus a break-even framework for local versus cloud API spending.

What this covers: the confirmed specifications and what remains unpublished, the unified memory architecture, the bandwidth-bound decode math with clearly labelled estimates, which models fit in 192 GB, where the NPU actually helps, local-versus-cloud economics, failure modes, and a practical buying and deployment checklist.

Context and Background

Local LLM inference has spent three years caught between two hardware camps. Discrete GPUs from NVIDIA offer enormous memory bandwidth, on the order of hundreds of gigabytes per second to multiple terabytes per second, but their video memory is small and expensive per gigabyte on the consumer and workstation tiers. Apple’s M-series chips took the opposite route, pairing a wide unified memory interface with large capacity, which is why a high-memory Mac became the default recommendation for people who wanted to load a 70-billion-parameter model on a desk. AMD’s Ryzen AI Max line, first sold as the Ryzen AI Max+ 395 with up to 128 GB, was the x86 answer to that design.

The 400 series extends that answer in two directions. Per AMD’s launch material as reported by ServeTheHome and AMD’s own blog, maximum memory rises from 128 GB to 192 GB, memory speed moves from LPDDR5X-8000 to LPDDR5X-8533, and bandwidth rises roughly seven percent to 273 GB/s. AMD says up to 160 GB of that pool can be allocated to the GPU. The company also frames the top part as the first x86 client processor able to run 300-billion-parameter models locally at 4-bit quantization, a claim we will test against simple arithmetic later rather than take on faith.

The workstation framing matters. The “Pro” suffix in AMD’s naming signals the commercial platform: manageability features, longer validation cycles and OEM workstation designs from partners AMD names as ASUS, HP and Lenovo. That is a different buyer from the enthusiast building a mini PC, and it changes the questions. IT departments ask about fleet management, data-residency policy and total cost of ownership, not only about tokens per second.

It also helps to place this class of device relative to the rest of the inference landscape. Datacenter accelerators such as the parts covered in our AMD MI400 Instinct architecture analysis sit at the other end of the spectrum, with HBM stacks delivering multiple terabytes per second. Embedded modules, surveyed in edge AI inference on NVIDIA Jetson, Intel Movidius and Arm NPUs, occupy the low-power tier. The Ryzen AI Max+ Pro 400 workstation tier sits in between: tens of watts to roughly 120 W, hundreds of gigabytes per second, and capacity that neither of its neighbours offers at the price point of a laptop or small desktop.

For the primary vendor material, see AMD’s own announcement, AMD Powers Next-Generation Agent Computers with New Ryzen AI Halo Developer Platform and Ryzen AI Max PRO 400 Series Processors, which is the source for the SKU table and software claims used below.

What is confirmed, and what is not

Before the analysis, a note on epistemic status, because launch coverage of new silicon is noisy. The facts in this article come from AMD’s blog and from ServeTheHome’s reporting of the launch. Where those two sources disagree, we say so. For example, cache totals are presented differently: ServeTheHome lists L3 figures of 64 MB for the top two SKUs and 32 MB for the smallest, while AMD’s table lists larger totals of 80, 76 and 40 MB, which is consistent with those being combined L2 plus L3 figures. We treat the cache numbers as unresolved and do not rely on them anywhere in the analysis.

Several things were not published in the sources we could verify: street pricing, independent benchmark results, sustained-power behaviour under long inference runs, and the exact shipping dates by OEM beyond a “Q3 2026” window. Any tokens-per-second figure in this article is therefore an estimate derived from first principles and is labelled as such. Treat vendor “up to” numbers, and ours, as ceilings rather than promises.

The Reference Architecture: One Pool, Three Engines

The Ryzen AI Max+ Pro 400 combines Zen 5 CPU cores, an RDNA 3.5 integrated GPU and an XDNA 2 NPU on one package that shares a single 256-bit LPDDR5X-8533 memory pool of up to 192 GB. AMD says the GPU can address up to 160 GB. Peak bandwidth is 273 GB/s, which, rather than compute, usually limits decode speed.

Ryzen AI Max+ Pro 400 unified memory architecture with CPU, iGPU and NPU sharing one LPDDR5X pool

Figure 1: The three compute engines share one fabric and one memory pool, so model weights are not copied between separate CPU and GPU memories.

The diagram shows the defining property of the design. There is no separate video memory. The CPU cores, the integrated Radeon GPU and the NPU all reach the same LPDDR5X through an on-die fabric. A model loaded once is visible to whichever engine runs a given layer or operator, with no PCIe transfer in the path. The operating system and applications take their share, and AMD’s “up to 160 GB” figure for GPU-visible memory means roughly 32 GB of a 192 GB machine remains for everything else.

The SKU lineup as published

Three processors were named in the launch material. The table below reproduces the figures that both sources agree on, plus the AMD-only details marked accordingly.

SKU Cores / threads Max boost Radeon iGPU GPU CUs NPU (up to) Memory
Ryzen AI Max+ PRO 495 16C / 32T 5.2 GHz 8065S 40 55 TOPS up to 192 GB
Ryzen AI Max PRO 490 12C / 24T 5.0 GHz 8050S 32 50 TOPS up to 192 GB
Ryzen AI Max PRO 485 8C / 16T 5.0 GHz 8050S 32 50 TOPS up to 192 GB

All three are configurable between 45 W and 120 W, per both sources. Memory is LPDDR5X-8533 across the line. Only the 495 carries the “Max+” designation, which is the focus keyphrase of this article; the other two are “Max” parts. The architectural building blocks are Zen 5 for the CPU, RDNA 3.5 for graphics and XDNA 2 for the NPU, as AMD describes them.

One implication is easy to miss. Because all three SKUs share the same memory subsystem, the bandwidth ceiling for token generation is the same on the cheapest and the most expensive part. The extra GPU compute units and CPU cores of the 495 help prompt processing, batching and mixed workloads. They do not raise the decode ceiling for a single stream of generated text. Buying the top SKU for single-user chat is mostly a purchase of headroom, not of speed.

Why unified memory changes the model-sizing question

On a discrete-GPU workstation the question is “does the model fit in VRAM?”, and the answer forces compromises: aggressive quantization, splitting layers between GPU and system memory, or buying a second card. Partial offload to system memory over PCIe is slow because the link and the DDR channels are far narrower than the GPU’s own memory bus.

On a unified-memory machine the question becomes “does the model plus its key-value cache fit under the GPU-addressable ceiling, and can the memory bus feed it fast enough?” The first half has a generous answer at 160 GB. The second half has a modest one at 273 GB/s. The whole design trades speed for capacity, and the honest way to evaluate it is to separate those two quantities rather than collapse them into a single “AI performance” number.

How the workload divides across CPU, iGPU and NPU

Inference has two phases with opposite characters. Prefill processes the whole prompt in parallel, performing large matrix-matrix multiplications that are limited by compute throughput. Decode generates one token at a time, performing matrix-vector multiplications that must stream the whole set of active weights from memory for every token, which makes it limited by bandwidth.

The integrated GPU is the main engine for both phases in popular runtimes such as llama.cpp and vLLM-style stacks on ROCm, because it has the matrix throughput for prefill and the widest path to memory for decode. The CPU handles tokenization, sampling, orchestration and any layers or operators the GPU backend does not cover. The NPU is a different tool, and we give it a section of its own below, because its role is the most commonly misunderstood part of the “AI PC” story.

Deeper Analysis: The Bandwidth Math Behind Local LLM Inference

This section derives what the platform can plausibly do, using only published memory figures and standard model arithmetic. Everything numeric here is an estimate, not a measurement. Independent benchmarks for the 400 series were not available when we wrote this, and we will not invent them.

Decode is a memory-streaming problem

During token generation, each new token requires reading every weight that participates in the forward pass. For a dense model that is the entire weight file. The throughput ceiling is therefore:

tokens_per_second_max = memory_bandwidth / bytes_read_per_token

With 273 GB/s and a 4-bit-class quantization averaging about 4.8 bits per weight (typical of llama.cpp Q4_K_M-style formats, which keep some tensors at higher precision), a dense 70-billion-parameter model weighs roughly 42 GB. The ceiling is 273 / 42, or about 6.5 tokens per second. Real systems do not reach the ceiling. Cache effects, kernel inefficiency, KV-cache reads and fabric contention typically leave a fraction of peak bandwidth usable. We assume 60 to 70 percent as a planning range, an assumption rather than a measured value, which puts the dense 70B case near 4 tokens per second.

Local LLM inference on Ryzen AI Max+ Pro 400 showing prefill compute-bound and decode bandwidth-bound phases

Figure 2: Prefill is limited by iGPU compute; decode loops over the weights and is limited by the 273 GB/s memory ceiling.

The diagram makes the asymmetry concrete. Prefill happens once per request and writes the KV cache. Decode repeats for every output token and re-reads the weights each time, so the bandwidth term dominates perceived responsiveness. A user waiting for a long answer feels decode; a user pasting a 50-page document feels prefill.

Estimated decode ceilings by model class

The table applies the formula to representative model shapes. Sizes are approximate and depend on the quantization format. “Planning” uses 65 percent of peak bandwidth, our assumption. Active-parameter counts for the mixture-of-experts examples come from the model publishers’ own descriptions as we understand them; verify against current model cards before relying on them.

Model shape Weights in memory (approx.) Bytes read per token (approx.) Ceiling at 273 GB/s Planning at 65%
8B dense, 4-bit class 4.5 GB 4.5 GB ~60 tok/s ~39 tok/s
32B dense, 4-bit class 19 GB 19 GB ~14 tok/s ~9 tok/s
70B dense, 4-bit class 42 GB 42 GB ~6.5 tok/s ~4 tok/s
70B dense, 8-bit 70 GB 70 GB ~3.9 tok/s ~2.5 tok/s
30B MoE, ~3B active, 4-bit class 18 GB ~2 GB ~137 tok/s ~89 tok/s
~117B MoE, ~5B active, ~4-bit 60-65 GB ~2.7 GB ~100 tok/s ~65 tok/s
300B dense, 4-bit 150 GB 150 GB ~1.8 tok/s ~1.2 tok/s

Two features of the table deserve attention. First, dense models above roughly 40 GB are usable for patient, document-style work but sluggish for conversation. Second, sparse models are the sweet spot. A mixture-of-experts model that stores 100 GB or more of weights but activates only a small fraction per token reads far fewer bytes, so it decodes at a speed closer to a small dense model while delivering the knowledge capacity of a large one. The MoE rows are optimistic in one respect: routing is data-dependent, and shared attention weights and the KV cache add reads that this simple model ignores, so treat those ceilings as upper bounds.

Why the 300-billion-parameter claim needs context

AMD’s claim that the top part can run 300-billion-parameter models locally at 4-bit quantization holds as a capacity statement. Three hundred billion parameters at 4 bits is about 150 GB of weights, which fits under the 160 GB GPU-addressable limit with only about 10 GB left for the KV cache and runtime buffers. That is tight, and a long context would not fit.

As a speed statement the claim is silent. A dense 300B model would decode at roughly 1 to 2 tokens per second by our arithmetic, which is fine for overnight batch jobs and poor for interactive use. If the model is a sparse mixture-of-experts with, hypothetically, 30 billion active parameters, the bytes read per token fall to about 15 GB and the ceiling rises to around 18 tokens per second. The architecture of the model, not the label “300B”, determines whether the experience is usable. We could not verify which specific 300B-class model AMD used in its demonstrations, and we flag that as unknown.

Prefill: where compute and the iGPU matter

Prefill cost scales with prompt length. A rule of thumb is that processing a prompt costs about two floating-point operations per parameter per token. An 8B model on an 8,192-token prompt therefore needs roughly 1.3 times ten to the fourteenth operations. If the iGPU sustained 20 teraflops of useful half-precision throughput, a purely illustrative figure rather than an AMD specification, prefill would take around 6.6 seconds. At 5 teraflops sustained, it would take 26 seconds.

This is why the 40-CU 495 and the 32-CU 490 and 485 differ meaningfully for retrieval-augmented generation and coding assistants that send long prompts, even though their decode ceilings are identical. If your workload is long-context summarization, the GPU compute gap between SKUs matters more than the headline memory number. If it is short-prompt chat, it matters much less.

The KV cache is the hidden memory consumer

The key-value cache grows with context length and competes with weights for the same pool. Its size per token is 2 × layers × KV heads × head dimension × bytes per element. For a 70B model with 80 layers, 8 grouped-query KV heads and a head dimension of 128 at 16-bit precision, that is 2 × 80 × 8 × 128 × 2 bytes, about 328 KB per token. At 32,000 tokens the cache is roughly 10.7 GB; at 128,000 tokens it is about 43 GB.

On a 192 GB machine with 160 GB visible to the GPU, a 42 GB 4-bit 70B model plus a 128k-token cache totals about 85 GB and fits comfortably. The same 128k context on a 300B dense model would not. KV-cache quantization to 8-bit halves those figures at some quality cost, and runtimes differ in how well they support it. For the broader serving-side picture of how caches and request scheduling interact, see our guide to continuous batching in LLM inference.

Decision flow for which models fit and run well on Ryzen AI Max+ Pro 400 dense versus mixture-of-experts

Figure 3: Dense and mixture-of-experts models read very different numbers of bytes per token, which decides whether a model that fits in 192 GB is also pleasant to use.

The decision flow reduces the sizing exercise to two questions. Does the model fit with its cache under the GPU ceiling? And how many bytes does it read per token? A model can pass the first test and fail the second, which is exactly the trap a capacity-only marketing claim invites.

Quantization choices on this platform

Quantization trades quality for memory and bandwidth, and on a bandwidth-bound platform the speed benefit is nearly proportional: halving bits per weight roughly doubles the decode ceiling. Moving a 70B model from 8-bit to a 4-bit class format raises the ceiling from about 3.9 to about 6.5 tokens per second in our table, a smaller ratio than the weight sizes alone suggest because the 4-bit class formats average closer to 4.8 bits than 4. Formats native to recent hardware, such as FP8 and the microscaling 4-bit formats, depend on kernel support in the ROCm stack for this specific GPU generation; we could not verify the status of each, so check the runtime release notes. Our FP8 versus INT8 versus INT4 quantization benchmark explains the quality side of the trade in detail.

Where the NPU Actually Fits

AMD quotes up to 55 TOPS for the 495 and up to 50 TOPS for the 490 and 485 from the XDNA 2 NPU. TOPS means trillions of operations per second, typically measured on 8-bit integer operations, though the precision behind AMD’s “up to” figure was not specified in the sources we checked. Microsoft’s Copilot+ PC programme has set a 40 TOPS NPU floor, so all three parts clear that bar, which is relevant to Windows-side features but says little about LLM serving.

The NPU’s real value is efficiency, not peak speed. It is designed to run fixed, quantized neural network graphs at low power, which suits always-on tasks such as background transcription, noise suppression, camera effects and small classification models. For these, running on the NPU leaves the iGPU free and keeps the fan quiet on battery. That is a legitimate benefit, but it is a different benefit from large-model throughput.

Why the NPU does not raise the decode ceiling

The NPU sits on the same fabric and reads the same LPDDR5X pool. If decode is bandwidth-bound, adding a second engine that draws from the same 273 GB/s cannot raise the ceiling. Two engines reading simultaneously split the bus; they do not add to it. This is the most important correction to the “50 TOPS AI PC” marketing frame: for large-model token generation, TOPS is the wrong unit, and gigabytes per second is the right one.

Where the NPU can help is prefill and power. Prefill is compute-bound, and an NPU can execute those matrix operations at lower energy per operation than a general-purpose GPU. AMD’s Ryzen AI software documents hybrid execution modes in which the NPU and iGPU cooperate on a model, as we understand the vendor’s documentation, using the NPU for the prompt phase and the iGPU for token generation. The practical conditions are that the model is converted into the supported format and operator set, and that the runtime supports your model architecture. We could not verify, for the 400 series specifically, which model families and sizes the hybrid path supports, so the safe planning assumption is that the headline large-model workflow runs on the iGPU through ROCm-backed llama.cpp, vLLM or similar, while the NPU handles smaller or specially prepared models.

A division of labour that works

A sensible configuration for a mixed workstation looks like this. The large chat or coding model runs on the iGPU with the weights resident in the shared pool. A small embedding or reranking model for retrieval runs on the NPU or CPU, so it does not compete with the main model for GPU time. Speech-to-text and camera pipelines run on the NPU. The CPU’s 8 to 16 Zen 5 cores handle tokenization, vector search, document parsing and the application itself. Memory budgeting is the discipline that holds this together, since every resident model draws on the same 192 GB.

The Software Stack: What Runs Today

AMD lists support for PyTorch, vLLM, llama.cpp, Ollama, ComfyUI and LM Studio, with ROCm optimization, on both Windows and Linux. That list covers the mainstream local-inference toolchain, and it is the single biggest practical difference between this platform and the NPU-first “AI PC” laptops of 2024 and 2025, which often required vendor-specific model conversion to do anything useful.

Choosing a runtime

llama.cpp and the tools built on it, such as Ollama and LM Studio, offer the broadest model coverage through GGUF files and run well on unified-memory systems. They are the easiest path to a working setup and handle partial quantization formats gracefully. vLLM targets serving: continuous batching, paged attention and an OpenAI-compatible API. It becomes relevant when several people or agents share one machine, because batching raises aggregate throughput when the bottleneck is bandwidth. The mechanism is that one weight read can serve several concurrent sequences, so aggregate tokens per second rises with batch size until compute or cache limits intervene. Our article on continuous batching covers the scheduling detail.

PyTorch matters for fine-tuning and research code, though full fine-tuning of large models is out of reach at this bandwidth and capacity. Parameter-efficient methods on small and mid-sized models are realistic. ComfyUI covers image and video generation workflows, a reminder that the same memory pool serves diffusion models, which are compute-heavier and less bandwidth-bound than LLM decode.

Memory allocation in practice

Unified memory is not automatically all available to the GPU. The “up to 160 GB” figure implies a configurable split, set through firmware options and driver behaviour that differ by OEM and operating system. Linux and Windows handle dynamic allocation differently, and some runtimes require the allocation to be raised before launch or fail with out-of-memory errors despite free system RAM. We could not verify the exact configuration procedure for the 400 series on shipping workstations, so test early with your target model rather than assuming the spec-sheet ceiling is reachable on your image.

Edge deployment parallels

Teams coming from embedded work will recognise the pattern. The same sizing discipline appears in our edge LLM benchmark on Jetson Orin: capacity decides what loads, bandwidth decides how fast it runs, and thermal limits decide how long it keeps running. For vendor-specific runtime and NPU tooling comparisons on the Intel side, see OpenVINO 2026.4 versus 2025.4 for edge LLMs and NPUs. If you are comparing approaches to shrinking models for such targets, the extreme end is covered in our look at ternary and 1-bit LLMs for edge inference, which attacks exactly the bandwidth term discussed here.

Local Versus Cloud API Economics

The question buyers actually face is not “is the chip fast?” but “does owning inference capacity beat renting it?” The answer depends on token volume, utilisation and what quality of model you need. We give a framework and a worked example with placeholder inputs that you should replace with real quotes, because workstation pricing for the 400 series had not been published in the sources we verified and cloud prices change frequently.

Break-even framework comparing local Ryzen AI Max+ Pro 400 workstation inference cost with cloud API per-token cost

Figure 4: Local cost is mostly fixed and cloud cost scales with volume, so a break-even monthly token volume separates the two regimes.

The break-even formula

Let H be the workstation’s purchase price, L its useful life in months, P the monthly power and maintenance cost, and V the monthly token volume in millions of tokens. Let C be the blended cloud price per million tokens for a comparable model. Local monthly cost is H / L + P. Cloud monthly cost is C × V. Break-even volume is V* = (H / L + P) / C.

For illustration only, suppose a workstation costs 4,000 dollars, lasts 36 months and costs 15 dollars per month in electricity at light duty. Local monthly cost is about 126 dollars. If the comparable API charges an effective 1 dollar per million blended tokens, break-even is about 126 million tokens per month. At 5 dollars per million, it falls to about 25 million. These inputs are hypothetical, not quotes. The point is the structure: the cheaper the API model you would otherwise use, the harder it is for local hardware to win on cost alone.

What the formula leaves out

Throughput caps the numerator. At a planning rate of 4 tokens per second, a dense 70B model generates at most about 10 million tokens in a 30-day month of continuous operation, which is below the break-even in the 1-dollar example. A 65-token-per-second MoE model could generate roughly 168 million tokens per month at full utilisation, but real utilisation is nowhere near continuous. So for the same hardware, model choice moves the economics by more than an order of magnitude.

The formula also omits quality. Frontier hosted models outperform anything that fits on a desk for hard reasoning and agentic coding tasks. Comparing a local 30B model against a frontier API on price alone compares different products. Our analysis of AI inference cost optimization discusses routing, caching and model-tiering strategies that narrow this gap from the cloud side.

When local wins regardless of price

Three non-price factors often decide the case. Data governance: if prompts contain source code, patient data or product designs that policy forbids sending to a third party, the API is not an option and the comparison ends. Latency and availability: a local model has no network round trip and works offline, which matters on factory floors and in air-gapped sites. Predictability: fixed capital cost is easier to budget than usage-based billing for agentic workloads, whose token consumption can spike unpredictably. For industrial and product-lifecycle environments that handle proprietary engineering data, these factors frequently outweigh the per-token arithmetic.

A hybrid pattern

The pragmatic architecture is a router. Routine, privacy-sensitive or high-volume tasks go to the local MoE model; rare hard problems go to a hosted frontier model after redaction. This keeps most tokens at near-zero marginal cost while preserving access to the best reasoning when it is worth paying for. A 192 GB workstation can also act as a small team server, running a batched vLLM-style endpoint for several users at once, which improves utilisation and therefore the economics.

Trade-offs, Gotchas, and What Goes Wrong

The design’s central compromise is that capacity is cheap relative to bandwidth. Anyone expecting the responsiveness of a high-end discrete GPU on a large dense model will be disappointed, because 273 GB/s is a fraction of what a current workstation GPU offers from dedicated memory. The platform wins on “can I run it at all,” not on “how fast.”

Thermal and power behaviour. The 45 to 120 W configurable range is wide, and the same silicon will behave differently in a thin laptop and a desktop-class chassis. Sustained inference is a continuous load, unlike bursty office work. We found no independent data on throttling under long generations, so ask OEMs for sustained rather than peak figures, and test with a realistic 30-minute run before standardising on a chassis.

The memory split is a configuration hazard. Reserving 160 GB for the GPU leaves about 32 GB for the OS, browser, IDE and anything else. An engineer running a large model alongside a CAD or simulation package can exhaust the remainder, triggering swapping that wrecks both applications. Budget memory for the whole workload, not just the model.

Soldered memory is not upgradeable. LPDDR5X in this class is typically soldered to the board, which is how it reaches its speed. The 192 GB decision is therefore made at purchase and cannot be corrected later. Under-buying is expensive; over-buying is also expensive. Match capacity to the largest model plus cache you expect to use over the machine’s life.

Software maturity still lags CUDA. ROCm support has improved and the listed runtimes are covered, but kernel coverage for newer quantization formats, attention variants and multimodal models can trail NVIDIA by weeks or months. A model released this morning may run on CUDA today and on this platform later. If you depend on bleeding-edge model support, validate before committing a fleet.

Vendor “up to” numbers. The 55 TOPS and 160 GB figures are ceilings. The precision behind TOPS, the real usable allocation, and the sustained clocks are all details that vary. Our own tokens-per-second figures are bandwidth-ceiling estimates; real results depend on the runtime, quantization format and context length.

Anti-patterns to avoid. Do not size a purchase by parameter count alone, since total and active parameters differ widely in sparse models. Do not buy the top SKU for single-user decode speed, because the memory subsystem is identical across the line. Do not assume the NPU accelerates your large model unless your runtime documents it. And do not compare a local quantized model against a frontier API on price without comparing output quality on your own tasks.

Practical Recommendations

Start from the workload, not the chip. If you mainly run 7B to 32B dense models for coding assistance and chat, a 64 or 96 GB configuration is probably enough, if OEMs offer it, and the extra memory of the 192 GB option buys little. If you want a large sparse model with long context, or several models resident at once, the 192 GB option is the reason to buy this platform and nothing else in the price class offers it.

Pick the SKU by prompt shape. For long-prompt retrieval and code-context workloads, the 40-CU 495 gives more prefill headroom. For short chat, the 32-CU parts deliver the same decode ceiling at lower cost. Treat the NPU as a bonus for background features, not a reason to choose a SKU.

Prefer mixture-of-experts models for interactive work and reserve dense 70B-plus models for batch tasks. Quantize to a 4-bit class format by default, then test whether an 8-bit variant of a smaller model gives better quality per second on your task. Quantize the KV cache if you need long contexts.

Before deploying widely, run a short pilot. Measure real tokens per second on your actual models, sustained for at least 30 minutes, and compare with the planning figures in this article. Put the break-even formula into a spreadsheet with your real quotes.

  • Confirm memory configuration options and the GPU-addressable limit with the OEM, in writing.
  • Test your target model with a realistic context length, not a short prompt.
  • Measure sustained, not peak, throughput and note fan noise and throttling.
  • Check runtime support for your model architecture and quantization format on ROCm.
  • Define a routing policy for what stays local and what may go to a hosted API.
  • Budget memory for the full workload, leaving at least 32 GB for the OS and applications.
  • Re-check pricing and independent benchmarks once units ship, since this article’s numbers are estimates.

Frequently Asked Questions

What is the AMD Ryzen AI Max+ Pro 400?

It is AMD’s workstation-class processor family for local AI, announced with three models: the Max+ PRO 495 with 16 Zen 5 cores, and the Max PRO 490 and 485 with 12 and 8 cores. All combine an RDNA 3.5 integrated GPU and an XDNA 2 NPU on one package, share up to 192 GB of LPDDR5X-8533 memory, and run at a configurable 45 to 120 W. They are sold through OEMs such as HP, Lenovo and ASUS.

How much memory can the GPU use?

AMD says up to 160 GB of the 192 GB unified pool can be allocated to the GPU. That leaves roughly 32 GB for the operating system and applications. The split is configurable, and exact procedures vary by OEM and operating system, so confirm the option with your vendor before purchase. Memory is shared, so every model, cache and application you run draws on the same pool, and it is generally soldered and not upgradeable later.

Can it really run a 300-billion-parameter model?

As a capacity claim, yes: 300 billion parameters at 4-bit quantization is about 150 GB, which fits under the 160 GB limit with little room for context. As a speed claim, it depends on the architecture. A dense 300B model would decode at only around 1 to 2 tokens per second by our bandwidth estimate, while a sparse mixture-of-experts model with far fewer active parameters could be several times faster.

Does the 50 to 55 TOPS NPU speed up large language models?

Not for token generation, which is limited by shared memory bandwidth that the NPU cannot increase. The NPU is most useful for low-power background tasks such as transcription and camera effects, and potentially for the prompt-processing phase in supported hybrid modes. For large models, plan around the integrated GPU running through ROCm-supported runtimes, and verify NPU support for your specific model before counting on it.

Is local inference cheaper than a cloud API?

Only above a break-even token volume. Divide the amortised hardware and power cost per month by your effective cloud price per million tokens to find it. With the hypothetical inputs in this article, break-even ranged from about 25 to 126 million tokens per month. Cost is rarely the only factor: data governance, offline operation and predictable budgets often justify local hardware even below break-even, while frontier quality often favours the cloud.

Which models run best on it?

Sparse mixture-of-experts models give the best experience, because they store a lot of knowledge but read few bytes per token. By our estimates, models in the 30B to roughly 120B total-parameter range with a few billion active parameters can reach tens of tokens per second. Dense 7B to 32B models are comfortable, and dense 70B models are usable at a few tokens per second. Dense models above 100 GB suit batch work only.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *