LLM Quantization Formats Compared: GGUF vs AWQ vs GPTQ vs FP8
A 70-billion-parameter model stored in 16-bit floats needs roughly 141 GB for its weights alone, which is more than any single GPU on the market holds. The same model in a 4-bit format fits in about 40 GB, and in FP8 it fits in about 71 GB. Yet the file you download is not “the model at 4 bits”. It is a specific packing scheme, with a specific calibration story and a specific set of kernels that can read it, and those details decide both accuracy and speed.
The four names that dominate Hugging Face model cards are GGUF, AWQ, GPTQ and FP8. The current catalogue of LLM quantization formats looks like a menu of equivalent options, but they differ on three axes: what gets quantized (weights only, or weights and activations), how the rounding error is minimized, and which runtime can execute the result efficiently. Choosing on file size alone is the most common mistake.
This article takes a mechanism-first view. You will leave with worked memory arithmetic for 8B and 70B models, a clear picture of how each format represents numbers, the exact commands to produce and serve each one, and a decision matrix tied to hardware rather than folklore.
What this covers: the weight-only versus weight-and-activation split, group size and block scales, GGUF k-quants, AWQ and GPTQ calibration, FP8 E4M3 and E5M2, the NVFP4 and MXFP4 status, engine support in llama.cpp, vLLM and TensorRT-LLM, failure modes, and a decision matrix.
Context and Background
Quantization replaces high-precision numbers with a small set of low-precision codes plus a scale that maps them back. For large language models the dominant cost at inference time is moving weights from memory to compute units, so shrinking weights directly raises decode throughput. Decode, the token-by-token phase, is memory-bandwidth bound: each generated token reads essentially every weight once. Cutting the bytes per weight by four cuts the bytes read by four, and speed follows if the kernel is good.
Two broad families exist. Weight-only quantization (written W4A16: 4-bit weights, 16-bit activations) stores weights in few bits but dequantizes them to 16-bit just before the matrix multiply. It saves memory and bandwidth, but the arithmetic still runs at 16-bit speed. Weight-and-activation quantization (W8A8) quantizes both operands so the matrix multiply itself runs on 8-bit hardware, which helps the compute-bound prefill phase as well.
The four formats map onto this split. GGUF, AWQ and GPTQ are overwhelmingly used as weight-only schemes. FP8 as deployed in vLLM through LLM Compressor is a W8A8 scheme. The vLLM documentation lists exactly this: LLM Compressor offers FP8 W8A8, INT4 W4A16, INT8 W4A8 and INT8 W8A8, while AutoAWQ and GPTQModel are separate integrations (vLLM quantization docs).
If you want measured speed and accuracy numbers across precisions rather than mechanisms, our FP8 vs INT8 vs INT4 LLM quantization benchmark covers that ground. This article is the companion piece on container formats and algorithms. For constrained silicon, the INT4 vs INT8 vs FP8 edge NPU quantization guide applies the same ideas to NPUs.
One framing point before the details. A “format” bundles three separable things: a number representation, an algorithm that chose the values, and a file layout that a kernel can read. GPTQ and AWQ are algorithms that happen to emit similar 4-bit group-quantized layouts. GGUF is a file container that holds many different quantization types. FP8 is a number representation that needs almost no algorithm at all. Keeping these layers apart removes most of the confusion in forum arguments.
The Core Mechanism: What Each Format Actually Stores
The short answer: GGUF stores weights in blocks with per-block scales and is built for llama.cpp on CPUs and mixed devices. AWQ and GPTQ store 4-bit integers with per-group scales, chosen by activation-aware scaling or Hessian-based rounding respectively, and target GPU engines. FP8 stores 8-bit floats and usually quantizes activations too, targeting recent NVIDIA and AMD GPUs.

Figure 1: How the four LLM quantization formats split by what is quantized and which runtime consumes them.
The diagram shows a trained checkpoint branching first on whether activations are quantized, then on the algorithm. The three weight-only branches converge on different engines: GGUF files go to llama.cpp, while AWQ, GPTQ and FP8 checkpoints are loaded by GPU serving engines such as vLLM and TensorRT-LLM.
Group size and block scales: the shared idea
All integer schemes share one trick. Quantizing a whole tensor with a single scale is crude, because one outlier stretches the range and wastes codes on everyone else. Instead the weights are split into small groups, typically 32 to 128 consecutive values, and each group gets its own scale (and sometimes a zero point).
The cost is metadata. A 4-bit scheme with a 16-bit scale per group of 128 weights spends 4 + 16/128 = 4.125 bits per weight. Add a 4-bit zero point and it is 4.156 bits. A group of 32 with the same 16-bit scale costs 4.5 bits, which is why smaller groups buy accuracy at the price of file size and some kernel efficiency. Remember this arithmetic, because it explains why “4-bit” files are never exactly a quarter of the FP16 size.
GGUF and the k-quant family
GGUF is the single-file binary format used by llama.cpp. It stores tensors plus metadata such as the tokenizer, so one file is self-contained. Inside it, each tensor carries its own quantization type, which means a “Q4_K_M” file is actually a mix of types.
The k-quants use a two-level scheme. Q4_K packs weights into super-blocks of 256 values, split into eight sub-blocks of 32. Each sub-block has its own 6-bit scale and 6-bit minimum, and each super-block has a 16-bit scale and 16-bit minimum for those. The arithmetic gives (256 x 4 bits + 8 x 6 + 8 x 6 + 2 x 16) / 256 = 4.5 bits per weight. I derived this from the block layout as I understand it from the llama.cpp source; the llama.cpp quantize README confirms that the measured file averages are higher (4.89 bits for Q4_K_M) because the “M” mix keeps some sensitive tensors at higher precision.
Per the llama.cpp quantize README for Llama 3.1 8B, the reported sizes include Q4_K_S at 4.67 bits per weight (4.36 GiB), Q4_K_M at 4.89 (4.58 GiB), Q5_K_M at 5.70 (5.33 GiB), Q6_K at 6.56 (6.14 GiB), Q8_0 at 8.50 (7.95 GiB) and F16 at 16.0 (14.96 GiB) (llama.cpp quantize README). Q8_0 is a simpler scheme: blocks of 32 int8 values with one 16-bit scale, giving (32 x 8 + 16)/32 = 8.5 bits.
The suffixes S, M and L denote small, medium and large mixes. The “I” family (IQ4_XS, IQ3_S and so on) uses non-linear codebooks and generally requires an importance matrix for best results, covered below.
AWQ: activation-aware scaling
Activation-aware Weight Quantization starts from an observation in the paper: not all weights matter equally, and protecting roughly 1% of salient weights greatly reduces quantization error (Lin et al., arXiv:2306.00978). Crucially, saliency is read from activation magnitudes, not weight magnitudes. A weight channel that multiplies consistently large activations contributes more error when rounded.
Keeping 1% of weights in FP16 would give mixed-precision tensors that kernels hate. AWQ avoids that by an equivalent transformation: scale the salient input channel of the weight up by a factor s and scale the corresponding activation down by 1/s. The product is unchanged mathematically, but the scaled-up weight occupies a larger part of the quantization grid, so its relative rounding error shrinks. The scales come from offline activation statistics, with no backpropagation or reconstruction, which the authors argue reduces overfitting to the calibration set.
In practice that means AWQ quantization is fast, needs only a small calibration batch, and produces a plain group-quantized INT4 layout. The common deployment form is W4 with group size 128.
GPTQ: second-order rounding
GPTQ treats quantization as a layer-wise reconstruction problem. For each linear layer it wants quantized weights whose output on calibration inputs is as close as possible to the original. It uses approximate second-order information, the Hessian of the layer’s squared error, to decide how to round each weight and how to adjust the not-yet-quantized weights to compensate (Frantar et al., arXiv:2210.17323).
Concretely, GPTQ quantizes a weight column, measures the error that rounding introduced, and spreads a correction across remaining columns using the inverse Hessian. The paper reports quantizing a 175B model in about four GPU hours with accuracy staying close to baseline at 3 to 4 bits, and notes usable if degraded results at 2 bits.
The result is again a packed integer layout with per-group scales. What differs is how the integers were chosen: GPTQ minimizes reconstruction error directly, where AWQ reshapes the problem with channel scales first and then rounds plainly.
FP8: E4M3 and E5M2
FP8 is not a grouping trick but a different number system. The reference paper defines two encodings: E4M3 with 4 exponent bits and 3 mantissa bits, and E5M2 with 5 exponent bits and 2 mantissa bits (Micikevicius et al., arXiv:2209.05433). E4M3 gives up infinities and keeps one NaN pattern to stretch its range; E5M2 follows IEEE 754 conventions for special values.
The tradeoff is range versus precision. E5M2 has a wider dynamic range and coarser steps, so it suits gradients during training. E4M3 has finer steps and a narrower range, so it is the usual choice for weights and activations in inference. Because floating point spaces its codes logarithmically, FP8 handles outliers more gracefully than INT8, whose 256 codes are evenly spaced. A per-tensor or per-channel scale still maps the data into FP8’s representable range, but no calibration search is needed beyond measuring maxima.
NVFP4 and MXFP4: the 4-bit floating point arrival
A newer branch replaces integer 4-bit with 4-bit floats. NVIDIA’s NVFP4 stores each value as E2M1 (1 sign, 2 exponent, 1 mantissa bit), groups values into micro-blocks of 16, and gives each block an FP8 E4M3 scale plus one FP32 scale per tensor. The Open Compute Project’s MXFP4 uses blocks of 32 with a power-of-two E8M0 scale. NVIDIA’s post puts NVFP4 at roughly 4.5 bits per value including block scales and reports hardware acceleration in Blackwell fifth-generation Tensor Cores (NVIDIA NVFP4 introduction).
The same post claims 1% or less degradation versus FP8 on several DeepSeek-R1-0528 benchmarks with post-training quantization. That is a vendor-reported result on one model family, so treat it as a promising signal rather than a general guarantee. The vLLM docs list Marlin kernels covering GPTQ, AWQ, FP8 and FP4 on Turing and newer (not MXFP4 on Turing), and show an NVFP4 W4A16 backend option, so FP4 support exists but is the youngest and fastest-moving part of the stack. Verify your specific engine version before committing.
Deeper Analysis: Calibration, Arithmetic, and Kernels
Calibration is where the formats diverge in cost and risk. The next diagram contrasts how each method consumes a small calibration set.

Figure 2: Calibration pipelines. AWQ searches channel scales, GPTQ updates columns using the Hessian, and the GGUF importance matrix weights block error.
All three methods first push a calibration set through the full-precision model to collect activation statistics. AWQ then searches per-channel scales; GPTQ walks columns with error compensation; llama.cpp’s importance matrix (imatrix) records how strongly each weight column is excited and weights the block-wise rounding error accordingly. GGUF can also be produced with no calibration at all, which is faster but loses accuracy at the lowest bit widths.
Worked memory arithmetic for 8B and 70B
Use the Llama 3.1 family as the concrete case. The README’s F16 size of 14.96 GiB for the 8B model implies about 8.03 billion parameters (14.96 x 2^30 / 2 bytes). The 70B model has about 70.6 billion. The formula is bytes = parameters x bits per weight / 8.
| Format | Bits per weight | 8B weights | 70B weights |
|---|---|---|---|
| BF16 or F16 | 16 | 16.1 GB | 141 GB |
| FP8 (8 bit, scales negligible) | about 8 | 8.0 GB | 70.6 GB |
| GGUF Q8_0 | 8.50 | 8.5 GB | 75 GB |
| GGUF Q6_K | 6.56 | 6.6 GB | 58 GB |
| GGUF Q5_K_M | 5.70 | 5.7 GB | 50 GB |
| GGUF Q4_K_M | 4.89 | 4.9 GB | 43 GB |
| AWQ or GPTQ W4 g128 | about 4.13 to 4.25 | about 4.2 GB | about 37 GB |
The 8B bits-per-weight figures come from the llama.cpp README; the 70B column applies the same ratios to 70.6B parameters, so it is an estimate. The AWQ and GPTQ row follows from the group arithmetic above and varies slightly with whether zero points are stored and whether the embedding and output layers are left in 16-bit. Real checkpoints are typically a few percent larger because embeddings and the output head are often kept at higher precision.
Weights are only part of the bill. The key-value (KV) cache grows with context and batch size. For Llama 3.1 8B (32 layers, 8 KV heads, head dimension 128, as I recall from its published configuration) the cache costs 2 x 32 x 8 x 128 x 2 bytes = 128 KiB per token at FP16. At 32,768 tokens that is 4 GiB per sequence. For 70B (80 layers, same 8 KV heads and head dimension) it is 320 KiB per token, so 10 GiB per 32k sequence.
Two consequences follow. A 70B Q4_K_M model at 43 GB weights plus a single 32k sequence needs about 53 GB, which fits one 80 GB GPU with room for several concurrent sequences but not a 48 GB card. And at high batch sizes the KV cache, not the weights, dominates memory, which is why weight-only 4-bit alone cannot rescue a long-context, high-concurrency deployment. An FP8 KV cache halves that cost, and vLLM lists quantized KV cache as a separate feature.
Why speed is not proportional to size
The intuition “4 bits is 4x faster” holds only in the memory-bound decode regime with a kernel that dequantizes efficiently. The next diagram shows the two phases of a request.

Figure 3: Prefill is compute bound and benefits from FP8 tensor cores; decode is memory bound and benefits from 4-bit weight reads.
Prefill processes the whole prompt in parallel and is limited by arithmetic throughput. W4A16 formats do not speed that up, because they dequantize to FP16 and use FP16 math; W8A8 FP8 does, because Hopper, Ada and Blackwell tensor cores execute FP8 matrix multiplies natively. Decode generates one token per step per sequence, reads all weights, and is limited by bandwidth, so 4-bit weights help most there.
At large batch sizes decode becomes compute bound too, and the advantage of W4A16 shrinks while FP8 keeps its edge. This is the central practical tension: W4A16 wins at low concurrency (a single user, a laptop, an edge box), FP8 tends to win at high-throughput server batch sizes. I am stating this as the mechanism-level expectation; the crossover batch size depends on GPU, model and kernel, and you should measure it.
Kernel quality can swamp format differences. The vLLM docs show Marlin kernels shared by GPTQ and AWQ checkpoints, so those two formats can run at nearly the same speed when both are repacked into Marlin layout. A GPTQ checkpoint run through a slow legacy kernel can lose badly to the same bits in a fast kernel. Judging formats by headline accuracy and ignoring which kernel your engine will select is how teams end up with a 4-bit model that is slower than FP16.
Engine and hardware support
The vLLM documentation gives a compatibility table. AWQ runs on NVIDIA Turing and newer, Intel GPU and x86 CPU; GPTQ on NVIDIA Volta and newer, Intel GPU and x86 CPU; LLM Compressor FP8 W8A8 on NVIDIA Ada and Hopper and on AMD GPU; GGUF on NVIDIA Volta and newer and AMD GPU. The page warns the chart may change, so check the current version.
In that table GGUF support in vLLM exists, but llama.cpp is the format’s native home and offers the broadest backend range: CPU with SIMD, Apple Metal, CUDA, ROCm, Vulkan and others. It can also offload some layers to the GPU and keep the rest in system RAM, which no GPU-only engine matches. TensorRT-LLM, NVIDIA’s engine, emphasizes FP8 and FP4 paths on recent NVIDIA hardware plus INT4 AWQ and GPTQ weight-only modes; consult its documentation for the exact precision matrix of your release rather than relying on a blog summary.
Runnable commands
Producing a GGUF k-quant with llama.cpp starts from a Hugging Face checkpoint. The sequence is convert to a 16-bit GGUF, optionally compute an importance matrix, then quantize.
# 1. Convert HF weights to an F16 GGUF
python convert_hf_to_gguf.py ./Llama-3.1-8B-Instruct --outtype f16 \
--outfile llama31-8b-f16.gguf
# 2. Compute an importance matrix on a calibration text file
./build/bin/llama-imatrix -m llama31-8b-f16.gguf -f calib.txt -o imatrix.gguf
# 3. Quantize to Q4_K_M using the imatrix
./build/bin/llama-quantize --imatrix imatrix.gguf \
llama31-8b-f16.gguf llama31-8b-Q4_K_M.gguf Q4_K_M
# 4. Run it, offloading all layers to the GPU
./build/bin/llama-cli -m llama31-8b-Q4_K_M.gguf -ngl 99 -c 8192 -p "Hello"
The quantize syntax (llama-quantize [--imatrix file] input.gguf output.gguf TYPE) is taken from the README. The conversion script name and the llama-imatrix flags match recent llama.cpp builds as I know them, but flag spellings change, so run each tool with --help.
For GPU serving, the simplest FP8 path is dynamic quantization by LLM Compressor, which needs no calibration data, followed by a vLLM launch.
# pip install llmcompressor
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
model_id = "meta-llama/Llama-3.1-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto")
tok = AutoTokenizer.from_pretrained(model_id)
recipe = QuantizationModifier(
targets="Linear", scheme="FP8_DYNAMIC", ignore=["lm_head"]
)
oneshot(model=model, recipe=recipe)
model.save_pretrained("Llama-3.1-8B-FP8-Dynamic")
tok.save_pretrained("Llama-3.1-8B-FP8-Dynamic")
# Serve the FP8 checkpoint
vllm serve ./Llama-3.1-8B-FP8-Dynamic --max-model-len 16384
# Serve a published AWQ or GPTQ checkpoint; the format is read from its config
vllm serve <org>/<model>-AWQ --max-model-len 16384
The LLM Compressor class names above reflect its documented API at the time of my knowledge; verify them against the installed version. vLLM normally detects the quantization method from the checkpoint’s quantization_config, so an explicit --quantization flag is usually unnecessary for pre-quantized models.
A Format-by-Format Comparison
With mechanisms and numbers in hand, the formats can be compared on the dimensions that matter operationally.
| Dimension | GGUF (k-quants) | AWQ | GPTQ | FP8 W8A8 |
|---|---|---|---|---|
| What is quantized | Weights | Weights | Weights | Weights and activations |
| Typical bits | 2 to 8, mixed per tensor | 4 | 2 to 8, mostly 4 | 8 |
| Selection method | Block min/max, optional imatrix | Activation-aware channel scaling | Hessian-based rounding | Max or percentile scaling |
| Calibration need | Optional | Small set | Small set | None for dynamic, small set for static |
| Quantization cost | Minutes | Minutes to under an hour for 7B-class | Longer, grows with size | Minutes |
| Best engine | llama.cpp | vLLM, TensorRT-LLM | vLLM, TensorRT-LLM | vLLM, TensorRT-LLM |
| Strongest on | CPU, Apple silicon, mixed | Low-concurrency GPU | Low-concurrency GPU | High-throughput GPU |
| Needs recent GPU | No | Turing and newer in vLLM | Volta and newer in vLLM | Ada, Hopper or newer |
The timing cells are qualitative. I did not find authoritative cross-method timing benchmarks to cite, apart from the GPTQ paper’s four GPU hours for a 175B model, so the rows reflect order of magnitude rather than measured values.
Accuracy: what is actually known
Published accuracy comparisons depend on the model, the benchmark and the calibration set, and results from one family do not transfer cleanly to another. The claims that survive across papers are limited. Both AWQ and GPTQ stay close to baseline at 4 bits on language-modeling perplexity; the GPTQ paper notes near-baseline accuracy at 3 to 4 bits; and FP8 post-training quantization works well on models where INT8 struggled, according to the FP8 formats paper.
Perplexity is a blunt instrument. The llama.cpp README itself names perplexity and KL divergence as the usual measures, and KL divergence against the full-precision token distribution is the better one for quantization, because it captures shifts in the whole distribution rather than only the probability of the reference token. Neither tells you about long-form reasoning, code, tool-calling or multilingual quality, where small distribution shifts can compound over a long generation.
I deliberately give no league table of “Q4_K_M scores X on benchmark Y”. Any such number I could write without a source would be fabricated, and the spread between different evaluation harnesses is larger than the gaps people argue about. Run your own evaluation instead, as the recommendations below describe.
Where the formats overlap and where they truly differ
Operationally, AWQ and GPTQ are close cousins: 4-bit, group-quantized, same kernel family, similar footprint. The deciding factors are usually whether a well-tested checkpoint for your exact model already exists and how that checkpoint behaves on your tasks. AWQ’s lack of reconstruction training is sometimes claimed to make it more robust to distribution shift between calibration and deployment data, which matches the paper’s argument but is not a guarantee.
GGUF differs structurally. Its mixed-precision design (keeping sensitive tensors at 6 bits inside a Q4_K_M file) is a form of sensitivity-aware allocation that the uniform AWQ and GPTQ layouts do not do by default. That is one reason a 4.89-bit GGUF can rival a 4.25-bit GPTQ file despite the extra size. FP8 differs in kind: it trades half the compression for a much simpler, lower-risk path and for hardware-native arithmetic.
Trade-offs, Gotchas, and What Goes Wrong
Quantization failures are rarely dramatic. A broken model that emits gibberish is caught on day one. The dangerous failures are quiet regressions that survive a smoke test and surface weeks later in production.
Calibration mismatch. AWQ, GPTQ and imatrix all learn from a calibration set. If you calibrate on English Wikipedia and serve code, structured JSON or Hindi, the statistics that protected “salient” channels may be the wrong ones. The usual symptom is degraded tool-call formatting or multilingual quality while English chat looks fine. Calibrate on text resembling your traffic, and keep the sample count in the low hundreds so you do not overfit.
Outliers and activation quantization. Weight-only schemes sidestep the hardest problem in LLM quantization, activation outliers, because activations stay in 16-bit. W8A8 cannot. FP8’s logarithmic spacing tolerates outliers better than INT8, but static per-tensor activation scales can still clip rare large values. Dynamic per-token scaling avoids the clipping at the cost of a little runtime overhead, which is why the dynamic scheme is a safe default.
Sensitive layers. The embedding table and output head (lm_head) are routinely left in higher precision because errors there hit every token. The ignore=["lm_head"] line in the recipe above does this. Mixture-of-experts routers and the first and last transformer blocks are other candidates for exclusion. Our analysis of Tencent’s HY4 770B MoE and Sherry quantization shows how expert-heavy models change the sensitivity picture.
Kernel and hardware mismatch. A checkpoint can load and still run on a slow fallback path. FP8 W8A8 in vLLM is listed for Ada, Hopper and AMD GPUs; on an Ampere card the same checkpoint may fall back to weight-only behavior or fail, depending on version. Always confirm the kernel actually selected from the engine’s startup log, then compare tokens per second against the FP16 baseline on your hardware.
Re-quantizing a quantized model. Converting an AWQ checkpoint to GGUF, or a GGUF down to a smaller GGUF, compounds error. Quantize from the highest-fidelity source you have, ideally the original BF16 weights. The llama.cpp README’s examples start from an F32 or F16 GGUF for this reason.
Format sprawl in the supply chain. Third-party quantized uploads vary in quality and provenance. Pin exact repository revisions, check the config for group size and whether the output head was quantized, and prefer a checkpoint from the model publisher or a quantization tool you can reproduce yourself. Pickle-based formats carry code-execution risk, so favor safetensors and GGUF over arbitrary pickles.
The memory trap. Weight savings can vanish into the KV cache. A team that quantizes a 70B model to 4 bits and then serves 128k-token contexts at batch 16 will still run out of memory. Budget weights, KV cache at your real context length, activations and engine overhead separately, as in the arithmetic above.
Quality cliffs below 4 bits. Moving from 8 to 6 to 5 to 4 bits is usually gentle. Below 4 bits, degradation accelerates, and the llama.cpp 2- and 3-bit types lean on imatrix and non-linear codebooks to stay usable. Small models suffer first: an 8B model at 3 bits is far more fragile than a 70B model at 3 bits, because it has less redundancy to spend.
Practical Recommendations
Start from the hardware and the traffic shape, not from the format name. The decision flow below summarizes the choices.

Figure 4: A decision flow from serving target to format, ending in validation on your own evaluations.
For a laptop, workstation, Apple silicon or CPU-and-GPU split, use GGUF with Q4_K_M as the default and Q5_K_M or Q6_K if memory allows. For a modern data-center GPU serving many concurrent users, FP8 W8A8 is the lowest-risk first step: half the memory of BF16, native tensor-core math, minimal calibration. If FP8 does not fit the model on your GPUs, step down to a W4A16 AWQ or GPTQ checkpoint and accept the lower compute speedup.
For older NVIDIA GPUs without FP8 hardware, AWQ or GPTQ 4-bit through a Marlin-class kernel is the practical route, subject to the architecture limits in the vLLM table. If you are on Blackwell and chasing density, evaluate NVFP4 against FP8 on your own quality tests, given the vendor-reported results are limited to specific models.
A decision matrix for common use cases:
| Use case | First choice | Alternative | Why |
|---|---|---|---|
| Local assistant on a laptop | GGUF Q4_K_M | GGUF Q5_K_M | Single user, memory bound, CPU and Metal support |
| Single 24 GB GPU, 8B to 14B model | GGUF Q5_K_M or AWQ | GPTQ | Fits with context headroom |
| Production API at high concurrency on H100 class | FP8 W8A8 | AWQ W4A16 | Compute-bound batches favor native FP8 |
| 70B model on 1x 80 GB GPU | AWQ or GPTQ W4 | GGUF Q4_K_M | FP8 weights alone are about 71 GB |
| Older GPU without FP8 | AWQ W4 | GPTQ W4 | Kernel support on Turing and newer |
| Edge device with NPU | See the edge guide | INT8 | Vendor toolchains dominate |
A short checklist before you ship:
- Quantize from the original BF16 weights, not from another quantized file.
- Calibrate on data that resembles production traffic.
- Keep
lm_headand embeddings at higher precision unless you have measured the effect. - Compare KL divergence against the full-precision model on a held-out sample, and run your task-level evaluations.
- Check the engine startup log to confirm the intended kernel, and benchmark at your real batch size and context length.
- Budget KV cache memory separately and consider an FP8 KV cache.
- Pin the checkpoint revision and record the quantization recipe for reproducibility.
Frequently Asked Questions
Which is better, GGUF or AWQ?
Neither is better in general; they target different runtimes. GGUF is the native format of llama.cpp and runs well on CPUs, Apple silicon and mixed CPU-GPU setups. AWQ is a 4-bit GPU format served mainly through vLLM and TensorRT-LLM. If your hardware is a single machine with limited VRAM, choose GGUF. If you are serving from data-center NVIDIA GPUs and want a 4-bit checkpoint, choose AWQ, then verify accuracy on your own tasks.
Is GPTQ outdated now that AWQ and FP8 exist?
GPTQ is older, but it is still supported: the vLLM table lists it for NVIDIA Volta and newer, which is broader than AWQ’s Turing-and-newer listing. Its Hessian-based rounding remains a strong accuracy baseline, and many published checkpoints exist. Newer tooling such as GPTQModel and LLM Compressor keeps the algorithm alive. The practical decision is usually which tested checkpoint and kernel exist for your model, rather than the algorithm’s age.
How much accuracy do I lose with 4-bit quantization?
It depends on model, method and benchmark, so no single figure is honest. Papers on AWQ and GPTQ report accuracy close to the full-precision baseline at 4 bits, and larger models tolerate it better than small ones. Measure it yourself: compute KL divergence against the BF16 model on a held-out sample, then run task-level evaluations such as code, tool use and your domain data. Small perplexity gaps can hide larger task regressions.
Can I run FP8 on an RTX 3090 or A100?
Not with native speedups. Ampere GPUs lack FP8 tensor cores, and vLLM lists FP8 W8A8 for Ada, Hopper and AMD GPUs. Some engines can store FP8 weights and dequantize on the fly for memory savings, but you lose the compute advantage, and support varies by version. On Ampere, AWQ or GPTQ 4-bit through Marlin kernels is the usual route. Confirm in your engine’s release notes before planning around it.
What is the difference between Q4_K_M and Q4_K_S in GGUF?
Both use the 4-bit k-quant block scheme, but they allocate precision differently across tensors. Q4_K_M keeps selected sensitive tensors at higher precision, so it averages 4.89 bits per weight (4.58 GiB for Llama 3.1 8B) versus 4.67 bits (4.36 GiB) for Q4_K_S, according to the llama.cpp README. The extra roughly 0.2 GiB buys better accuracy. Q4_K_M is the common default for that reason.
Are NVFP4 and MXFP4 ready to replace FP8?
Not universally. Both are 4-bit float formats with block scales: NVFP4 uses 16-value blocks with FP8 scales, MXFP4 uses 32-value blocks with power-of-two scales. NVFP4 is hardware-accelerated on Blackwell, and NVIDIA reports small degradation versus FP8 on some models. Engine support is still maturing and results are model-specific, so treat FP4 as something to evaluate on your workload rather than a drop-in replacement for FP8 today.
Further Reading
- FP8 vs INT8 vs INT4 LLM quantization benchmark 2026: measured speed and accuracy across precisions.
- INT4 vs INT8 vs FP8 for edge NPU quantization: the same trade-offs on constrained silicon.
- Tencent HY4 770B MoE and Sherry quantization explained: quantization for very large mixture-of-experts models.
- vLLM quantization documentation: the current hardware support table.
- llama.cpp quantize README: quantization types, sizes and imatrix usage.
By Riju — about
