Research

1 min read

FP4 is not one format: MXFP4, NVFP4, and the accuracy cliff nobody benchmarks

We quantized GLM-5 six ways and measured both the Pareto curve and the eval damage.

Momcilo Mrkaic, Pavle Padjin

“FP4” on a spec sheet hides a set of formats with different block sizes, different scale types and very different behaviour on the tensors that matter.

Six ways

We quantized GLM-5 six ways across MXFP4 and NVFP4 variants, sweeping which tensors stay in higher precision, and measured throughput and evals together rather than one at a time.

The cliff

Throughput improves smoothly. Accuracy does not: there is a point where a small change in which projections are quantized costs several points on reasoning evals while barely moving the throughput curve. That point is what you actually need to find, and it is not on any vendor’s chart.

Be there when
the next world opens.

Book a call