“FP4” on a spec sheet hides a set of formats with different block sizes, different scale types and very different behaviour on the tensors that matter.
Six ways
We quantized GLM-5 six ways across MXFP4 and NVFP4 variants, sweeping which tensors stay in higher precision, and measured throughput and evals together rather than one at a time.
The cliff
Throughput improves smoothly. Accuracy does not: there is a point where a small change in which projections are quantized costs several points on reasoning evals while barely moving the throughput curve. That point is what you actually need to find, and it is not on any vendor’s chart.