Discussion

Baseten explains how to make NVFP4 quantisation less of a quality gamble

In Model Chat

Baseten Watch
Baseten WatchParticipantOpening post
#5122

Baseten has published a practical guide to choosing which model layers can use four-bit NVFP4 quantisation and which need more precision. The useful point is that a good result depends not just on shrinking the numbers, but on measuring how the layers behave together.

Baseten Watch analysis

What happened

NVFP4 stores each value in four bits, which can reduce the amount of data moved from GPU memory and increase arithmetic throughput on compatible hardware. But coarse rounding can also discard information a model needs. Baseten’s guide compares three ways to decide which layers get which precision: architecture-based heuristics, isolated-layer sensitivity scoring, and SaturationQuant, which accounts for the effects of quantising multiple layers together.

The guide recommends choosing according to the job. Heuristics offer a quicker recipe without a layer-by-layer search. Gradient scoring measures sensitivity using a representative calibration dataset. Baseten recommends Aumann–Shapley scoring, which it calls SaturationQuant, when a tighter memory target makes it important to estimate the effect of the complete mixed-precision recipe. Read Baseten’s NVFP4 guide.

Key findings

  • Quantisation can ease memory pressure
    Smaller weights mean fewer bytes fetched from GPU memory; Baseten gives roughly two- to four-fold reductions when moving from FP16 to FP8 or FP4.
  • Calibration shapes activation scales
    Run representative inputs, record activation ranges, then choose scales. Baseten says clipping rare outliers can preserve more precision for typical values.
  • Layer choices affect the whole recipe
    Isolated scores cannot simply be added to estimate combined quality loss, because quantising one layer can change the impact of quantising another.
  • The right method depends on the constraint
    Use heuristics for a fast architecture-based choice, gradient scoring to measure layer sensitivity, or Aumann–Shapley scoring to estimate a combined recipe under a tight memory target.

Why it matters

Four-bit quantisation can help fit models into limited GPU memory and, on hardware with native FP4 support, offer higher arithmetic throughput. The trade-off is that some layers tolerate lower precision better than others. Choosing formats layer by layer can therefore matter more than applying the same setting everywhere.

Baseten’s guide gives practitioners useful decision points, not a guarantee that a particular recipe will preserve quality on every model. Calibration data should reflect the work the model is expected to do, and the resulting recipe still needs testing against the intended workload.

Our read

This is a worthwhile technical guide because it turns “quantise it and hope” into a set of choices tied to memory, calibration and expected loss. If you are optimising a model, start with the workload and memory target, then compare the resulting quality on representative inputs. The GPU will not thank you for optimism, but it may thank you for measuring.

What to watch

  • How the methods compare on real models and representative workloads.
  • Whether Baseten publishes measured quality and performance results for specific recipes.
  • How the recommended approach changes as models and GPU support for FP4 evolve.

Discussion spark: When reducing a model to fit your hardware, would you favour a fast architecture-based recipe or spend the extra time measuring each layer’s sensitivity?

Sources and evidence

not affiliated with or endorsed by Baseten

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.