Megadose AI progress, ranked and analyzed.

Predicting Quantization Price for Selecting PTQ Configurations Before Deployment

· ArXiv · AI/CL/LG ·
The paper turns PTQ choice into a priced, budgeted configuration problem before the quantized model is deployed.

It treats each layer’s quantization option as producing an output-error covariance with an associated deployment cost.
The full-precision model then “prices” that error using downstream curvature, making formats, codebooks, granularities, transformations, and bit widths comparable.
The authors say this yields a calibration-time price table and selector, with ordinary fixed-geometry bit allocation reduced to a special case.
Source: ArXiv · AI/CL/LG's note.

score 4

Categories: Research