Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
WeightPE trains network weights to compress into a smaller grammar, trading some accuracy for more reusable structure.
The method flattens int8 weights into a string, then uses a lossy Re-Pair compressor inside a straight-through estimator. Near-matching patterns are forced to match within a global L2 budget, while the network trains through the rewritten weights. On ViT-B/16 and ViT-L/16 MLP weights finetuned on CIFAR-10, the resulting grammars were 0.43x and 0.38x the size of an equivalent int8 QAT baseline, with 1.9 and 1.1 accuracy-point losses. The authors say the effect also carries to LZ78 and SEQUITUR, even though training targeted Re-Pair. ArXiv · AI/CL/LG's note
The method flattens int8 weights into a string, then uses a lossy Re-Pair compressor inside a straight-through estimator. Near-matching patterns are forced to match within a global L2 budget, while the network trains through the rewritten weights. On ViT-B/16 and ViT-L/16 MLP weights finetuned on CIFAR-10, the resulting grammars were 0.43x and 0.38x the size of an equivalent int8 QAT baseline, with 1.9 and 1.1 accuracy-point losses. The authors say the effect also carries to LZ78 and SEQUITUR, even though training targeted Re-Pair. ArXiv · AI/CL/LG's note
score 4