Planning to Learn
The paper argues that cross-entropy beats exact policy gradient because it prices in future learning, not just the immediate reward.
Ian Osband frames classification as a policy-gradient problem and says the exact gradient is too myopic even when the correct label is known. The proposed “horizon loss” truncates cross-entropy according to how much training remains, shifting toward exact policy gradient late in training. In the paper’s tests on MNIST and ImageNet models, it improves top-1 accuracy over cross-entropy at a flat learning rate, with larger gains under label noise. ArXiv · AI/CL/LG's note
Ian Osband frames classification as a policy-gradient problem and says the exact gradient is too myopic even when the correct label is known. The proposed “horizon loss” truncates cross-entropy according to how much training remains, shifting toward exact policy gradient late in training. In the paper’s tests on MNIST and ImageNet models, it improves top-1 accuracy over cross-entropy at a flat learning rate, with larger gains under label noise. ArXiv · AI/CL/LG's note
score 4