How Does Distribution Shift Shape Pretraining Gains in Neural PDE Surrogates?
Pretraining helped most when target data was scarce, but the size of the gain changed with both airfoil coverage and modeled physics.
The paper pretrains a neural PDE surrogate on 254,909 RANS solutions from one airfoil family, then fine-tunes on another. With 1,000 target samples, pretraining matched scratch training with 3.25x more data for the same Spalart-Allmaras setup, versus 2.58x for the transition-modeled target. At 5,000 samples, that ordering flipped, with the transition-modeled target showing the larger relative gain. The authors also report that sampling more distinct target airfoils helped both targets at 1,000 samples, but the gain was clearly above draw-to-draw variation only in the same-SA case. ArXiv · AI/CL/LG's note
The paper pretrains a neural PDE surrogate on 254,909 RANS solutions from one airfoil family, then fine-tunes on another. With 1,000 target samples, pretraining matched scratch training with 3.25x more data for the same Spalart-Allmaras setup, versus 2.58x for the transition-modeled target. At 5,000 samples, that ordering flipped, with the transition-modeled target showing the larger relative gain. The authors also report that sampling more distinct target airfoils helped both targets at 1,000 samples, but the gain was clearly above draw-to-draw variation only in the same-SA case. ArXiv · AI/CL/LG's note
score 4