Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models
Nine tabular foundation models all lost performance under real-world distribution shifts.
The paper tests TabPFN variants, TabICL variants, Mitra, LimiX, and TabFM on three TableShift datasets: HELOC, Voting, and Childhood Lead. The reported OOD gaps range from 0.003 to 0.060, depending on the shift type. The authors say stronger in-distribution performance still tends to track stronger OOD performance, as with classical tabular models. They also flag a deployment problem: some of the best-performing models require more memory and compute than standard infrastructure can support. ArXiv · AI/CL/LG's note
The paper tests TabPFN variants, TabICL variants, Mitra, LimiX, and TabFM on three TableShift datasets: HELOC, Voting, and Childhood Lead. The reported OOD gaps range from 0.003 to 0.060, depending on the shift type. The authors say stronger in-distribution performance still tends to track stronger OOD performance, as with classical tabular models. They also flag a deployment problem: some of the best-performing models require more memory and compute than standard infrastructure can support. ArXiv · AI/CL/LG's note
score 4