ScAn-Bench: Evaluating Scaling Analysis Methodology
The paper introduces two surrogate benchmarks to test how scaling-law analysis itself holds up.
ScAn-Bench-LLM and ScAn-Bench-VLM are built from 4,524 language-model checkpoints and 8,024 vision-language checkpoints. The authors frame this as a missing evaluation layer for scaling prescriptions across architecture, data, and hyperparameters. They say the benchmarks support a systematic look at data acquisition and extrapolation methods across modalities. ArXiv · AI/CL/LG's note
ScAn-Bench-LLM and ScAn-Bench-VLM are built from 4,524 language-model checkpoints and 8,024 vision-language checkpoints. The authors frame this as a missing evaluation layer for scaling prescriptions across architecture, data, and hyperparameters. They say the benchmarks support a systematic look at data acquisition and extrapolation methods across modalities. ArXiv · AI/CL/LG's note
score 6