Megadose AI progress, ranked and analyzed.

Telescopic Language Models

· ArXiv · AI/CL/LG ·
A single trained model is reported to work across all 20 layer-prefix budgets without extra inference machinery.

The paper trains a nested Transformer with stochastic prefix supervision plus a full-capacity anchor, using two forward-backward passes per step. In a 200M-parameter proxy setup on 20B FineWeb-Edu tokens, the authors say it reduces the area under the quality-budget curve by 43-44% versus fixed-exit suites while matching full-capacity quality. They argue fixed-exit supervision leaves unsupervised intermediate prefixes unusable, with baseline perplexities at chance-level ranges. The sampling density can be shifted toward selected depths, trading continuum performance for stronger fixed operating points. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research