Megadose Built for builders and researchers.

Learning to Learn a Language

· HF Daily Papers ·
A 300M-parameter byte model trained on synthetic non-language data can infer real languages from context with frozen weights.

The paper introduces PFLM, a byte-level transformer trained only on sequences from freshly sampled structural causal models. It never sees the same synthetic “language” twice, forcing it to infer the continuation rule from the prefix. On Wikipedia in six languages, its compression improves with long context, reaching 0.9 to 2.4 bits per byte at one million bytes. The authors also report transfer to numerals, deterministic sequences, and six non-text domains where it beats gzip and PPMd.

Source: HF Daily Papers' note

score 6

Categories: Research