Learning Latent Protein Languages for Autoregressive Generation
The paper argues that learned latent tokens make autoregressive protein generation scale and sample better than raw amino acid tokens.
The authors introduce PLL for sequences and SLL for structures, then train next-token transformer models on each. PLLM shows a stronger fitted compute-scaling exponent than an amino-acid autoregressive baseline and cuts low-entropy generated samples by 54%. SLL improves validation perplexity for sequence-to-structure tokenization and supports backbone generation with favorable diversity and novelty results. The paper also reports roughly 1,000x faster latent-token sampling than MSA-based AlphaFold2 for long proteins in their measurements. HF Daily Papers' note
The authors introduce PLL for sequences and SLL for structures, then train next-token transformer models on each. PLLM shows a stronger fitted compute-scaling exponent than an amino-acid autoregressive baseline and cuts low-entropy generated samples by 54%. SLL improves validation perplexity for sequence-to-structure tokenization and supports backbone generation with favorable diversity and novelty results. The paper also reports roughly 1,000x faster latent-token sampling than MSA-based AlphaFold2 for long proteins in their measurements. HF Daily Papers' note
score 4