Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale
Fixed token IDs carried real language-modeling ability at 1.7B scale, but did not match learned embeddings.
The paper trained three decoder-only models from scratch under the same recipe, swapping only the input interface. Two used fixed token codes instead of a trainable input embedding table, removing 100.7 million trainable parameters. Those fixed-code models still reached measurable results, including 52.40% HellaSwag normalized accuracy for canonical codes. The learned-input control remained stronger on several evaluations, so the claim is viability, not parity. Source: HF Daily Papers' note
The paper trained three decoder-only models from scratch under the same recipe, swapping only the input interface. Two used fixed token codes instead of a trainable input embedding table, removing 100.7 million trainable parameters. Those fixed-code models still reached measurable results, including 52.40% HellaSwag normalized accuracy for canonical codes. The learned-input control remained stronger on several evaluations, so the claim is viability, not parity. Source: HF Daily Papers' note
score 5