Disentangling Paradigm, Identifier, and Decoding in Generative Retrieval
Decoding choices, not just model architecture, account for large swings in diffusion-based generative retrieval results.
The paper tests autoregressive, masked-diffusion, and block-diffusion retrievers across fixed budgets, identifier types, and decoding methods. It finds diffusion Hit@1 can move by 6.6 to 13.7 points from decoding alone, while one-pass scoring matches or beats generate-and-match in 11 of 12 settings. Autoregressive models still lead on Hit@1, and on NQ320K the authors say that lead comes from the model rather than beam search. Random identifiers retaining 83–90% of residual-quantised performance suggests much of the task is memorizing query-to-identifier mappings. ArXiv · AI/CL/LG's note
The paper tests autoregressive, masked-diffusion, and block-diffusion retrievers across fixed budgets, identifier types, and decoding methods. It finds diffusion Hit@1 can move by 6.6 to 13.7 points from decoding alone, while one-pass scoring matches or beats generate-and-match in 11 of 12 settings. Autoregressive models still lead on Hit@1, and on NQ320K the authors say that lead comes from the model rather than beam search. Random identifiers retaining 83–90% of residual-quantised performance suggests much of the task is memorizing query-to-identifier mappings. ArXiv · AI/CL/LG's note
score 4