Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories
Embedding models found the right structural match nearby, then ranked the more literal match first.
The paper tests retrieval with meaning and wording deliberately separated in competition math and embodied-agent trajectories. In math, strict Hit@1 at the hardest disguise tier was 0.0% for both production embedders, even though the correct item was usually still in the top 10. In trajectories, models fell to chance or below chance when the matching task required different objects or receptacles. LLM rerankers recovered part of the gap, but the authors say some math gains appear tied to memorized well-known competitions. ArXiv · AI/CL/LG's note
The paper tests retrieval with meaning and wording deliberately separated in competition math and embodied-agent trajectories. In math, strict Hit@1 at the hardest disguise tier was 0.0% for both production embedders, even though the correct item was usually still in the top 10. In trajectories, models fell to chance or below chance when the matching task required different objects or receptacles. LLM rerankers recovered part of the gap, but the authors say some math gains appear tied to memorized well-known competitions. ArXiv · AI/CL/LG's note
score 5