Megadose Built for builders and researchers.

Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels

· ArXiv · AI/CL/LG ·
PixelTriage tries to spend full image tokens only on the retrieved memories likely to matter.

The paper says most of the value from opening images comes from one or two retrieved memories, not the whole set. Its plug-in reads the dialogue, a short note, and a thumbnail before the answering model runs, then predicts which memories deserve pixels. In tests with a 7B answering model, it used 11-23% of the visual tokens without a significant accuracy loss, and on DMV answered 2.9 times faster than opening all images. ArXiv · AI/CL/LG's note

score 5

Categories: Research