Megadose AI progress, ranked and analyzed.

Post-Training VLMs for Video Mistake Detection

· ArXiv · AI/CL/LG ·
The paper tests mistake detection as a video-question answering problem, including actions the model has not seen before.

The authors introduce MD-VQA, a protocol and benchmark for judging whether a video step matches its written instruction. They argue that closed-set mistake detectors are too tied to specific procedures, because new tasks can require new data and retraining. Their post-training method uses a reward function aimed at spotting mismatches between the instruction and the video. In evaluations, it beats zero-shot, supervised fine-tuning, and other post-training baselines, with up to an 11.6% gain on unseen procedures in EP-VQA. ArXiv · AI/CL/LG's note

score 4

Categories: Research