Megadose AI progress, ranked and analyzed.

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

· HF Daily Papers ·
The benchmark makes embodied MLLMs act through sudden household hazards, then scores the physical consequences.

ReactHuman uses more than 1,000 reproducible simulated scenes across 17 event families, with ground truth from 240 Hz rigid-body simulation. The authors include adversarial objects, such as a foam anvil or steel apple, to test whether models follow physics rather than appearance. In tests of seven multimodal LLMs, models mishandled roughly one in three hazards and often missed interception points even when choosing the right kind of action. HF Daily Papers' note

score 5

Categories: Research