MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
The paper introduces a 30,000-sample benchmark for testing business-idea agents on image-plus-text inputs.
MBA-Bench spans six domains where visual cues are meant to matter, not just captions. The authors generate reference ideas using GPT-4o with retrieval and evidence-augmented synthesis, then judge agents on six business criteria. They also train two agents, MBA-b and MBA-k, for blind and disclosed-criteria settings, reporting gains over caption-only and multimodal baselines. HF Daily Papers' note
MBA-Bench spans six domains where visual cues are meant to matter, not just captions. The authors generate reference ideas using GPT-4o with retrieval and evidence-augmented synthesis, then judge agents on six business criteria. They also train two agents, MBA-b and MBA-k, for blind and disclosed-criteria settings, reporting gains over caption-only and multimodal baselines. HF Daily Papers' note
score 4