MintAct: A Unified Visual Agent for Digital Environments
MintAct is pitched as one model family for UI grounding, navigation, and visual tool use across mobile, desktop, and web.
The paper says the models were trained at 2B, 4B, and 8B scales and are meant to match domain-specific agents across those tasks. Its infrastructure runs hundreds of concurrent digital-environment instances for trajectory collection and online reinforcement learning. The authors also describe an asynchronous RL setup designed to control cross-domain training mix and handle noisy feedback. They report state-of-the-art results, including 48.9 on OSWorld-Verified, at comparable model sizes. HF Daily Papers' note
The paper says the models were trained at 2B, 4B, and 8B scales and are meant to match domain-specific agents across those tasks. Its infrastructure runs hundreds of concurrent digital-environment instances for trajectory collection and online reinforcement learning. The authors also describe an asynchronous RL setup designed to control cross-domain training mix and handle noisy feedback. They report state-of-the-art results, including 48.9 on OSWorld-Verified, at comparable model sizes. HF Daily Papers' note
score 5