DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
The benchmark’s best tested setup only reached 66.34% accuracy on verifiable analytics tasks spanning mixed files, databases, documents, and video.
DataSpace contains 410 cross-language tasks over 7,439 artifacts totaling 15.01 GB. Agents get only a question and a local workspace, then must return the full requested table. The paper says harness choice alone produced a 15.36-point accuracy spread with the same model backbone. Multimodal evidence integration and joins reduced accuracy across all six tested backbones. HF Daily Papers' note
DataSpace contains 410 cross-language tasks over 7,439 artifacts totaling 15.01 GB. Agents get only a question and a local workspace, then must return the full requested table. The paper says harness choice alone produced a 15.36-point accuracy spread with the same model backbone. Multimodal evidence integration and joins reduced accuracy across all six tested backbones. HF Daily Papers' note
score 5