Megadose Built for builders and researchers.

D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

· ArXiv · AI/CL/LG ·
D2K-Bench tests whether agents can turn supplied expert GPU-design guidance into faster, correct kernels.

The benchmark covers 26 tasks and 85 workloads, with guidance split into algorithmic insight, dataflow design, and low-level optimization. In paired runs on NVIDIA B200 GPUs, correctness rose from 93.1% to 98.5%, and the overall Performance Score moved from 1.46 to 1.95. For three frontier models that solved all tasks in both settings, geometric mean speedup increased from 1.69x to 2.49x. The authors say the gains show expert guidance helps, but generated code still leaves some design properties unimplemented. ArXiv · AI/CL/LG's note

score 6

Categories: Research