PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
PTXBench finds current LLMs still unreliable at turning architecture-specific PTX into faster GPU kernels.
The benchmark tests correctness, runtime use of target instructions, and speedups against frontier libraries on GEMM and attention workloads for H100 and B200 GPUs. The paper says performance drops sharply on harder attention backward tasks, and even using the intended PTX instructions does not guarantee competitive speed. No evaluated model consistently matches the libraries across the suite. Fine-tuning Qwen3.6-27B helped on some repair-conditioned tasks, but generalization stayed uneven. HF Daily Papers' note
The benchmark tests correctness, runtime use of target instructions, and speedups against frontier libraries on GEMM and attention workloads for H100 and B200 GPUs. The paper says performance drops sharply on harder attention backward tasks, and even using the intended PTX instructions does not guarantee competitive speed. No evaluated model consistently matches the libraries across the suite. Fine-tuning Qwen3.6-27B helped on some repair-conditioned tasks, but generalization stayed uneven. HF Daily Papers' note
score 5