每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

arXiv:2608.173792026-08-17

We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads,

作者 · Genghan Zhang

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道