← Search

Junfeng Gong

1 accepted papers

2026

From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph

ICLR 2026poster

Despite significant evolution of CUDA programming and domain-specific libraries, effectively utilizing GPUs with massively parallel engines remains difficult. Large language models (LLMs) show strong potential in generating optimized CUDA code from sequential code. However, using LLMs in practice fa…

Cited by 0SourceScholar