← Search

Xinwei Qiang

2 accepted papers

2026

DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training

ICLR 2026poster

Determinism is indispensable for reproducibility in large language model (LLM) training, yet it often exacts a steep performance cost. In widely used attention implementations such as FlashAttention-3, the deterministic backward pass can incur up to a 37.9% throughput reduction relative to its non‑d…

Cited by 0SourcecodeScholar
2026

TritonGym: A Benchmark for Agentic LLM Workflows in Triton GPU Code Generation

ICML 2026poster

Large language models (LLMs) can already draft plausible Triton kernels, yet most existing evaluations still focus on single-shot generation and underplay tool use and feedback. We introduce *TritonGym*, a benchmark and orchestration framework for evaluating agentic workflows in GPU code generation.…

Cited by 0SourceScholar