← Search

Guangrui Ma

2 accepted papers

2026

THETA: Threshold-Based Exclusive Batching for Memory-Bandwidth-Constrained LLM Inference

ICML 2026poster

Chunked prefill has become the dominant scheduling strategy for large language model (LLM) inference, interleaving prefill and decode operations to improve GPU utilization. However, this approach does not universally outperform exclusive batching: on bandwidth-constrained GPUs, mixed batches can int…

Cited by 0SourceScholar
2025

TARFVAE: Efficient One-Step Generative Time Series Forecasting via TARFLOW based VAE

NeurIPS 2025poster

Time series data is ubiquitous, with forecasting applications spanning from finance to healthcare. Beyond popular deterministic methods, generative models are gaining attention due to advancements in areas like image synthesis and video generation, as well as their inherent ability to provide probab…

Cited by 0SourcecodeScholar