← Search

Hantao Hua

3 accepted papers

2026

Gram2Token: Enabling Run-time GPU-Native Grammar-Constrained Decoding for LLMs

ICML 2026poster

Grammar-constrained decoding is essential for enabling large language models (LLMs) to efficiently generate structured outputs in applications, such as JSON objects for parameter passing. Existing approaches typically execute grammar constraint masking on the CPU, while LLM inference is performed on…

Cited by 0SourceScholar
2026

Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference

IJCAI 2026

In simulation-in-the-loop decision-making systems, reinforcement learning (RL) inference is often constrained by simulator-side execution overhead, where workloads are highly dynamic and sensitive to runtime thread configurations. Existing multithreaded strategies struggle to match thread resources

Cited by 0Scholar