← Search

Guangming Tan

8 accepted papers

2026

JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials

ICML 2026poster

Discovering atom-level phenomena requires molecular dynamics (MD) simulations with ab initio accuracy. Machine learning interatomic potentials (MLIPs) enable stable, high-accuracy MD simulations, and their models exhibit scaling-law trends similar to large language models. However, the lack of scala…

Cited by 0SourceScholar
2026

Mastering Sparse CUDA Generation through Pretrained Models and Deep Reinforcement Learning

ICLR 2026oral

Code generation is a crucial research area in the field of artificial intelligence, holding the potential to revolutionize software development and streamline programming processes. However, generating the high-performance code, which need to be executed in a shorter time for the low-latency scenari…

Cited by 0SourcecodeScholar
2026

MatRIS: Toward Reliable and Efficient Pretrained Machine Learning Interaction Potentials

ICLR 2026poster

Universal MLIPs (uMLIPs) demonstrate broad applicability across diverse material systems and have emerged as a powerful and transformative paradigm in chemical and computational materials science. Equivariant uMLIPs achieve state-of-the-art accuracy in a wide range of benchmarks by incorporating equ…

Cited by 0SourceScholar
2026

RCMoE: A Communication-Efficient Random Compression Framework for Resource-Constrained Mixture-of-Experts Training

AAAI 2026technical

Mixture-of-Experts (MoE) architecture with experts parallelism scales LLMs efficiently by activating only a subset of experts per input, avoiding proportional training costs. However, the intensive and heterogeneous communication substantially hinders the efficiency and scalability of MoE training i

Cited by 0SourcePDFScholar
2025

ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism

NeurIPS 2025oral

Multimodal large language models (MLLMs) extend LLMs to handle images, videos, and audio by incorporating feature extractors and projection modules. However, these additional components—combined with complex inference pipelines and heterogeneous workloads—introduce significant inference overhead. Th…

Cited by 0SourceScholar
2023

RLEKF: An Optimizer for Deep Potential with Ab Initio Accuracy

AAAI 2023technical

It is imperative to accelerate the training of neural network force field such as Deep Potential, which usually requires thousands of images based on first-principles calculation and a couple of days to generate an accurate potential energy surface. To this end, we propose a novel optimizer named re…

Cited by 5SourcePDFScholar