← Search

HongXin Xu

3 accepted papers

2025

DynaPipe: Dynamic Layer Redistribution for Efficient Serving of LLMs with Pipeline Parallelism

NeurIPS 2025poster

To accelerate large language model (LLM) inference, pipeline parallelism partitions model layers into sequential stages, each assigned to a different device for concurrent execution. However, this method often suffers from pipeline bubbles caused by imbalanced computation in the tail stage. While up…

Cited by 0SourceScholar
2024

Enhancing 3D Single Object Tracking with Efficient Point Cloud Segmentation

IROS 2024poster

3D single object tracking (SOT) based on point cloud has attracted much attention due to its important role in machine vision and autonomous driving. Recently, M2-Track proposes a two-stage tracking structure centered on motion, but they ignore the effect of segmentation errors in sparse point cloud…

Cited by 0SourceScholar