← Search

Mingyu Gao

6 accepted papers

2026

Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance

CVPR 2026

Multimodal continual instruction tuning enables multimodal large language models to sequentially adapt to new tasks while building upon previously acquired knowledge. However, this continual learning paradigm faces the significant challenge of catastrophic forgetting, where learning new tasks leads

Cited by 0SourcecodeScholar
2025

Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning

NeurIPS 2025spotlight

Leveraging attention sparsity to accelerate long-context large language models (LLMs) has been of great importance recently. However, most existing sparse attention algorithms use a fixed budget of how many tokens to use in their computations. This simple static decision raises critical issues in re…

Cited by 0SourceScholar
2024

Seesaw: Compensating for Nonlinear Reduction with Linear Computations for Private Inference

ICML 2024poster

With increasingly serious data privacy concerns and strict regulations, privacy-preserving machine learning (PPML) has emerged to securely execute machine learning tasks without violating privacy. Unfortunately, the computational cost to securely execute nonlinear computations in PPML remains signif…

Cited by 5SourcePDFScholar
2023

GLT-T: Global-Local Transformer Voting for 3D Single Object Tracking in Point Clouds

AAAI 2023technical

Current 3D single object tracking methods are typically based on VoteNet, a 3D region proposal network. Despite the success, using a single seed point feature as the cue for offset learning in VoteNet prevents high-quality 3D proposals from being generated. Moreover, seed points with different impor…

2023

OSP2B: One-Stage Point-to-Box Network for 3D Siamese Tracking

IJCAI 2023poster

Two-stage point-to-box network acts as a critical role in the recent popular 3D Siamese tracking paradigm, which first generates proposals and then predicts corresponding proposal-wise scores. However, such a network suffers from tedious hyper-parameter tuning and task misalignment, limiting the tra…

2023

ST${2}$: Spatial-Temporal State Transformer for Crowd-Aware Autonomous Navigation

RA-L 2023

Empowering an intelligent agent with the ability of autonomous navigation in complex and dynamic environments is an important and active research topic in embodied artificial intelligence. In this letter, we address this challenging task from the view of exploiting both the spatial and temporal stat

Cited by 34SourceScholar