← Search

Mu Yang

7 accepted papers

2025

MCTrack: A Unified 3D Multi-Object Tracking Framework for Autonomous Driving

IROS 2025

This paper introduces MCTrack, a new 3D multi-object tracking method that achieves performance across KITTI, nuScenes, and Waymo datasets. Addressing the gap in existing tracking paradigms, which often perform well on specific datasets but lack generalizability, MCTrack offers a unified solution. Ad

Cited by 27SourcecodeScholar
2025

UniScene: Unified Occupancy-centric Driving Scene Generation

CVPR 2025poster

Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles…

2024

Diarist: Streaming Speech Translation with Speaker Diarization

ICASSP 2024accepted

End-to-end speech translation (ST) for conversation recordings involves several under-explored challenges such as speaker diarization (SD) without accurate word time stamps and handling of overlapping speech in a streaming fashion. In this work, we propose DiariST, the first streaming ST and SD solu…

Cited by 0SourceScholar
2023

Learning ASR Pathways: A Sparse Multilingual ASR Model

ICASSP 2023accepted

Neural network pruning compresses automatic speech recognition (ASR) models effectively. However, in multilingual ASR, language-agnostic pruning may lead to severe performance drops on some languages because language-agnostic pruning masks may not fit all languages and discard important language-spe…

Cited by 0SourceScholar
2022

Joint Hypoglycemia Prediction and Glucose Forecasting via Deep Multi-Task Learning

ICASSP 2022accepted

We present a multitask learning approach to the problem of hypoglycemia (HG) prediction in diabetes. The approach is based on a state-of-the-art time series forecasting model, N-BEATS, and extends it by adding a classification task so that the model performs both glucose forecasting (i.e., predictin…

Cited by 0SourceScholar
2022

Towards Lifelong Learning of Multilingual Text-to-Speech Synthesis

ICASSP 2022accepted

This work presents a lifelong learning approach to train a multilingual Text-To-Speech (TTS) system, where each language was seen as an individual task and was learned sequentially and continually. It does not require pooled data from all languages altogether, and thus alleviates the storage and com…

Cited by 0SourceScholar
2021

EventPlus: A Temporal Event Understanding Pipeline

NAACL 2021system demonstrations

We present EventPlus, a temporal event understanding pipeline that integrates various state-of-the-art event understanding components including event trigger and type detection, event argument detection, event duration and temporal relation extraction. Event information, especially event temporal kn…