← Search

Lisha Chen

14 accepted papers

2026

BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition

ICASSP 2026oral

Speech is a rich signal, and labeled audio-text pairs are costly, making self-supervised learning essential for scalable representation learning. A core challenge in speech SSL is generating pseudo-labels that are both informative and efficient: strong labels, such as those used in HuBERT, improve d…

Cited by 0SourcePDFScholar
2025

Beyond Value Functions: Single-Loop Bilevel Optimization under Flatness Conditions

NeurIPS 2025poster

Bilevel optimization, a hierarchical optimization paradigm, has gained significant attention in a wide range of practical applications, notably in the fine-tuning of generative models. However, due to the nested problem structure, most existing algorithms require either the Hessian vector calculatio…

Cited by 0SourceScholar
2025

Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference Guidance

ICML 2025spotlight

Multi-objective learning under user-specified preference is common in real-world problems such as multi-lingual speech recognition under fairness. In this work, we frame such a problem as a semivectorial bilevel optimization problem, whose goal is to optimize a pre-defined preference function, subje…

Cited by 0SourcePDFScholar
2025

Objective Soups: Multilingual Multi-Task Modeling for Speech Processing

NeurIPS 2025poster

The need for training multilingual multi-task speech processing (MSP) models that perform both automatic speech recognition and speech-to-text translation is increasingly evident. However, a significant challenge arises from the conflicts among multiple objectives when using a single model. Multi-ob…

Cited by 0SourceScholar
2024

FERERO: A Flexible Framework for Preference-Guided Multi-Objective Learning

NeurIPS 2024poster

Finding specific preference-guided Pareto solutions that represent different trade-offs among multiple objectives is critical yet challenging in multi-objective problems. Existing methods are restrictive in preference definitions and/or their theoretical guarantees. In this work, we introduce a Fle…

2024

Variance Reduction Can Improve Trade-Off in Multi-Objective Learning

ICASSP 2024accepted

Many machine learning problems today have multiple objective functions, which are often tackled by the multi-objective learning (MOL) framework. Albeit many encouraging results are obtained by MOL algorithms, a recent theoretical study [1] revealed that these gradient-based MOL methods (e.g., MGDA,…

Cited by 0SourceScholar
2023

Three-Way Trade-Off in Multi-Objective Learning: Optimization, Generalization and Conflict-Avoidance

NeurIPS 2023poster

Multi-objective learning (MOL) often arises in emerging machine learning problems when multiple learning criteria or tasks need to be addressed. Recent works have developed various _dynamic weighting_ algorithms for MOL, including MGDA and its variants, whose central idea is to find an update direc…

2022

Is Bayesian Model-Agnostic Meta Learning Better than Model-Agnostic Meta Learning, Provably?

AISTATS 2022poster

Meta learning aims at learning a model that can quickly adapt to unseen tasks. Widely used meta learning methods include model agnostic meta learning (MAML), implicit MAML, Bayesian MAML. Thanks to its ability of modeling uncertainty, Bayesian MAML often has advantageous empirical performance. Howev…

2022

Sharp-MAML: Sharpness-Aware Model-Agnostic Meta Learning

ICML 2022spotlight

Model-agnostic meta learning (MAML) is currently one of the dominating approaches for few-shot meta-learning. Albeit its effectiveness, the optimization of MAML can be challenging due to the innate bilevel problem structure. Specifically, the loss landscape of MAML is much more complex with possibly…

2021

Uncertain Graph Neural Networks for Facial Action Unit Detection

AAAI 2021technical

Capturing the dependencies among different facial action units (AU) is extremely important for the AU detection task. Many studies have employed graph-based deep learning methods to exploit the dependencies among AUs. However, the dependencies among AUs in real world data are often noisy and the unc…

Cited by 88SourcePDFScholar