← Search

Lie Lu

9 accepted papers

2026

A robust PPG foundation model using multimodal physiological supervision

ICML 2026poster

Photoplethysmography (PPG), a non-invasive measure of changes in blood volume, is widely used in both wearable devices and clinical settings. Recent PPG foundation models either use open-source ICU datasets with pretraining paradigms that require high-quality data and thus complicate generalization …

Cited by 0SourceScholar
2025

RapVerse: Coherent Vocals and Whole-Body Motion Generation from Text

ICCV 2025poster

In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existing works that typically address these two modalities in isolation. To facilitate this, we first collect the RapVerse dat…

Cited by 0SourcePDFScholar
2025

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval

ICASSP 2025accepted

Content creators often use music to enhance their videos, from soundtracks in movies to background music in video blogs and social media content. However, identifying the best music for a video can be a difficult and time-consuming task. To address this challenge, we propose a novel framework for au…

Cited by 0SourceScholar
2025

XAttnMark: Learning Robust Audio Watermarking with Cross-Attention

ICML 2025poster

The rapid proliferation of generative audio synthesis and editing technologies has raised significant concerns about copyright infringement, data provenance, and the spread of misinformation through deepfake audio. Watermarking offers a proactive solution by embedding imperceptible, identifiable, an…

Cited by 1SourcePDFScholar
2024

Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and Videos

ICASSP 2024accepted

A "match cut" is a common video editing technique where a pair of shots that have a similar composition transition fluidly from one to another. Although match cuts are often visual, certain match cuts involve the fluid transition of audio, where sounds from different sources merge into one indisting…

Cited by 0SourceScholar
2023

Deepspace: Dynamic Spatial and Source CUE Based Source Separation for Dialog Enhancement

ICASSP 2023accepted

Dialog Enhancement (DE) is a feature which allows a user to increase the level of dialog in TV or movie content relative to nondialog sounds. When only the original mix is available, DE is "unguided," and requires source separation. In this paper, we describe the DeepSpace system, which performs sou…

Cited by 0SourceScholar
2023

High Quality Audio Coding with Mdctnet

ICASSP 2023accepted

We propose a neural audio generative model, MDCTNet, operating in the perceptually weighted domain of an adaptive modified discrete cosine transform (MDCT). The architecture of the model captures correlations in both time and frequency directions with recurrent layers (RNNs). An audio coding system…

Cited by 0SourceScholar