← Search

Weituo Hao

8 accepted papers

2024

InstructME: An Instruction Guided Music Edit Framework with Latent Diffusion Models

IJCAI 2024poster

Music editing primarily entails the modification of instrument tracks or remixing in the whole, which offers a novel reinterpretation of the original piece through a series of operations. These music processing methods hold immense potential across various applications but demand substantial experti…

2023

Estimating Total Correlation with Mutual Information Estimators

AISTATS 2023poster

Total correlation (TC) is a fundamental concept in information theory that measures statistical dependency among multiple random variables. Recently, TC has shown noticeable effectiveness as a regularizer in many learning tasks, where the correlation among multiple latent embeddings requires to be j…

2023

Mitigating Test-Time Bias for Fair Image Retrieval

NeurIPS 2023poster

We address the challenge of generating fair and unbiased image retrieval results given neutral textual queries (with no explicit gender or race connotations), while maintaining the utility (performance) of the underlying vision-language (VL) model. Previous methods aim to disentangle learned represe…

2021

FairFil: Contrastive Neural Debiasing Method for Pretrained Text Encoders

ICLR 2021poster

Pretrained text encoders, such as BERT, have been applied increasingly in various natural language processing (NLP) tasks, and have recently demonstrated significant performance gains. However, recent studies have demonstrated the existence of social bias in these pretrained NLP models. Although pri…

Cited by 130SourcePDFScholar
2021

Improving Zero-Shot Voice Style Transfer via Disentangled Representation Learning

ICLR 2021poster

Voice style transfer, also called voice conversion, seeks to modify one speaker's voice to generate speech as if it came from another (target) speaker. Previous works have made progress on voice conversion with parallel training data and pre-known speakers. However, zero-shot voice style transfer, w…

Cited by 75SourcePDFScholar
2021

MixKD: Towards Efficient Distillation of Large-scale Language Models

ICLR 2021poster

Large-scale language models have recently demonstrated impressive empirical performance. Nevertheless, the improved results are attained at the price of bigger models, more power consumption, and slower inference, which hinder their applicability to low-resource (both memory and computation) platfor…

Cited by 90SourcePDFScholar
2020

CLUB: A Contrastive Log-ratio Upper Bound of Mutual Information

ICML 2020poster

Mutual information (MI) minimization has gained considerable interests in various machine learning tasks. However, estimating and minimizing MI in high-dimensional spaces remains a challenging problem, especially when only samples, rather than distribution forms, are accessible. Previous works mainl…

2020

Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-Training

CVPR 2020poster

Learning to navigate in a visual environment following natural-language instructions is a challenging task, because the multimodal inputs to the agent are highly variable, and the training data on a new task is often limited. In this paper, we present the first pre-training and fine-tuning paradigm…

Cited by 320PDFcodeScholar