← Search

Xiang Hao

7 accepted papers

2025

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding

NeurIPS 2025poster

Long video understanding is inherently challenging for vision-language models (VLMs) because of the extensive number of frames. With each video frame typically expanding into tens or hundreds of tokens, the limited context length of large language models (LLMs) forces the VLMs to perceive the frames…

Cited by 0SourceScholar
2023

Movies2Scenes: Using Movie Metadata To Learn Scene Representation

CVPR 2023poster

Understanding scenes in movies is crucial for a variety of applications such as video moderation, search, and recommendation. However, labeling individual scenes is a time-consuming process. In contrast, movie level metadata (e.g., genre, synopsis, etc.) regularly gets produced as part of the film p…

Cited by 19SourcePDFScholar
2023

Two-Stage Neural Network for ICASSP 2023 Speech Signal Improvement Challenge

ICASSP 2023accepted

In ICASSP 2023 speech signal improvement challenge, we developed a dual-stage neural model which improves speech signal quality induced by different distortions in a stage-wise divide-and-conquer fashion. Specifically, in the first stage, the speech improvement network focuses on recovering the miss…

Cited by 0SourceScholar
2021

Fullsubnet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement

ICASSP 2021accepted

This paper proposes a full-band and sub-band fusion model, named as FullSubNet, for single-channel real-time speech enhancement. Full-band and sub-band refer to the models that input full-band and sub-band noisy spectral feature, output full-band and sub-band speech target, respectively. The sub-ban…

Cited by 0SourceScholar
2020

Masking and Inpainting: A Two-Stage Speech Enhancement Approach for Low SNR and Non-Stationary Noise

ICASSP 2020accepted

Currently, low signal-to-noise ratio (SNR) and non-stationary noise cause severe performance degradation for most of speech enhancement models. For better speech enhancement at the above scenarios, this paper proposes a two-stage approach that consists of binary masking and spectrogram inpainting. I…

Cited by 0SourceScholar
2020

Time-Domain Neural Network Approach for Speech Bandwidth Extension

ICASSP 2020accepted

In this paper, we study the time-domain neural network approach for speech bandwidth extension. We propose a network architecture, named multi-scale fusion neural network (MfNet), that gradually restores the low-frequency signal and predicts the high-frequency signal through the exchange of informat…

Cited by 0SourceScholar
2019

An Attention-based Neural Network Approach for Single Channel Speech Enhancement

ICASSP 2019accepted

This paper proposes an attention-based neural network approach for single channel speech enhancement. Our work is inspired by the recent success of attention models in sequence-to-sequence learning. It is intuitive to use attention mechanism in speech enhancement as humans are able to focus on the i…

Cited by 56SourceScholar