← Search

Chang Zeng

7 accepted papers

2026

A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation

ICML 2026poster

Query-based universal sound separation is fundamental to intelligent auditory systems, aiming to isolate specific sources from mixtures. Despite recent advances, existing methods continue to suffer from residual interference in complex acoustic scenes. This performance limitation stems largely from …

Cited by 0SourcecodeScholar
2026

ComGS: Efficient 3D Object-Scene Composition via Surface Octahedral Probes

ICLR 2026poster

Gaussian Splatting (GS) enables immersive rendering, but realistic 3D object–scene composition remains challenging. Baked appearance and shadow information in GS radiance fields cause inconsistencies when combining objects and scenes. Addressing this requires relightable object reconstruction and sc…

Cited by 0SourcecodeScholar
2026

SpatialVID: A Large-Scale Video Dataset with Spatial Annotations

CVPR 2026

Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current models remain severely constrained by the scarcity of large-scale, high-quality training data. While several datasets pr

Cited by 0SourcecodeScholar
2025

SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios

ICLR 2025poster

Systematic evaluation of speech separation and enhancement models under moving sound source conditions requires extensive and diverse data. However, real-world datasets often lack sufficient data for training and evaluation, and synthetic datasets, while larger, lack acoustic realism. Consequently,…

2023

Cross-Modal Audio-Visual Co-Learning for Text-Independent Speaker Verification

ICASSP 2023accepted

Visual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learning paradigm. The primary motivation of our cross-modal co-learning method is mo…

Cited by 0SourceScholar
2023

SSI-Net: A Multi-Stage Speech Signal Improvement System for ICASSP 2023 SSI Challenge

ICASSP 2023accepted

The ICASSP 2023 Speech Signal Improvement (SSI) Challenge concentrates on improving the speech signal quality of real-time communication (RTC) systems. In this paper, we introduce the speech signal improvement network (SSI-Net) submitted to the ICASSP 2023 SSI Challenge, which satisfies the real-tim…

Cited by 5SourceScholar
2022

Attention Back-End for Automatic Speaker Verification with Multiple Enrollment Utterances

ICASSP 2022accepted

Probabilistic linear discriminant analysis (PLDA) or cosine similarity have been widely used in traditional speaker verification systems as back-end techniques to measure pairwise similarities. To make better use of multiple enrollment utterances, we propose a novel attention back-end model that can…

Cited by 0SourceScholar