← Search

Ziteng Wang

19 accepted papers

2026

SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear Attention

ICLR 2026poster

In Diffusion Transformer (DiT) models, particularly for video generation, attention latency is a major bottleneck due to the long sequence length and the quadratic complexity. Interestingly, we find that attention weights can be decoupled into two matrices: a small fraction of large weights with hig…

Cited by 44SourcecodeScholar
2025

FACET: Fast and Accurate Event-Based Eye Tracking Using Ellipse Modeling for Extended Reality

ICRA 2025

Eye tracking is a key technology for gaze-based interactions in Extended Reality (XR), but traditional frame-based systems struggle to meet XR's demands for high accuracy, low latency, and power efficiency. Event cameras offer a promising alternative due to their high temporal resolution and low pow

Cited by 9SourcecodeScholar
2025

VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction-Editing Data and Long Captions

NeurIPS 2025poster

Despite the success of Vision-Language Models (VLMs) like CLIP in aligning vision and language, their proficiency in detailed, fine-grained visual comprehension remains a key challenge. We present CLIP-IN, a novel framework that bolsters CLIP's fine-grained perception through two core innovations. F…

Cited by 0SourcecodeScholar
2024

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

NeurIPS 2024oral

We introduce Cambrian-1, a family of multimodal LLMs (MLLMs) designed with a vision-centric approach. While stronger language models can enhance multimodal capabilities, the design choices for vision components are often insufficiently explored and disconnected from visual representation learning re…

2023

Self-Paced Learning Based Graph Convolutional Neural Network for Mixed Integer Programming (Student Abstract)

AAAI 2023technical

Graph convolutional neural network (GCN) based methods have achieved noticeable performance in solving mixed integer programming problems (MIPs). However, the generalization of existing work is limited due to the problem structure. This paper proposes a self-paced learning (SPL) based GCN network (S…

Cited by 3SourcePDFScholar
2023

The WHU-Alibaba Audio-Visual Speaker Diarization System for the MISP 2022 Challenge

ICASSP 2023accepted

This paper describes the system developed by the WHU-Alibaba team for the Multimodal Information Based Speech Processing (MISP) 2022 Challenge. We extend the Sequence-to-Sequence Target-Speaker Voice Activity Detection framework to simultaneously detect multiple speakers’ voice activities from audio…

Cited by 0SourceScholar
2022

Causality Inspired Representation Learning for Domain Generalization

CVPR 2022oral

Domain generalization (DG) is essentially an out-of-distribution problem, aiming to generalize the knowledge learned from multiple source domains to an unseen target domain. The mainstream is to leverage statistical models to model the dependence between data and labels, intending to learn represent…

Cited by 213PDFcodeScholar
2022

Check and Link: Pairwise Lesion Correspondence Guides Mammogram Mass Detection

ECCV 2022poster

"Detecting mass in mammogram is significant due to the high occurrence and mortality of breast cancer. In mammogram mass detection, modeling pairwise lesion correspondence explicitly is particularly important. However, most of the existing methods build relatively coarse correspondence and have not…

Cited by 6SourcePDFScholar
2022

Multi-Task Deep Residual Echo Suppression with Echo-Aware Loss

ICASSP 2022accepted

This paper introduces the NWPU Team’s entry to the ICASSP 2022 AEC Challenge. We take a hybrid approach that cascades a linear AEC with a neural post-filter. The former is used to deal with the linear echo components while the latter suppresses the residual non-linear echo components. We use gated c…

Cited by 0SourceScholar
2022

NN3A: Neural Network Supported Acoustic Echo Cancellation, Noise Suppression and Automatic Gain Control for Real-Time Communications

ICASSP 2022accepted

Acoustic echo cancellation (AEC), noise suppression (NS) and automatic gain control (AGC) are three often required modules for real-time communications (RTC). This paper proposes a neural network supported algorithm for RTC, namely NN3A, which incorporates an adaptive filter and a multi-task model f…

Cited by 0SourceScholar
2021

Weighted Recursive Least Square Filter and Neural Network Based Residual ECHO Suppression for the AEC-Challenge

ICASSP 2021accepted

This paper presents a real-time Acoustic Echo Cancellation (AEC) algorithm submitted to the AEC-Challenge. The algorithm consists of three modules: Generalized Cross-Correlation with PHAse Transform (GCC-PHAT) based time delay compensation, weighted Recursive Least Square (wRLS) based linear adaptiv…

Cited by 0SourceScholar
2020

A Visual-Pilot Deep Fusion for Target Speech Separation in Multitalker Noisy Environment

ICASSP 2020accepted

Separating the target speech in multi-talker noisy environment is a challenging problem for audio-only source separation algorithms. The major problem behind is that the separated speech from the same talker can switch among the outputs across consecutive segments, causing the talker permutation iss…

Cited by 5SourceScholar
2018

On SDW-MWF and Variable Span Linear Filter with Application to Speech Recognition in Noisy Environments

ICASSP 2018accepted

Neural network based spectral mask estimation for acoustic beamforming, which consists of linear filtering and mask estimation, has shown to be a promising approach for robust speech recognition in noisy environments. Nevertheless, few improvements are made on the linear filtering. In this paper, we…

Cited by 0SourceScholar
2018

Semi-Supervised Learning with Deep Neural Networks for Relative Transfer Function Inverse Regression

ICASSP 2018accepted

Prior knowledge of the relative transfer function (RTF) is useful in many applications but remains little studied. In this paper, we propose a semi-supervised learning algorithm based on deep neural networks (DNNs) for RTF inverse regression, that is to generate the full-band RTF vector directly fro…

Cited by 0SourceScholar
2015

Fast Second Order Stochastic Backpropagation for Variational Inference

NeurIPS 2015poster

We propose a second-order (Hessian or Hessian-free) based optimization method for variational inference inspired by Gaussian backpropagation, and argue that quasi-Newton optimization can be developed as well. This is accomplished by generalizing the gradient computation in stochastic backpropagatio…

Cited by 52SourcePDFScholar