← Search

ShiLiang Zhang

65 accepted papers

2026

DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation

CVPR 2026

Pixel diffusion aims to generate images directly in pixel space in an end-to-end fashion. This approach avoids the limitations of VAE in the two-stage latent diffusion, offering higher model capacity. Existing pixel diffusion models suffer from slow training and inference, as they usually model both

Cited by 0SourcecodeScholar
2026

MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling

ICLR 2026poster

Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, long-range dynamics, and fine-grained human–environment interactions. This often leads to unrealistic or physically implaus…

Cited by 0SourceScholar
2026

SCAN: Self-Calibrated AutoregressioN for High-Quality Visual Generation

AAAI 2026technical

Human artists can continuously refine their coarse sketches during artistic creation. This is quite different from existing autoregressive generation, where a token is determined once sampled. Aiming to flexibly refine the generated contents, this paper presents a Self-Calibrated AutoregressioN (SCA

Cited by 0SourcePDFScholar
2026

When Person Re-Identification Meets Event Camera: A Benchmark Dataset and an Attribute-Guided Re-Identification Framework

AAAI 2026technical

Recent researchers have proposed using event cameras for person re-identification (ReID) due to their promising performance and better balance in terms of privacy protection, event camera-based person ReID has attracted significant attention. Currently, mainstream event-based person ReID algorithms

Cited by 0SourcePDFScholar
2025

3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization

ICASSP 2025accepted

We introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the combined strengths of acoustic, semantic, and visual data, se…

Cited by 0SourceScholar
2025

Efficient Multi-modal Long Context Learning for Training-free Adaptation

ICML 2025poster

Traditional approaches to adapting multi-modal large language models (MLLMs) to new tasks have relied heavily on fine-tuning. This paper introduces Efficient Multi-Modal Long Context Learning (EMLoC), a novel training-free alternative that embeds demonstration examples directly into the model input.…

2025

Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap

ICASSP 2025accepted

While automatic speech recognition (ASR) systems have achieved remarkable performance with large-scale datasets, their efficacy remains inadequate in low-resource settings, encompassing dialects, accents, minority languages, and long-tail hotwords, domains with significant practical relevance. With…

Cited by 0SourceScholar
2025

Generalizable Object Keypoint Localization from Generative Priors

CVPR 2025poster

Generalizable object keypoint localization is a fundamental computer vision task in understanding the object structure. It is challenging for existing keypoint localization methods because their limited training data cannot provide generalizable shape and semantic cues, leading to inferior performan…

Cited by 0SourcePDFScholar
2025

MV-VTON: Multi-View Virtual Try-On with Diffusion Models

AAAI 2025technical

The goal of image-based virtual try-on is to generate an image of the target person naturally wearing the given clothing. However, existing methods solely focus on the frontal try-on using the frontal clothing. When the views of the clothing and person are significantly inconsistent, particularly wh…

2025

MagCache: Fast Video Generation with Magnitude-Aware Cache

NeurIPS 2025poster

Existing acceleration techniques for video diffusion models often rely on uniform heuristics or time-embedding variants to skip timesteps and reuse cached features. These approaches typically require extensive calibration with curated prompts and risk inconsistent outputs due to prompt-specific over…

Cited by 0SourceScholar
2025

NN-Former: Rethinking Graph Structure in Neural Architecture Representation

CVPR 2025poster

The growing use of deep learning necessitates efficient network design and deployment, making neural predictors vital for estimating attributes such as accuracy and latency. Recently, Graph Neural Networks (GNNs) and transformers have shown promising performance in representing neural architectures.…

2025

OmniAudio: Generating Spatial Audio from 360-Degree Video

ICML 2025poster

Traditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, \textbf{360V2SA}, to generate spa…

2025

OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

ACL 2025long

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and natural interactions in full-duplex dialogue systems remains a sig…

2025

Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision

ICASSP 2025accepted

Training speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persistent challenge. In this paper, we propose a novel self-supervised speaker verification approach, Self-Distillation Prototypes Network (SDPN), which effectively facilitates self-sup…

Cited by 0SourceScholar
2025

Speech Recognition Meets Large Language Model: Benchmarking, Models, and Exploration

AAAI 2025technical

In this paper, we focus on prompting one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Despite the growing body of research in this area, we find that many crucial design decis…

2025

UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook

ACL 2025long

The emergence of audio language models is empowered by neural audio codecs, which establish critical mappings between continuous waveforms and discrete tokens compatible with language model paradigms. The evolutionary trends from multi-layer residual vector quantizer to single-layer quantizer are be…

2025

UniSpeaker: A Unified Approach for Multimodality-driven Speaker Generation

EMNLP 2025

While recent advances in reference-based speaker cloning have significantly improved the authenticity of synthetic speech, speaker generation driven by multimodal cues such as visual appearance, textual descriptions, and other biometric signals remains in its early stages. To pioneer truly multimoda

2025

Unified Video Generation via Next-Set Prediction in Continuous Domain

ICCV 2025poster

Existing video generation strategies can be categorized into two categories, i.e., the diffusion and autoregressive (AR) methods. While AR methods achieves high efficiency by predicting the next token based on known visual cues, they generally fall short of diffusion models in terms of video quality…

Cited by 0SourcePDFScholar
2024

Decoupled Optimisation for Long-Tailed Visual Recognition

AAAI 2024technical

When training on a long-tailed dataset, conventional learning algorithms tend to exhibit a bias towards classes with a larger sample size. Our investigation has revealed that this biased learning tendency originates from the model parameters, which are trained to disproportionately contribute to the…

Cited by 6SourcePDFScholar
2024

FunCodec: A Fundamental, Reproducible and Integrable Open-Source Toolkit for Neural Speech Codec

ICASSP 2024accepted

This paper presents FunCodec, a fundamental neural speech codec toolkit, which is an extension of the open-source speech processing toolkit FunASR. FunCodec provides reproducible training recipes and inference scripts for the latest neural speech codec models, such as SoundStream and Encodec. Thanks…

Cited by 0SourceScholar
2024

Hourglass-AVSR: Down-Up Sampling-Based Computational Efficiency Model for Audio-Visual Speech Recognition

ICASSP 2024accepted

Recently audio-visual speech recognition (AVSR), which better leverages video modality as additional information to extend automatic speech recognition (ASR), has shown promising results in complex acoustic environments. However, there is still substantial space to improve as complex computation of…

Cited by 0SourceScholar
2024

Leveraging Speech PTM, Text LLM, And Emotional TTS For Speech Emotion Recognition

ICASSP 2024accepted

In this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. First, we investigated the representation ability of different speech self-supervi…

Cited by 0SourceScholar
2024

LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model

CVPR 2024highlight

The capacity of existing human keypoint localization models is limited by keypoint priors provided by the training data. To alleviate this restriction and pursue more general model this work studies keypoint localization from a different perspective by reasoning locations based on keypiont clues in…

2024

Loss Masking Is Not Needed In Decoder-Only Transformer For Discrete-Token-Based ASR

ICASSP 2024accepted

Recently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and use a shared vocabulary for both text and speech tokens. Then they train a singl…

Cited by 0SourceScholar
2024

OVMR: Open-Vocabulary Recognition with Multi-Modal References

CVPR 2024poster

The challenge of open-vocabulary recognition lies in the model has no clue of new categories it is applied to. Existing works have proposed different methods to embed category cues into the model e.g. through few-shot fine-tuning providing category names or textual descriptions to Vision-Language Mo…

2024

Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs

CVPR 2024poster

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities in various multi-modal tasks. Nevertheless their performance in fine-grained image understanding tasks is still limited. To address this issue this paper proposes a new framework to enhance the fine-grained image understand…

2024

Recognizing Ultra-High-Speed Moving Objects with Bio-Inspired Spike Camera

AAAI 2024technical

Bio-inspired spike camera mimics the sampling principle of primate fovea. It presents high temporal resolution and dynamic range, showing great promise in fast-moving object recognition. However, the physical limit of CMOS technology in spike cameras still hinders their capability of recognizing ult…

2024

SeACo-Paraformer: A Non-Autoregressive ASR System with Flexible and Effective Hotword Customization Ability

ICASSP 2024accepted

Hotword customization is one of the concerned issues remained in ASR field - it is of value to enable users of ASR systems to customize names of entities, persons and other phrases to obtain better experience. The past few years have seen effective modeling strategies for ASR contextualization devel…

Cited by 0SourceScholar
2024

SlideSpeech: A Large Scale Slide-Enriched Audio-Visual Corpus

ICASSP 2024accepted

Multi-Modal automatic speech recognition (ASR) techniques aim to leverage additional modalities to improve the performance of speech recognition systems. While existing approaches primarily focus on video or contextual information, the utilization of extra supplementary textual information has been…

Cited by 0SourceScholar
2024

emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

ACL 2024findings

We propose emotion2vec, a universal speech emotion representation model. emotion2vec is pre-trained on open-source unlabeled emotion data through self-supervised online distillation, combining utterance-level loss and frame-level loss during pre-training. emotion2vec outperforms state-of-the-art pre…

2023

Speech and Noise Dual-Stream Spectrogram Refine Network With Speech Distortion Loss For Robust Speech Recognition

ICASSP 2023accepted

In recent years, the joint training of speech enhancement front-end and automatic speech recognition (ASR) back-end has been widely used to improve the robustness of ASR systems. Traditional joint training methods only use enhanced speech as input for the backend. However, it is difficult for speech…

Cited by 0SourceScholar
2023

Unleashing the Full Potential of Product Quantization for Large-Scale Image Retrieval

NeurIPS 2023poster

Due to its promising performance, deep hashing has become a prevalent method for approximate nearest neighbors search (ANNs). However, most of current deep hashing methods are validated on relatively small-scale datasets, leaving potential threats when are applied to large-scale real-world scenarios…

2022

M2Met: The Icassp 2022 Multi-Channel Multi-Party Meeting Transcription Challenge

ICASSP 2022accepted

Recent development of speech signal processing, such as speech recognition, speaker diarization, etc., has inspired numerous applications of speech technologies. The meeting scenario is one of the most valuable and, at the same time, most challenging scenarios for the deployment of speech technologi…

Cited by 0SourceScholar
2022

MDERank: A Masked Document Embedding Rank Approach for Unsupervised Keyphrase Extraction

ACL 2022findings

Keyphrase extraction (KPE) automatically extracts phrases in a document that provide a concise summary of the core content, which benefits downstream information retrieval and NLP tasks. Previous state-of-the-art methods select candidate keyphrases based on the similarity between learned representat…

2022

Modeling The Detection Capability Of High-Speed Spiking Cameras

ICASSP 2022accepted

The novel working principle enables spiking cameras to capture high-speed moving objects. However, the applications of spiking cameras can be affected by many factors, such as brightness intensity, detectable distance, and the maximum speed of moving targets. Improper settings such as weak ambient b…

Cited by 0SourceScholar
2022

Prosospeech: Enhancing Prosody with Quantized Vector Pre-Training in Text-To-Speech

ICASSP 2022accepted

Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted pitch used in previous prosody modeling works have inevitable errors, which hurts the prosody modeling; 2) different attr…

Cited by 0SourceScholar
2022

Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis

EMNLP 2022main

Recently, hybrid systems of clustering and neural diarization models have been successfully applied in multi-party meeting analysis. However, current models always treat overlapped speaker diarization as a multi-label classification problem, where speaker dependency and overlaps are not well conside…

2022

Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge

ICASSP 2022accepted

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic spe…

Cited by 0SourceScholar
2021

Graph Consistency Based Mean-Teaching for Unsupervised Domain Adaptive Person Re-Identification

IJCAI 2021poster

Recent works show that mean-teaching is an effective framework for unsupervised domain adaptive person re-identification. However, existing methods perform contrastive learning on selected samples between teacher and student networks, which is sensitive to noises in pseudo labels and neglects the re…

2021

Robust Pose Estimation in Crowded Scenes with Direct Pose-Level Inference

NeurIPS 2021poster

Multi-person pose estimation in crowded scenes is challenging because overlapping and occlusions make it difficult to detect person bounding boxes and infer pose cues from individual keypoints. To address those issues, this paper proposes a direct pose-level inference strategy that is free of boundi…

2020

Joint Visual and Temporal Consistency for Unsupervised Domain Adaptive Person Re-Identification

ECCV 2020poster

Unsupervised domain adaptive person Re-IDentification (ReID) is challenging because of the large domain gap between source and target domains, as well as the lackage of labeled data on the target domain. This paper tackles this challenge through jointly enforcing visual and temporal consistency in t…

Cited by 194SourcePDFScholar
2019

Bi-Directional Cascade Network for Perceptual Edge Detection

CVPR 2019poster

Exploiting multi-scale representations is critical to improve edge detection for objects at different scales. To extract edges at dramatically different scales, we propose a Bi-Directional Cascade Network (BDCN) structure, where an individual layer is supervised by labeled edges at its specific scal…

Cited by 551PDFcodeScholar
2019

Global-Local Temporal Representations for Video Person Re-Identification

ICCV 2019poster

This paper proposes the Global-Local Temporal Representation (GLTR) to exploit the multi-scale temporal cues in video sequences for video person Re-Identification (ReID). GLTR is constructed by first modeling the short-term temporal cues among adjacent frames, then capturing the long-term relations…

Cited by 286PDFScholar
2019

Investigation of Modeling Units for Mandarin Speech Recognition Using Dfsmn-ctc-smbr

ICASSP 2019accepted

The choice of acoustic modeling units is critical to acoustic modeling in large vocabulary continuous speech recognition (LVCSR) tasks. The recent connectionist temporal classification (CTC) based acoustic models have more options for the choice of modeling units. In this work, we propose a DFSMN-CT…

Cited by 0SourceScholar
2019

Robust Audio-visual Speech Recognition Using Bimodal Dfsmn with Multi-condition Training and Dropout Regularization

ICASSP 2019accepted

Audio-visual speech recognition (AVSR) is thought to be one of the potential solutions for robust speech recognition, especially in noisy environments. Compared to audio only speech recognition, the major issues of AVSR include the lack of publicly available audio-visual corpora and the need of robu…

Cited by 0SourceScholar
2018

Deep Feed-Forward Sequential Memory Networks for Speech Synthesis

ICASSP 2018accepted

The Bidirectional LSTM (BLSTM) RNN based speech synthesis system is among the best parametric Text-to-Speech (TTS) systems in terms of the naturalness of generated speech, especially the naturalness in prosody. However, the model complexity and inference cost of BLSTM prevents its usage in many runt…

Cited by 0SourceScholar
2018

Person Transfer GAN to Bridge Domain Gap for Person Re-Identification

CVPR 2018poster

Although the performance of person Re-Identification (ReID) has been significantly boosted, many challenging issues in real scenarios have not been fully investigated, e.g., the complex scenes and lighting variations, viewpoint and pose changes, and the large number of identities in a camera network…

Cited by 2297SourcePDFScholar
2017

Pose-Driven Deep Convolutional Model for Person Re-Identification

ICCV 2017poster

Feature extraction and matching are two crucial components in person Re-Identification (ReID). The large pose deformations and the complex view variations exhibited by the captured person images significantly increase the difficulty of learning and matching of the features from person images. To ove…

Cited by 999PDFScholar
2015

Multi-Task Learning With Low Rank Attribute Embedding for Person Re-Identification

ICCV 2015poster

We propose a novel Multi-Task Learning with Low Rank Attribute Embedding (MTL-LORAE) framework for person re-identification. Re-identifications from multiple cameras are regarded as related tasks to exploit shared information to improve re-identification accuracy. Both low level features and semanti…

Cited by 204PDFScholar