← Search

Jian Xue

10 accepted papers

2026

AUHead: Realistic Emotional Talking Head Generation via Action Units Control

ICLR 2026poster

Realistic talking-head video generation is critical for virtual avatars, film production, and interactive systems. Current methods struggle with nuanced emotional expressions due to the lack of fine-grained emotion control. To address this issue, we introduce a novel two-stage method (AUHead) to dis…

Cited by 0SourcecodeScholar
2026

Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image Compression

CVPR 2026

Multi-view image compression (MIC) aims to achieve high compression efficiency by exploiting inter-image correlations, playing a crucial role in 3D applications. As a subfield of MIC, distributed multi-view image compression (DMIC) offers performance comparable to MIC while eliminating the need for

Cited by 0SourceScholar
2025

Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation

ACL 2025finding

Generative Error Correction (GEC) has emerged as a powerful post-processing method to boost the performance of Automatic Speech Recognition (ASR) systems. In this paper, we first show that GEC models struggle to generalize beyond the specific types of errors encountered during training, limiting the…

2025

Text-to-Any-Skeleton Motion Generation Without Retargeting

ICCV 2025poster

Recent advances in text-driven motion generation have shown notable advancements. However, these works are typically limited to standardized skeletons and rely on a cumbersome retargeting process to adapt to varying skeletal configurations of diverse characters. In this paper, we present OmniSkel, a…

Cited by 0SourcePDFScholar
2024

Diarist: Streaming Speech Translation with Speaker Diarization

ICASSP 2024accepted

End-to-end speech translation (ST) for conversation recordings involves several under-explored challenges such as speaker diarization (SD) without accurate word time stamps and handling of overlapping speech in a streaming fashion. In this work, we propose DiariST, the first streaming ST and SD solu…

Cited by 0SourceScholar
2024

Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation

ICASSP 2024accepted

The growing need for instant spoken language transcription and translation is driven by increased global communication and cross-lingual interactions. This has made offering translations in multiple languages essential for user applications. Traditional approaches to automatic speech recognition (AS…

Cited by 0SourceScholar
2023

Fast and Accurate Factorized Neural Transducer for Text Adaption of End-to-End Speech Recognition Models

ICASSP 2023accepted

Neural transducer is now the most popular end-to-end model for speech recognition, due to its naturally streaming ability. However, it is challenging to adapt it with text-only data. Factorized neural transducer (FNT) model was proposed to mitigate this problem. The improved adaptation ability of FN…

Cited by 0SourceScholar
2022

Shape-Adaptive Selection and Measurement for Oriented Object Detection

AAAI 2022technical

The development of detection methods for oriented object detection remains a challenging task. A considerable obstacle is the wide variation in the shape (e.g., aspect ratio) of objects. Sample selection in general object detection has been widely studied as it plays a crucial role in the performanc…

2019

A Markerless Body Motion Capture System for Character Animation Based on Multi-view Cameras

ICASSP 2019accepted

A novel application system is proposed in this paper to achieve the generation of 3D character animation driven by markerless human body motion capture. The whole pipeline of the system consists of four parts: capturing motion data by multiple cameras, detecting 2D human body joints and estimating 3…

Cited by 0SourceScholar
2015

Investigating online low-footprint speaker adaptation using generalized linear regression and click-through data

ICASSP 2015accepted

To develop speaker adaptation algorithms for deep neural network (DNN) that are suitable for large-scale online deployment, it is desirable that the adaptation model be represented in a compact form and learned in an unsupervised fashion. In this paper, we propose a novel low-footprint adaptation te…

Cited by 0SourceScholar