← Search

Peng Chang

9 accepted papers

2026

COGNOS: Universal Enhancement for Time Series Anomaly Detection via Constrained Gaussian-Noise Optimization and Smoothing

ICML 2026poster

Reconstruction-based methods are a dominant paradigm in time series anomaly detection (TSAD), however, their near-universal reliance on Mean Squared Error (MSE) loss results in statistically flawed reconstruction residuals. This fundamental weakness leads to noisy, unstable anomaly scores, hindering…

Cited by 0SourceScholar
2025

Co-Speech Gesture Video Generation with Implicit Motion-Audio Entanglement

CVPR 2025poster

Co-speech gestures are essential to non-verbal communication, enhancing both the naturalness and effectiveness of human interaction. Although recent methods have made progress in generating co-speech gesture videos, many rely on strong visual controls, such as pose images or TPS keypoint movements,…

2024

Co-speech Gesture Video Generation with 3D Human Meshes

ECCV 2024poster

"Co-speech gesture video generation is an enabling technique for many digital human applications. Substantial progress has been made in creating high-quality talking head videos. However, existing hand gesture video generation methods are primarily limited by the widely adopted 2D skeleton-based ges…

Cited by 1SourcePDFScholar
2024

Dialogue Cross-Enhanced Central Engagement Attention Model for Real-Time Engagement Estimation

IJCAI 2024poster

Real-time engagement estimation has been an important research topic in human-computer interaction in recent years. The emergence of the NOvice eXpert Interaction (NOXI) dataset, enriched with frame-wise engagement annotations, has catalyzed a surge in research efforts in this domain. Existing featu…

2022

Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis

EMNLP 2022finding

Multimodal speech emotion recognition (SER) and sentiment analysis (SA) are important techniques for human-computer interaction. Most existing multimodal approaches utilize either shallow cross-modal fusion of pretrained features, or deep cross-modal fusion with raw features. Recently, attempts have…

2022

Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment

ICASSP 2022accepted

Automatic pronunciation assessment is an important technology to help self-directed language learners. While pronunciation quality has multiple aspects including accuracy, fluency, completeness, and prosody, previous efforts typically only model one aspect (e.g., accuracy) at one granularity (e.g.,…

Cited by 0SourceScholar
2021

CASS-NAT: CTC Alignment-Based Single Step Non-Autoregressive Transformer for Speech Recognition

ICASSP 2021accepted

We propose a CTC alignment-based single step non-autoregressive transformer (CASS-NAT) for speech recognition. Specifically, the CTC alignment contains the information of (a) the number of tokens for decoder input, and (b) the time span of acoustics for each token. The information are used to extrac…

Cited by 0SourceScholar
2021

Leveraging Large-Scale Weakly Labeled Data for Semi-Supervised Mass Detection in Mammograms

CVPR 2021poster

Mammographic mass detection is an integral part of a computer-aided diagnosis system. Annotating a large number of mammograms at pixel-level in order to train a mass detection model in a fully supervised fashion is costly and time-consuming. This paper presents a novel self-training framework for se…

Cited by 14PDFScholar