← Search

Pu Wang

40 accepted papers

2026

CAST-LUT: Tokenizer-Guided HSV Look-Up Tables for Purple Flare Removal

AAAI 2026technical

Purple flare, a diffuse chromatic aberration artifact commonly found around highlight areas, severely degrades the tone transition and color of the image. Existing traditional methods are based on hand-crafted features, which lack flexibility and rely entirely on fixed priors, while the scarcity of

Cited by 0SourcePDFScholar
2026

DCA-LUT: Deep Chromatic Alignment with 5D LUT for Purple Fringing Removal

AAAI 2026technical

Purple fringing, a persistent artifact caused by Longitudinal Chromatic Aberration (LCA) in camera lenses, has long degraded the clarity and realism of digital imaging. Traditional solutions rely on complex and expensive apochromatic (APO) lens hardware and the extraction of handcrafted features, ig

Cited by 0SourcePDFScholar
2026

LiveGesture: Streamable Co-Speech Gesture Generation Model

CVPR 2026

We propose LiveGesture, the first fully streamable, speech-driven full-body gesture generation framework that operates with zero look-ahead and supports arbitrary sequence length. Unlike existing co-speech gesture methods--which are designed for offline generation and either treat body regions indep

Cited by 0SourceScholar
2026

MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos

CVPR 2026

Temporal Action Detection (TAD) in untrimmed videos poses significant challenges, particularly for Activities of Daily Living (ADL) requiring models to (1) process long-duration videos, (2) capture temporal variations in actions, and (3) simultaneously detect dense overlapping actions. Existing CNN

Cited by 0SourcecodeScholar
2026

SSVD-O: PARAMETER-EFFICIENT FINE-TUNING WITH STRUCTURED SVD FOR SPEECH RECOGNITION

ICASSP 2026oral

Parameter-efficient fine-tuning (PEFT) is a scalable approach for adapting large speech foundation models to new domains. While methods such as LoRA and its state-of-the-art variants reduce adaptation costs, they typically allocate parameters uniformly across model subspaces, which limits their effi…

Cited by 0SourcePDFScholar
2026

Walk Before You Dance: High-fidelity and Editable Dance Synthesis via Generative Masked Motion Prior

AAAI 2026technical

Recent advances in dance generation have enabled the automatic synthesis of 3D dance motions. However, existing methods still face significant challenges in simultaneously achieving high realism, precise dance-music synchronization, diverse motion expression, and physical plausibility. To address th

Cited by 0SourcePDFScholar
2025

GenHMR: Generative Human Mesh Recovery

AAAI 2025technical

Human mesh recovery (HMR) is crucial in many computer vision applications; from health to arts and entertainment. HMR from monocular images has predominantly been addressed by deterministic methods that output a single prediction for a given 2D image. However, HMR from a single image is an ill-posed…

Cited by 0SourcePDFScholar
2025

How to Auto-optimize Prompts for Domain Tasks? Adaptive Prompting and Reasoning through Evolutionary Domain Knowledge Adaptation

NeurIPS 2025poster

Designing optimal prompts and reasoning processes for large language models (LLMs) on domain-specific tasks is both necessary and challenging in real-world applications. Determining how to integrate domain knowledge, enhance reasoning efficiency, and even provide domain experts with refined knowledg…

Cited by 0SourceScholar
2025

Hybrid Learning-based Balance Function Assessment of Stroke Patients with a Single Ear-Worn IMU

IROS 2025

Rehabilitation robotics has attracted increasing attention due to its ability to provide continuous, precise, and adaptive treatment programs for stroke patients during their recovery. Accurately assessing lower-limb motor function is crucial in effectively implementing robot-assisted rehabilitation

Cited by 0SourceScholar
2025

LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living

CVPR 2025poster

Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant representation learning essential for Activities of Daily Living (ADL). This limitation s…

2025

MaskControl: Spatio-Temporal Control for Masked Motion Synthesis

ICCV 2025poster

Recent advances in motion diffusion models have enabled spatially controllable text-to-motion generation. However, these models struggle to achieve high-precision control while maintaining high-quality motion generation. To address these challenges, we propose MaskControl, the first approach to intr…

2025

MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the Wild

ICCV 2025poster

Reconstructing a 3D hand mesh from a single RGB image is challenging due to complex articulations, self-occlusions, and depth ambiguities. Traditional discriminative methods, which learn a deterministic mapping from a 2D image to a single 3D mesh, often struggle with the inherent ambiguities in 2D-t…

2025

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living

AAAI 2025technical

The introduction of vision-language models like CLIP has enabled the development of foundational video models capable of generalizing to unseen videos and human actions. However, these models are typically trained on web videos, which often fail to capture the challenges present in Activities of Dai…

2025

Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study

NeurIPS 2025spotlight

How to integrate and verify spatial intelligence in foundation models remains an open challenge. Current practice often proxies Visual-Spatial Intelligence (VSI) with purely textual prompts and VQA-style scoring, which obscures geometry, invites linguistic shortcuts, and weakens attribution to genui…

Cited by 0SourceScholar
2025

mmCooper: A Multi-agent Multi-stage Communication-efficient and Collaboration-robust Cooperative Perception Framework

ICCV 2025poster

Collaborative perception significantly enhances individual vehicle perception performance through the exchange of sensory information among agents. However, real-world deployment faces challenges due to bandwidth constraints and inevitable calibration errors during information exchange. To address t…

2024

BAMM: Bidirectional Autoregressive Motion Model

ECCV 2024poster

"Generating human motion from text has been dominated by denoising motion models either through diffusion or generative masking process. However, these models face great limitations in usability by requiring prior knowledge of the motion length. Conversely, autoregressive motion models address this…

2024

Object Trajectory Estimation with Multi-Band Wi-Fi Neural Dynamic Fusion

ICASSP 2024accepted

In contrast to existing multi-band Wi-Fi fusion in a frame-to-frame basis for simple classification, this paper considers asynchronous sequence-to-sequence fusion between sub-7GHz channel state information (CSI) and 60GHz beam SNR for more challenging downstream tasks such as continuous regression.…

Cited by 0SourceScholar
2024

Radar Perception with Scalable Connective Temporal Relations for Autonomous Driving

ICASSP 2024accepted

Due to the noise and low spatial resolution in automotive radar data, exploring temporal relations of learnable features over consecutive 2 radar frames has shown performance gain on downstream tasks (e.g., object detection and tracking) in our previous study [1]. In this paper, we further enhance r…

Cited by 0SourceScholar
2024

SIRA: Scalable Inter-frame Relation and Association for Radar Perception

CVPR 2024poster

Conventional radar feature extraction faces limitations due to low spatial resolution noise multipath reflection the presence of ghost targets and motion blur. Such limitations can be exacerbated by nonlinear object motion particularly from an ego-centric viewpoint. It becomes evident that to addres…

Cited by 3SourcePDFScholar
2024

WI-FI based Indoor Monitoring Enhanced by Multimodal Fusion

ICASSP 2024accepted

Indoor monitoring systems are in high demand to protect vulnerable people, especially when they are alone at home, in nursing homes, hospitals, etc. Although surveillance systems in public spaces use cameras and microphones to find incidents, indoor monitoring in personal spaces needs to protect pri…

Cited by 0SourceScholar
2023

Deep Proximal Gradient Method for Learned Convex Regularizers

ICASSP 2023accepted

We consider the problem of simultaneously learning a convex penalty function and its proximity operator for image reconstruction from incomplete measurements. Our goal is to apply Accelerated Proximal Gradient Method (APGM) using a learned proximity operator in place of the true proximity operator o…

Cited by 0SourceScholar
2023

Gaitmixer: Skeleton-Based Gait Representation Learning Via Wide-Spectrum Multi-Axial Mixer

ICASSP 2023accepted

Most existing gait recognition methods are appearance-based, which rely on the silhouettes extracted from the video data of human walking activities. The less-investigated skeleton-based gait recognition methods directly learn the gait dynamics from 2D/3D human skeleton sequences, which are theoreti…

Cited by 0SourceScholar
2023

Spatial-Domain Object Detection Under Mimo-Fmcw Automotive Radar Interference

ICASSP 2023accepted

This paper considers spatial-domain detector design for mutual interference mitigation among automotive MIMO-FMCW radars. This detector design is based on our previously derived interference signal model that fully accounts for the time-frequency incoherence and the slow-time code incoherence betwee…

Cited by 0SourceScholar
2023

mmWave Wi-Fi Trajectory Estimation with Continuous-Time Neural Dynamic Learning

ICASSP 2023accepted

We leverage standards-compliant beam training measurements from commercial-of-the-shelf (COTS) 802.11ad/ay devices for localization of a moving object. Two technical challenges need to be addressed: (1) the beam training measurements are intermittent due to beam scanning overhead control and content…

Cited by 0SourceScholar
2022

Adversarial Bi-Regressor Network for Domain Adaptive Regression

IJCAI 2022poster

Domain adaptation (DA) aims to transfer the knowledge of a well-labeled source domain to facilitate unlabeled target learning. When turning to specific tasks such as indoor (Wi-Fi) localization, it is essential to learn a cross-domain regressor to mitigate the domain shift. This paper proposes a nov…

Cited by 8SourcePDFScholar
2022

Incomplete Multi-View Domain Adaptation via Channel Enhancement and Knowledge Transfer

ECCV 2022poster

"Unsupervised domain adaptation (UDA) borrows well-labeled source knowledge to solve the specific task on unlabeled target domain with the assumption that both domains are from a single sensor, e.g., RGB or depth images. To boost model performance, multiple sensors are deployed on new-produced devic…

2022

Local Learning Matters: Rethinking Data Heterogeneity in Federated Learning

CVPR 2022oral

Federated learning (FL) is a promising strategy for performing privacy-preserving, distributed learning with a network of clients (i.e., edge devices). However, the data distribution among clients is often non-IID in nature, making efficient optimization difficult. To alleviate this issue, many FL a…

Cited by 217PDFcodeScholar
2021

A Consensus Equilibrium Solution For Deep Image Prior Powered By Red

ICASSP 2021accepted

Recent advances in solving imaging inverse problems have witnessed the combination of deep learning models with classical image models for better signal representation. One such approach, DeepRED, combines the deep image prior (DIP) with the regularization by denoising (RED) framework to boost the p…

Cited by 0SourceScholar
2021

Extended Object Tracking With Automotive Radar Using B-Spline Chained Ellipses Model

ICASSP 2021accepted

This paper introduces a B-spline chained ellipses model representation for extended object tracking (EOT) using high-resolution automotive radar measurements. With offline automotive radar training datasets, the proposed model parameters are learned using the expectation-maximization (EM) algorithm.…

Cited by 0SourceScholar
2020

Extended Object Tracking Using Hierarchical Truncation Measurement Model with Automotive Radar

ICASSP 2020accepted

Motivated by real-world automotive radar measurements that are distributed around object (e.g., vehicles) edges with a certain volume, a novel hierarchical truncated Gaussian measurement model is proposed to resemble the underlying spatial distribution of radar measurements. With the proposed measur…

Cited by 0SourceScholar
2020

Slow-Time MIMO-FMCW Automotive Radar Detection with Imperfect Waveform Separation

ICASSP 2020accepted

This paper considers object detection in the case of imperfect waveform separation, in the context of automotive radars with a slow-time MIMO-FMCW signaling scheme. We develop an explicit signal model that accounts for waveform separation residuals and propose a Kronecker subspace-based object detec…

Cited by 0SourceScholar
2019

Misspecified CRB on Parameter Estimation for a Coupled Mixture of Polynomial Phase and Sinusoidal FM Signals

ICASSP 2019accepted

This paper studies parameter estimation of a coupled mixture of polynomial phase signal (PPS) and sinusoidal frequency modulated (FM) signal, a newly introduced model motivated by industrial applications. Particularly, we analytically evaluate the estimation performance (or performance loss) via the…

Cited by 0SourceScholar
2018

Terahertz Imaging of Binary Reflectance with Variational Bayesian Inference

ICASSP 2018accepted

In this paper, we propose a Bayesian inference approach to extract the binary reflectance pattern of samples from compressed measurements in the terahertz (THz) frequency band. Compared with existing compressed THz imaging methods relying on the sparsity of the reflectance pattern, the proposed Baye…

Cited by 0SourceScholar
2016

Knowledge-aided hyperparameter-free Bayesian detection in stochastic homogeneous environments

ICASSP 2016accepted

This paper considers adaptive signal detection in stochastic homogeneous environments where the disturbance covariance matrix of both test and training signals, R, is assumed to be a random matrix with a priori knowledge of R. Unlike existing detectors assuming a known hyperparameter associated with…

Cited by 0SourceScholar