← Search

Shanshe Wang

9 accepted papers

2023

Diffusion-Based 3D Human Pose Estimation with Multi-Hypothesis Aggregation

ICCV 2023poster

In this paper, a novel Diffusion-based 3D Pose estimation (D3DP) method with Joint-wise reProjection-based Multi-hypothesis Aggregation (JPMA) is proposed for probabilistic 3D human pose estimation. On the one hand, D3DP generates multiple possible 3D pose hypotheses for a single 2D observation. It…

Cited by 125PDFcodeScholar
2022

AIMNet: Adaptive Image-Tag Merging Network For Automatic Medical Report Generation

ICASSP 2022accepted

In recent years, medical report generation has received increasing research interest with the goal of automatically generating long and coherent descriptive paragraphs that can de-scribe in detail the observations of normal and abnormal regions in the input medical images. Unlike general image capti…

Cited by 0SourceScholar
2022

P-STMO: Pre-trained Spatial Temporal Many-to-One Model for 3D Human Pose Estimation

ECCV 2022poster

"This paper introduces a novel Pre-trained Spatial Temporal Many-to-One (P-STMO) model for 2D-to-3D human pose estimation task. To reduce the difficulty of capturing spatial and temporal information, we divide this task into two stages: pre-training (Stage I) and fine-tuning (Stage II). In Stage I,…

2022

STRPM: A Spatiotemporal Residual Predictive Model for High-Resolution Video Prediction

CVPR 2022poster

Although many video prediction methods have obtained good performance in low-resolution (64 128) videos, predictive models for high-resolution (512 4K) videos have not been fully explored yet, which are more meaningful due to the increasing demand for high-quality videos. Compared with low-resolutio…

Cited by 68PDFScholar
2021

Evolutionary Quantization of Neural Networks with Mixed-Precision

ICASSP 2021accepted

Quantization is an effective way for reducing the memory and computation costs of deep neural networks. Most of existing methods exploit the fixed-precision quantization approach, e.g., weights and activations (i.e., output features) are represented as 8-bit values. Although mixed-precision quantiza…

Cited by 0SourceScholar
2021

MAU: A Motion-Aware Unit for Video Prediction and Beyond

NeurIPS 2021poster

Accurately predicting inter-frame motion information plays a key role in video prediction tasks. In this paper, we propose a Motion-Aware Unit (MAU) to capture reliable inter-frame motion information by broadening the temporal receptive field of the predictive units. The MAU consists of two modules,…

2021

Teacher-Student Learning With Multi-Granularity Constraint Towards Compact Facial Feature Representation

ICASSP 2021accepted

In this paper, we propose a novel end-to-end feature compression scheme by leveraging the representation and learning capability of deep neural networks, towards intelligent front-end equipped analysis with promising accuracy and efficiency. In particular, the extracted features are compactly coded…

Cited by 0SourceScholar
2018

Cluster-Based Point Cloud Coding with Normal Weighted Graph Fourier Transform

ICASSP 2018accepted

Point cloud has attracted more and more attention in 3D object representation, especially in free-view rendering. However, it is challenging to efficiently deploy the point cloud due to its huge data amount with multiple attributes including coordinates, normal and color. In order to represent point…

Cited by 0SourceScholar