← Search

Wenming Yang

64 accepted papers

2026

Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy

ICLR 2026poster

Human image animation aims to generate a video from a static reference image, guided by pose information extracted from a driving video. Existing approaches often rely on pose estimators to extract intermediate representations, but such signals are prone to errors under occlusion or complex poses. B…

Cited by 0SourcecodeScholar
2026

CausalLens: Sensitivity-Guided Multi-Head Causal Intervention for Hallucination Mitigation in Large Vision-Language Models

CVPR 2026

Recent Large Vision-Language Models (LVLMs) have shown impressive capabilities in multimodal understanding and generation. Despite this progress, they remain prone to *hallucination*, where model outputs conflict with the visual input due to an over-reliance on textual priors. Existing inference-tim

Cited by 0SourceScholar
2026

CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation

CVPR 2026

Co-speech gesture generation has significantly advanced human-computer interaction, yet speaker movements remain constrained due to the omission of text-driven non-spontaneous gestures (e.g., bowing while talking). Existing methods face two key challenges: 1) the semantic prior gap due to the lack o

Cited by 0SourceScholar
2026

EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video

ICML 2026spotlight

Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware. Existing vision-based methods predominantly rely on planar surfaces or fingertip contacts, failing to generalize to complex 3…

Cited by 0SourcecodeScholar
2026

ICDiffAD: Implicit Conditioning Diffusion Model for Time Series Anomaly Detection

ICLR 2026poster

Time series anomaly detection (TSAD) faces critical challenges from intrinsic data noisiness and temporal heterogeneity, which undermine the reconstruction fidelity of prevailing generative approaches. While diffusion models offer theoretical advantages in capturing complex temporal dynamics, their…

Cited by 0SourceScholar
2026

OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text

ICLR 2026poster

Composed video retrieval presents a complex challenge: retrieving a target video based on a source video and a textual modification instruction. This task demands fine-grained reasoning over multimodal transformations. However, existing benchmarks predominantly focus on vision–text alignment, largel…

Cited by 0SourceScholar
2026

RAR: Reversing Visual Attention Re-Sinking for Unlocking Potential in Multimodal Large Language Models

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks, yet they frequently exhibit suboptimal output layers, where intermediate decoder layers outperform the final ones, signaling underutilized model capacity. In this work, we delve into the root causes a…

Cited by 0SourceScholar
2026

SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection

CVPR 2026

Existing monocular 3D detectors typically tame the pronounced nonlinear regression of 3D bounding box through decoupled prediction paradigm, which employs multiple branches to estimate geometric center, depth, dimensions, and rotation angle separately.Although this decoupling strategy simplifies the

Cited by 0SourceScholar
2026

TLDIFFGAN: A LATENT DIFFUSION-GAN FRAMEWORK WITH TEMPORAL INFORMATION FUSION FOR ANOMALOUS SOUND DETECTION

ICASSP 2026poster

Existing generative models for unsupervised anomalous sound detection are limited by their inability to fully capture the complex feature distribution of normal sounds, while the potential of powerful diffusion models in this domain remains largely unexplored. To address this challenge, we propose a…

Cited by 0SourcePDFScholar
2026

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

CVPR 2026

Despite remarkable advancements in multimodal large language models (MLLMs), their fine-grained visual understanding is constrained by a primary reliance on sparse textual supervision. Existing efforts to introduce visual supervision typically do so during post-training, when visual representations

Cited by 0SourceScholar
2026

VIVIDVOICE: A UNIFIED FRAMEWORK FOR SCENE-AWARE VISUALLY-DRIVEN SPEECH SYNTHESIS

ICASSP 2026poster

We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive auditory experiences that align with the real physical world. To tackle the two core challenges of data scarcity and modal…

Cited by 0SourcePDFScholar
2025

DM-Adapter: Domain-Aware Mixture-of-Adapters for Text-Based Person Retrieval

AAAI 2025technical

Text-based person retrieval (TPR) has gained significant attention as a fine-grained and challenging task that closely aligns with practical applications. Tailoring CLIP to person domain is now a emerging research topic due to the abundant knowledge of vision-language pretraining, but challenges sti…

2025

Decoupling Appearance Variations with 3D Consistent Features in Gaussian Splatting

AAAI 2025technical

Gaussian Splatting has emerged as a prominent 3D representation in novel view synthesis, but it still suffers from appearance variations, which are caused by various factors, such as modern camera ISPs, different time of day, weather conditions, and local light changes. These variations can lead to…

Cited by 2SourcePDFScholar
2025

FEG-VON: Frontier Embedding Graph for Efficient Visual Object Navigation

IROS 2025

Visual object navigation, requiring agents to locate target objects in novel environments through egocentric visual observation, remains a critical challenge in Embodied AI. We propose FEG-VON, a training-free framework that constructs and maintains a Frontier Embedding Graph for efficient Visual Ob

Cited by 0SourceScholar
2025

GRADEO: Towards Human-Like Evaluation for Text-to-Video Generation via Multi-Step Reasoning

ICML 2025poster

Recent great advances in video generation models have demonstrated their potential to produce high-quality videos, bringing challenges to effective evaluation. Unlike human evaluation, existing automated evaluation metrics lack high-level semantic understanding and reasoning capabilities for video,…

Cited by 0SourcePDFScholar
2025

GaussianSR: High Fidelity 2D Gaussian Splatting for Arbitrary-Scale Image Super-Resolution

AAAI 2025technical

Implicit neural representations (INRs) have revolutionized arbitrary-scale super-resolution (ASSR) by modeling images as continuous functions. Most existing INR-based ASSR networks first extract features from the given low-resolution image using an encoder, and then render the super-resolved result…

Cited by 3SourcePDFScholar
2025

MonoDGP: Monocular 3D Object Detection with Decoupled-Query and Geometry-Error Priors

CVPR 2025poster

Perspective projection has been extensively utilized in monocular 3D object detection methods. It introduces geometric priors from 2D bounding boxes and 3D object dimensions to reduce the uncertainty of depth estimation. However, due to errors originating from the object's visual surface, the boundi…

2025

Pose Magic: Efficient and Temporally Consistent Human Pose Estimation with a Hybrid Mamba-GCN Network

AAAI 2025technical

Current state-of-the-art (SOTA) methods in 3D Human Pose Estimation (HPE) are primarily based on Transformers. However, existing Transformer-based 3D HPE backbones often encounter a trade-off between accuracy and computational efficiency. To resolve the above dilemma, in this work, we leverage recen…

Cited by 3SourcePDFScholar
2025

Region-Centric 6-Dof Grasp Detection: A Data-Efficient Solution for Cluttered Scenes

IROS 2025

Robotic grasping, serving as the cornerstone of robot manipulation, is fundamental for embodied intelligence. Manipulation in challenging scenarios demands grasp detection algorithms with higher efficiency and generalizability. However, for general 6-Dof grasp detection, most data-driven methods dir

Cited by 0SourceScholar
2025

SAP-SLAM: Semantic-Assisted Perception SLAM with 3D Gaussian Splatting

ICRA 2025

The integration of 3D Gaussians has introduced a novel scene representation in Simultaneous Localization and Mapping (SLAM), characterized by explicit representation and differentiable rendering capabilities that enhance scene reconstruction and understanding. However, most current SLAM systems only

Cited by 0SourceScholar
2025

SOVGaussian: Sparse-View 3D Gaussian Splatting for Open-Vocabulary Scene Understanding

AAAI 2025technical

Modeling 3D open-vocabulary language fields is challenging yet highly anticipated. Despite great progress, existing approaches heavily rely on a large number of training views to construct language-embedded 3D scenes, which is unfortunately impractical in real-world scenarios. This paper introduces…

2024

Bilateral Event Mining and Complementary for Event Stream Super-Resolution

CVPR 2024poster

Event Stream Super-Resolution (ESR) aims to address the challenge of insufficient spatial resolution in event streams which holds great significance for the application of event cameras in complex scenarios. Previous works for ESR often process positive and negative events in a mixed paradigm. This…

2024

Binding-Adaptive Diffusion Models for Structure-Based Drug Design

AAAI 2024technical

Structure-based drug design (SBDD) aims to generate 3D ligand molecules that bind to specific protein targets. Existing 3D deep generative models including diffusion models have shown great promise for SBDD. However, it is complex to capture the essential protein-ligand interactions exactly in 3D sp…

2024

Clip-Based Synergistic Knowledge Transfer for text-based Person Retrieval

ICASSP 2024accepted

Text-based Person Retrieval (TPR) aims to retrieve the target person images given a textual query. The primary challenge lies in bridging the substantial gap between vision and language modalities, especially when dealing with limited large-scale datasets. In this paper, we introduce a CLIP-based Sy…

Cited by 0SourceScholar
2024

DEGAN: Discrimination Enhanced GAN for Perceptual-Oriented Super-Resolution

ICASSP 2024accepted

Recent years, generative adversarial networks (GANs) have gained significant prominence in single image super-resolution (SISR) tasks. This can mainly be attributed to their exceptional ability to generate intricate details. However, the instability and lack of realism in the details generated by GA…

Cited by 0SourceScholar
2024

Diffusion-Based Pose Refinement and Multi-Hypothesis Generation for 3D Human Pose Estimation

ICASSP 2024accepted

Previous probabilistic models for 3D Human Pose Estimation (3DHPE) aimed to enhance pose accuracy by generating multiple hypotheses. However, most of the hypotheses generated deviate substantially from the true pose. Compared to deterministic models, the excessive uncertainty in probabilistic models…

Cited by 0SourceScholar
2024

Improving Diffusion-Based Image Restoration with Error Contraction and Error Correction

AAAI 2024technical

Generative diffusion prior captured from the off-the-shelf denoising diffusion generative model has recently attained significant interest. However, several attempts have been made to adopt diffusion models to noisy inverse problems either fail to achieve satisfactory results or require a few thousa…

Cited by 4SourcePDFScholar
2024

Interaction-based Retrieval-augmented Diffusion Models for Protein-specific 3D Molecule Generation

ICML 2024poster

Generating ligand molecules that bind to specific protein targets via generative models holds substantial promise for advancing structure-based drug design. Existing methods generate molecules from scratch without reference or template ligands, which poses challenges in model optimization and may yi…

2024

LLM-Empowered State Representation for Reinforcement Learning

ICML 2024poster

Conventional state representations in reinforcement learning often omit critical task-related details, presenting a significant challenge for value networks in establishing accurate mappings from states to task rewards. Traditional methods typically depend on extensive sample learning to enrich stat…

2024

NLSIT: A Non-Local Stereo Interaction Transformer for Stereo Image Super-Resolution

ICASSP 2024accepted

In recent years, although Transformer has been introduced into stereo image super-resolution and accomplished great advances, the long-range complementary information in stereo images hasn’t been fully utilized. In view of beneficial non-local prior knowledge in both intra-view and cross-view, we pr…

Cited by 0SourceScholar
2024

Protein-Ligand Interaction Prior for Binding-aware 3D Molecule Diffusion Models

ICLR 2024poster

Generating 3D ligand molecules that bind to specific protein targets via diffusion models has shown great promise for structure-based drug design. The key idea is to disrupt molecules into noise through a fixed forward process and learn its reverse process to generate molecules from noise in a denoi…

2024

RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation

CVPR 2024poster

Real-time multi-person pose estimation presents significant challenges in balancing speed and precision. While two-stage top-down methods slow down as the number of people in the image increases existing one-stage methods often fail to simultaneously deliver high accuracy and real-time performance.…

2024

Residual Dense Swin Transformer for Continuous Depth-Independent Ultrasound Imaging

ICASSP 2024accepted

Ultrasound imaging is crucial for evaluating organ morphology and function, yet depth adjustment can degrade image quality and field-of-view, presenting a depth-dependent dilemma. Traditional interpolation-based zoom-in techniques often sacrifice detail and introduce artifacts. Motivated by the pote…

Cited by 0SourceScholar
2024

VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction

CVPR 2024poster

Existing NeRF-based methods for large scene reconstruction often have limitations in visual quality and rendering speed. While the recent 3D Gaussian Splatting works well on small-scale and object-centric scenes scaling it up to large scenes poses challenges due to limited video memory long optimiza…

Cited by 116SourcePDFScholar
2023

Basic Binary Convolution Unit for Binarized Image Restoration Network

ICLR 2023poster

Lighter and faster image restoration (IR) models are crucial for the deployment on resource-limited devices. Binary neural network (BNN), one of the most promising model compression methods, can dramatically reduce the computations and parameters of full-precision convolutional neural networks (CNN)…

2023

Crafting Training Degradation Distribution for the Accuracy-Generalization Trade-off in Real-World Super-Resolution

ICML 2023poster

Super-resolution (SR) techniques designed for real-world applications commonly encounter two primary challenges: generalization performance and restoration accuracy. We demonstrate that when methods are trained using complex, large-range degradations to enhance generalization, a decline in accuracy…

Cited by 24SourcePDFScholar
2023

DiffIR: Efficient Diffusion Model for Image Restoration

ICCV 2023poster

Diffusion model (DM) has achieved SOTA performance by modeling the image synthesis process into a sequential application of a denoising network. However, different from image synthesis generating each pixel from scratch, most pixels of image restoration (IR) are given. Thus, for IR, traditional DMs…

Cited by 291PDFcodeScholar
2023

Efficient Heatmap-Guided 6-Dof Grasp Detection in Cluttered Scenes

RA-L 2023

Fast and robust object grasping in clutter is a crucial component of robotics. Most current works resort to the whole observed point cloud for 6-Dof grasp generation, ignoring the guidance information excavated from global semantics, thus limiting high-quality grasp generation and real-time performa

Cited by 55SourcecodeScholar
2023

Flowpose: Conditional Normalizing Flows for 3D Human Pose and Shape Estimation from Monocular Videos

ICASSP 2023accepted

Human motion modeling is essential for video-based 3D human pose and shape estimation. Most existing methods model human motion by learning a deterministic mapping from the input videos to the human body parameters, while the uncertainties such as occlusions and depth ambiguities are ignored. To add…

Cited by 0SourceScholar
2023

Knowledge Distillation based Degradation Estimation for Blind Super-Resolution

ICLR 2023poster

Blind image super-resolution (Blind-SR) aims to recover a high-resolution (HR) image from its corresponding low-resolution (LR) input image with unknown degradations. Most of the existing works design an explicit degradation estimator for each degradation to guide SR. However, it is infeasible to pr…

2023

Local and Global Logit Adjustments for Long-Tailed Learning

ICCV 2023poster

Multi-expert ensemble models for long-tailed learning typically either learn diverse generalists from the whole dataset or aggregate specialists on different subsets. However, the former is insufficient for tail classes due to the high imbalance factor of the entire dataset, while the latter may bri…

Cited by 25PDFScholar
2023

Retiformer: Retinex-Based Enhancement In Transformer For Low-Light Image

ICASSP 2023accepted

Transformer-based methods have shown impressive potential in many low-level vision tasks but are rarely used for low-light image enhancement (LLIE). Direct use of Transformer in LLIE will bring unnatural visual effects. This phenomenon encourages us to attempt to learn from the theory of Retinex. Af…

Cited by 0SourceScholar
2023

Robust Content-Variant Reference Image Quality Assessment Via Similar Patch Matching

ICASSP 2023accepted

Although image quality assessment (IQA) methods have achieved remarkable success in the past decades, full-reference IQA is limited to reference images, while no-reference IQA has relatively poor performance. To boost the performance of IQA models in the no-reference scenario, a new class of IQA met…

Cited by 0SourceScholar
2023

Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short Video

ICCV 2023poster

Synthesizing realistic videos according to a given speech is still an open challenge. Previous works have been plagued by issues such as inaccurate lip shape generation and poor image quality. The key reason is that only motions and appearances on limited facial areas (e.g., lip area) are mainly dri…

Cited by 17PDFcodeScholar
2023

Structured Sparsity Learning for Efficient Video Super-Resolution

CVPR 2023poster

The high computational costs of video super-resolution (VSR) models hinder their deployment on resource-limited devices, e.g., smartphones and drones. Existing VSR models contain considerable redundant filters, which drag down the inference efficiency. To prune these unimportant filters, we develop…

2022

Coarse-to-Fine Embedded PatchMatch and Multi-Scale Dynamic Aggregation for Reference-Based Super-resolution

AAAI 2022technical

Reference-based super-resolution (RefSR) has made significant progress in producing realistic textures using an external reference (Ref) image. However, existing RefSR methods obtain high-quality correspondence matchings consuming quadratic computation resources with respect to the input size, limit…

2022

Efficient Non-local Contrastive Attention for Image Super-resolution

AAAI 2022technical

Non-Local Attention (NLA) brings significant improvement for Single Image Super-Resolution (SISR) by leveraging intrinsic feature correlation in natural images. However, NLA gives noisy information large weights and consumes quadratic computation resources with respect to the input size, limiting it…

2022

Pose-Invariant Face Recognition via Adaptive Angular Distillation

AAAI 2022technical

Pose-invariant face recognition is a practically useful but challenging task. This paper introduces a novel method to learn pose-invariant feature representation without normalizing profile faces to frontal ones or learning disentangled features. We first design a novel strategy to learn pose-invari…

Cited by 3SourcePDFScholar
2022

SCS-Co: Self-Consistent Style Contrastive Learning for Image Harmonization

CVPR 2022poster

Image harmonization aims to achieve visual consistency in composite images by adapting a foreground to make it compatible with a background. However, existing methods always only use the real image as the positive sample to guide the training, and at most introduce the corresponding composite image…

Cited by 53PDFcodeScholar
2022

Super-Resolution by Predicting Offsets: An Ultra-Efficient Super-Resolution Network for Rasterized Images

ECCV 2022poster

"Rendering high-resolution (HR) graphics brings substantial computational costs. Efficient graphics super-resolution (SR) methods may achieve HR rendering with small computing resources and have attracted extensive research interests in industry and research communities. We present a new method for…

Cited by 7SourcePDFScholar
2021

Amortized Bayesian Prototype Meta-learning: A New Probabilistic Meta-learning Approach to Few-shot Image Classification

AISTATS 2021poster

Probabilistic meta-learning methods recently have achieved impressive success in few-shot image classification. However, they introduce a huge number of random variables for neural network weights and thus severe computational and inferential challenges. In this paper, we propose a novel probabilist…

Cited by 26SourcePDFScholar
2021

Group Fisher Pruning for Practical Network Compression

ICML 2021spotlight

Network compression has been widely studied since it is able to reduce the memory and computation cost during inference. However, previous methods seldom deal with complicated structures like residual connections, group/depth-wise convolution and feature pyramid network, where channels of multiple l…

2021

Towards Impartial Multi-task Learning

ICLR 2021poster

Multi-task learning (MTL) has been widely used in representation learning. However, naively training all tasks simultaneously may lead to the partial training issue, where specific tasks are trained more adequately than others. In this paper, we propose to learn multiple tasks impartially. Specifica…

Cited by 198SourcePDFScholar
2018

Full-Reference Quality Assessment of Contrast Changed Images Based on Local Linear Model

ICASSP 2018accepted

This paper presents a new full-reference method to assess the quality of contrast changed images. In this method, we employ a linear model to describe the relationship between local patches of reference images and contrast changed images. With parameters of this model, three quality measures conside…

Cited by 0SourceScholar
2016

An efficient anomaly detection approach in surveillance video based on oriented GMM

ICASSP 2016accepted

The detection and localization of abnormal activities are considered in this work. An efficient approach called oriented G-MM(OGMM) is proposed. The approach uses optical flow as low-level feature and quantizes the orientation of optical flow into 8 sections. In training stage, the approach will lea…

Cited by 0SourceScholar
2016

Saliency detection based on integration of central bias, reweighting and multi-scale for superpixels

ICASSP 2016accepted

Saliency detection has been a significant problem in computer vision and helpful to object detection. In this paper, we propose a new computational saliency detection model under the Bayesian framework. First, central bias and the reweighting of the salient regions in the convex hull are applied to…

Cited by 0SourceScholar