← Search

Kaihao Zhang

36 accepted papers

2026

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling

ICML 2026poster

Preference optimization for diffusion models relies on reward functions that are both discriminative and computationally efficient. Vision-Language Models (VLMs) have emerged as powerful reward providers. However, their computation and memory cost can be substantial, and optimizing a latent diffusio…

Cited by 0SourceScholar
2026

InclusiveVidPose: Bridging the Pose Estimation Gap for Individuals with Limb Deficiencies in Video-Based Motion

ICLR 2026poster

Approximately 445.2 million individuals worldwide are living with traumatic amputations, and an estimated 31.64 million children aged 0–14 have congenital limb differences, yet they remain largely underrepresented in human pose estimation (HPE) research. Accurate HPE could significantly benefit this…

Cited by 0SourcecodeScholar
2026

MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding

CVPR 2026

Understanding high-resolution (HR) images remains a critical challenge for multimodal large language models (MLLMs). Recent approaches leverage vision-based retrieval-augmented generation (RAG) to retrieve query-relevant crops from HR images, improving understanding capacity of MLLMs. However, this

Cited by 0SourcecodeScholar
2026

Position: Human-Centric Vision Requires Topological Generalization Beyond Fixed Skeletal Topologies

ICML 2026poster

In this position paper, we argue that human-centric vision requires skeletal-topology generalization beyond fixed skeletons. Mainstream pose and body pipelines enforce a fixed skeleton graph with an indexed joint list and fixed adjacency, so the fixed joint inventory does not cover structural absenc…

Cited by 0SourceScholar
2026

ResiHMR: Residual-Limb Aware Single-Image 3D Human Mesh Recovery for Individuals with Limb Loss

CVPR 2026

Single-image human mesh recovery provides a compact 3D, person-centric representation that supports analysis, animation, AR and VR, rehabilitation, and human-computer interaction. However, prevailing systems impose an intact-limb prior and degrade on people with limb loss, because fixed-topology mod

Cited by 0SourceScholar
2025

Cross-View Isolated Sign Language Recognition via View Synthesis and Feature Disentanglement

ICCV 2025poster

Cross-view isolated sign language recognition (CV-ISLR) addresses the challenge of identifying isolated signs from viewpoints unseen during training, a problem aggravated by the scarcity of multi-view data in existing benchmarks. To bridge this gap, we introduce a novel two-stage framework comprisin…

Cited by 0SourcePDFScholar
2025

Gradient-Based Adversarial Attacks on Deep LiDAR Odometry

ICRA 2025

Adversarial attacks have been recently investigated in LiDAR perception problems for autonomous driving, where a small perturbation of source inputs can result in incorrect predictions. However, most previous studies focus on attacks on single-frame perception modules, lacking explorations of attack

Cited by 2SourceScholar
2025

LDPose: Towards Inclusive Human Pose Estimation for Limb-Deficient Individuals in the Wild

ICCV 2025poster

Human pose estimation aims to predict the location of body keypoints and enable various practical applications. However, existing research focuses solely on individuals with full physical bodies and overlooks those with limb deficiencies. As a result, current pose estimation methods cannot be genera…

Cited by 0SourcePDFScholar
2025

LLFA: Fusing Global Illumination and Local Priors for Low-Light Face Image Enhancement with Adaptor

ICASSP 2025accepted

Low-light image enhancement problem has been widely studied. However, most existing methods do not perform well on low-light face images due to no specific facial characteristic considerations. We first create large-scale low-light face datasets with synthesized and real-world images to address the…

Cited by 0SourceScholar
2025

MOERL: When Mixture-of-Experts Meet Reinforcement Learning for Adverse Weather Image Restoration

ICCV 2025poster

Adverse weather conditions, such as rain, snow, and haze, introduce complex degradations that present substantial challenges for effective image restoration. Existing all-in-one models often rely on fixed network structures, limiting their ability to adapt to the varying characteristics of different…

Cited by 0SourcePDFScholar
2025

MaterialMVP: Illumination-Invariant Material Generation via Multi-view PBR Diffusion

ICCV 2025poster

Physically-based rendering (PBR) has become a cornerstone in modern computer graphics, enabling realistic material representation and lighting interactions in 3D scenes. In this paper, we present MaterialMVP, a novel end-to-end model for generating PBR textures from 3D meshes and image prompts, addr…

Cited by 0SourcePDFScholar
2025

PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models

ACL 2025long

Large Language Models (LLMs) suffer severe performance degradation when facing extremely low-bit (sub 2-bit) quantization. Several existing sub 2-bit post-training quantization (PTQ) methods utilize a mix-precision scheme by leveraging an unstructured fine-grained mask to explicitly distinguish sali…

2025

Segmentation-Guided Sparse Transformer for Under-Display Camera Image Restoration

ICASSP 2025accepted

Under-display Camera is an emerging technology for full-screen display with a camera under the display. However, the current implementation of UDC causes serious image degradation. Incident light required for camera imaging undergoes attenuation and diffraction when passing through the display. Curr…

Cited by 0SourceScholar
2025

Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling

ICLR 2025poster

Controllable character image animation has a wide range of applications. Although existing studies have consistently improved performance, challenges persist in the field of character image animation, particularly concerning stability in complex backgrounds and tasks involving multiple characters. T…

Cited by 0SourcePDFScholar
2024

OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models

ECCV 2024poster

"Personalization is an important topic in text-to-image generation, especially the challenging multi-concept personalization. Current multi-concept methods are struggling with identity preservation, occlusion, and the harmony between foreground and background. In this work, we propose OMG, an occlus…

2024

View From Above: Orthogonal-View aware Cross-view Localization

CVPR 2024poster

This paper presents a novel aerial-to-ground feature aggregation strategy tailored for the task of cross-view image-based geo-localization. Conventional vision-based methods heavily rely on matching ground-view image features with a pre-recorded image database often through establishing planar homog…

Cited by 5SourcePDFScholar
2023

F&F Attack: Adversarial Attack against Multiple Object Trackers by Inducing False Negatives and False Positives

ICCV 2023poster

Multi-object tracking (MOT) aims to build moving trajectories for number-agnostic objects. Modern multi-object trackers commonly follow the tracking-by-detection strategy. Therefore, fooling detectors can be an effective solution but it usually requires attacks in multiple successive frames, resulti…

Cited by 10PDFScholar
2023

Homography Guided Temporal Fusion for Road Line and Marking Segmentation

ICCV 2023poster

Reliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occluded in the presence of moving vehicles, shadow, and glare and (2) highly structured with low intra-class shape variance a…

Cited by 5PDFcodeScholar
2023

InterTracker: Discovering and Tracking General Objects Interacting with Hands in the Wild

IROS 2023poster

Understanding human interaction with objects is an important research topic for embodied Artificial Intelligence and identifying the objects that humans are interacting with is a primary problem for interaction understanding. Existing methods rely on frame-based detectors to locate interacting objec…

Cited by 1SourceScholar
2023

MB-TaylorFormer: Multi-Branch Efficient Transformer Expanded by Taylor Formula for Image Dehazing

ICCV 2023poster

In recent years, Transformer networks are beginning to replace pure convolutional neural networks (CNNs) in the field of computer vision due to their global receptive field and adaptability to input. However, the quadratic computational complexity of softmax-attention limits the wide application in…

Cited by 124PDFcodeScholar
2023

Model Calibration in Dense Classification with Adaptive Label Perturbation

ICCV 2023poster

For safety-related applications, it is crucial to produce trustworthy deep neural networks whose prediction is associated with confidence that can represent the likelihood of correctness for subsequent decision-making. Existing dense binary classification models are prone to being over-confident. To…

Cited by 4PDFcodeScholar
2023

Punctuation-level Attack: Single-shot and Single Punctuation Can Fool Text Models

NeurIPS 2023poster

The adversarial attacks have attracted increasing attention in various fields including natural language processing. The current textual attacking models primarily focus on fooling models by adding character-/word-/sentence-level perturbations, ignoring their influence on human perception. In this p…

Cited by 3SourcePDFScholar
2023

Robust Single Image Reflection Removal Against Adversarial Attacks

CVPR 2023poster

This paper addresses the problem of robust deep single-image reflection removal (SIRR) against adversarial attacks. Current deep learning based SIRR methods have shown significant performance degradation due to unnoticeable distortions and perturbations on input images. For a comprehensive robustnes…

2023

Ultra-High-Definition Low-Light Image Enhancement: A Benchmark and Transformer-Based Method

AAAI 2023technical

As the quality of optical sensors improves, there is a need for processing large-scale images. In particular, the ability of devices to capture ultra-high definition (UHD) images and video places new demands on the image processing pipeline. In this paper, we consider the task of low-light image enh…

2021

ARVo: Learning All-Range Volumetric Correspondence for Video Deblurring

CVPR 2021poster

Video deblurring models exploit consecutive frames to remove blurs from camera shakes and object motions. In order to utilize neighboring sharp patches, typical methods rely mainly on homography or optical flows to spatially align neighboring blurry frames. However, such explicit approaches are less…

Cited by 83PDFScholar
2021

Benchmarking Ultra-High-Definition Image Super-Resolution

ICCV 2021poster

Increasingly, modern mobile devices allow capturing images at Ultra-High-Definition (UHD) resolution, which includes 4K and 8K images. However, current single image super-resolution (SISR) methods focus on super-resolving images to ones with resolution up to high definition (HD) and ignore higher-re…

Cited by 38PDFScholar
2021

Deep Two-View Structure-From-Motion Revisited

CVPR 2021poster

Two-view structure-from-motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM. Existing deep learning-based approaches formulate the problem in ways that are fundamentally ill-posed, relying on training data to overcome the inherent difficulties. In contrast, we propose a return to th…

Cited by 63PDFcodeScholar
2021

Pyramid Architecture Search for Real-Time Image Deblurring

ICCV 2021poster

Multi-scale and multi-patch deep models have been shown effective in removing blurs of dynamic scenes. However, these methods still have one major obstacle: manually designing a lightweight and high-efficiency network is challenging and time-consuming. To tackle this problem, we propose a novel debl…

Cited by 48PDFScholar
2020

Beyond Monocular Deraining: Stereo Image Deraining via Semantic Understanding

ECCV 2020poster

Rain is a common natural phenomenon. Taking images in the rain however often results in degraded quality of images, thus compromises the performance of many computer vision systems. Most existing de-rain algorithms use only one single input image and aim to recover a clean image. Few work has exploi…

Cited by 54SourcePDFScholar
2020

Displacement-Invariant Matching Cost Learning for Accurate Optical Flow Estimation

NeurIPS 2020poster

Learning matching costs has been shown to be critical to the success of the state-of-the-art deep stereo matching methods, in which 3D convolutions are applied on a 4D feature volume to learn a 3D cost volume. However, this mechanism has never been employed for the optical flow task. This is mainly…

2020

Human Parsing Based Texture Transfer from Single Image to 3D Human via Cross-View Consistency

NeurIPS 2020poster

This paper proposes a human parsing based texture transfer model via cross-view consistency learning to generate the texture of 3D human body from a single image. We use the semantic parsing of human body as input for providing both the shape and pose information to reduce the appearance variation…

2020

Single Image Super-Resolution via a Holistic Attention Network

ECCV 2020poster

Informative features play a crucial role in the single image super-resolution task. Channel attention has been demonstrated to be effective for preserving information-rich features in each layer. However, channel attention treats each convolution layer as a separate process, which is kind of missing…

Cited by 878SourcePDFScholar
2020

TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language Translation

NeurIPS 2020poster

Sign language translation (SLT) aims to interpret sign video sequences into text-based natural language sentences. Sign videos consist of continuous sequences of sign gestures with no clear boundaries in between. Existing SLT models usually represent sign visual features in a frame-wise manner so as…

2020

Unsupervised Domain Adaptation with Noise Resistible Mutual-Training for Person Re-identification

ECCV 2020poster

Unsupervised domain adaptation (UDA) in the task of person re-identification (re-ID) is highly challenging due to large domain divergence and no class overlap between domains. Pseudo-label based self-training is one of the representative techniques to address UDA. However, label noise caused by unsu…

Cited by 227SourcePDFScholar