← Search

Haoqiang Fan

37 accepted papers

2026

Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation

CVPR 2026

Bimanual manipulation requires policies that can reason about 3D geometry, anticipate how it evolves under action, and generate smooth, coordinated motions. However, existing methods typically rely on 2D features with limited spatial awareness, or require explicit point clouds that are difficult to

Cited by 0SourcecodeScholar
2026

BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation

ICRA 2026poster

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT ) tend to treat multi-view features equally and directly concatenate them for policy learning. How ever, it will introduce redundant visual information and bring hig…

2026

Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot Manipulation

CVPR 2026

While existing equivariant methods enhance data efficiency, they suffer from high computational intensity, reliance on single-modality inputs, and instability when combined with fast-sampling methods. In this work, we propose E3Flow, a novel framework that addresses the critical limitations of equiv

Cited by 0SourcecodeScholar
2026

HeRO: Hierarchical 3D Semantic Representation for Pose-Aware Object Manipulation

ICRA 2026poster

Imitation learning for robotic manipulation has progressed from 2D image policies to 3D representations that explicitly encode geometry. Yet purely geometric policies often lack explicit part-level semantics, which are critical for pose-aware manipulation (e.g., distinguishing a shoe's toe from heel…

2026

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

ICLR 2026poster

Temporal context is essential for robotic manipulation because such tasks are inherently non-Markovian, yet mainstream VLA models typically overlook it and struggle with long-horizon, temporally dependent tasks. Cognitive science suggests that humans rely on working memory to buffer short-lived repr…

Cited by 0SourcecodeScholar
2026

SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation

AAAI 2026technical

Robotic manipulation requires precise spatial understanding to interact with objects in the real world. Point-based methods suffer from sparse sampling, leading to the loss of fine-grained semantics. Image-based methods typically feed RGB and depth into 2D backbones pre-trained on 3D auxiliary tasks

Cited by 0SourcePDFScholar
2025

BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation

RA-L 2025

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT [1]) tend to treat multi-view features equally and directly concatenate them for policy learning. However, it will introduce redundant visual information and bring h

Cited by 8SourceScholar
2025

Diff-Shadow: Global-guided Diffusion Model for Shadow Removal

AAAI 2025technical

We propose Diff-Shadow, a global-guided diffusion model for high-quality shadow removal. Previous transformer-based approaches can utilize global information to relate shadow and non-shadow regions but are limited in their synthesis ability and recover images with obvious boundaries. In contrast, di…

2025

FlowPolicy: Enabling Fast and Robust 3D Flow-Based Policy via Consistency Flow Matching for Robot Manipulation

AAAI 2025technical

Robots can acquire complex manipulation skills by learning policies from expert demonstrations, which is often known as vision-based imitation learning. Generating policies based on diffusion and flow matching models has been shown to be effective, particularly in robotic manipulation tasks. However…

2025

MegActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer

AAAI 2025technical

Diffusion models have demonstrated superior performance in portrait animation. However, current approaches relied on either visual or audio modality to control character movements, failing to exploit the potential of mixed-modal control. This challenge arises from the difficulty in balancing the we…

2025

Realistic Noise Synthesis with Diffusion Models

AAAI 2025technical

Deep denoising models require extensive real-world training data, which is challenging to acquire. Current noise synthesis techniques struggle to accurately model complex noise distributions. We propose a novel Realistic Noise Synthesis Diffusor (RNSD) method using diffusion models to address these…

2025

Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

ICCV 2025poster

The rapid development of Large Multimodal Models (LMMs) for 2D images and videos has spurred efforts to adapt these models for interpreting 3D scenes. However, the absence of large-scale 3D vision-language datasets has posed a significant obstacle. To address this issue, typical approaches focus on…

Cited by 0SourcePDFScholar
2024

FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models

CVPR 2024highlight

Optical flow estimation a process of predicting pixel-wise displacement between consecutive frames has commonly been approached as a regression task in the age of deep learning. Despite notable advancements this de facto paradigm unfortunately falls short in generalization performance when trained o…

2024

Neural Spectral Decomposition for Dataset Distillation

ECCV 2024poster

"In this paper, we propose Neural Spectrum Decomposition, a generic decomposition framework for dataset distillation. Unlike previous methods, we consider the entire dataset as a high-dimensional observation that is low-rank across all dimensions. We aim to discover the low-rank representation of th…

2024

Sparse Beats Dense: Rethinking Supervision in Radar-Camera Depth Completion

ECCV 2024poster

"It is widely believed that sparse supervision is worse than dense supervision in the field of depth completion, but the underlying reasons for this are rarely discussed. To this end, we revisit the task of radar-camera depth completion and present a new method with sparse LiDAR supervision to outpe…

2024

You Only Look Around: Learning Illumination-Invariant Feature for Low-light Object Detection

NeurIPS 2024poster

In this paper, we introduce YOLA, a novel framework for object detection in low-light scenarios. Unlike previous works, we propose to tackle this challenging problem from the perspective of feature learning. Specifically, we propose to learn illumination-invariant features through the Lambertian ima…

2023

GAFlow: Incorporating Gaussian Attention into Optical Flow

ICCV 2023poster

Optical flow, or the estimation of motion fields from image sequences, is one of the fundamental problems in computer vision. Unlike most pixel-wise tasks that aim at achieving consistent representations of the same category, optical flow raises extra demands for obtaining local discrimination and s…

Cited by 32PDFcodeScholar
2023

Implicit Identity Leakage: The Stumbling Block to Improving Deepfake Detection Generalization

CVPR 2023poster

In this paper, we analyse the generalization ability of binary classifiers for the task of deepfake detection. We find that the stumbling block to their generalization is caused by the unexpected learned identity representation on images. Termed as the Implicit Identity Leakage, this phenomenon has…

2023

MEFLUT: Unsupervised 1D Lookup Tables for Multi-exposure Image Fusion

ICCV 2023poster

In this paper, we introduce a new approach for high-quality multi-exposure image fusion (MEF). We show that the fusion weights of an exposure can be encoded into a 1D lookup table (LUT), which takes pixel intensity value as input and produces fusion weight as output. We learn one 1D LUT for each exp…

Cited by 19PDFcodeScholar
2023

Supervised Homography Learning with Realistic Dataset Generation

ICCV 2023poster

In this paper, we propose an iterative framework, which consists of two phases: a generation phase and a training phase, to generate realistic training data and yield a supervised homography network. In the generation phase, given an unlabeled image pair, we utilize the pre-estimated dominant plane…

Cited by 11PDFcodeScholar
2022

D2C-SR: A Divergence to Convergence Approach for Real-World Image Super-Resolution

ECCV 2022poster

"In this paper, we present D2C-SR, a novel framework for the task of real-world image super-resolution. As an ill-posed problem, the key challenge in super-resolution related tasks is there can be multiple predictions for a given low-resolution input. Most classical deep learning based approaches ig…

2022

Deep Constrained Least Squares for Blind Image Super-Resolution

CVPR 2022poster

In this paper, we tackle the problem of blind image super-resolution(SR) with a reformulated degradation model and two novel modules. Following the common practices of blind SR, our method proposes to improve both the kernel estimation as well as the kernel-based high-resolution image restoration. T…

Cited by 131PDFcodeScholar
2022

Efficient One Pass Self-Distillation with Zipf’s Label Smoothing

ECCV 2022poster

"Self-distillation exploits non-uniform soft supervision from itself during training and improves performance without any runtime cost. However, the overhead during training is often overlooked, and yet reducing time and memory overhead during training is increasingly important in the giant models’…

2022

Explaining Deepfake Detection by Analysing Image Matching

ECCV 2022poster

"This paper aims to interpret how deepfake detection models learn artifact features of images when just supervised by binary labels. To this end, three hypotheses from the perspective of image matching are proposed as follows. 1. Deepfake detection models indicate real/fake images based on visual co…

2022

Learning Optical Flow with Adaptive Graph Reasoning

AAAI 2022technical

Estimating per-pixel motion between video frames, known as optical flow, is a long-standing problem in video understanding and analysis. Most contemporary optical flow techniques largely focus on addressing the cross-image matching with feature similarity, with few methods considering how to explici…

2022

Practical Stereo Matching via Cascaded Recurrent Network With Adaptive Correlation

CVPR 2022oral

With the advent of convolutional neural networks, stereo matching algorithms have recently gained tremendous progress. However, it remains a great challenge to accurately extract disparities from real-world image pairs taken by consumer-level devices like smartphones, due to practical complicating f…

Cited by 318PDFcodeScholar
2022

RealFlow: EM-Based Realistic Optical Flow Dataset Generation from Videos

ECCV 2022poster

"Obtaining the ground truth labels from a video is challenging since the manual annotation of pixel-wise flow labels is prohibitively expensive and laborious. Besides, existing approaches try to adapt the trained model on synthetic datasets to authentic videos, which inevitably suffers from domain d…

2021

FFB6D: A Full Flow Bidirectional Fusion Network for 6D Pose Estimation

CVPR 2021poster

In this work, we present FFB6D, a full flow bidirectional fusion network designed for 6D pose estimation from a single RGBD image. Our key insight is that appearance information in the RGB image and geometry information from the depth image are two complementary data sources, and it still remains un…

Cited by 364PDFcodeScholar
2021

Motion Basis Learning for Unsupervised Deep Homography Estimation With Subspace Projection

ICCV 2021poster

In this paper, we introduce a new framework for unsupervised deep homography estimation. Our contributions are 3 folds. First, unlike previous methods that regress 4 offsets for a homography, we propose a homography flow representation, which can be estimated by a weighted sum of 8 pre-defined homog…

Cited by 70PDFcodeScholar
2021

NBNet: Noise Basis Learning for Image Denoising With Subspace Projection

CVPR 2021poster

In this paper, we introduce NBNet, a novel framework for image denoising. Unlike previous works, we propose to tackle this challenging problem from a new perspective: noise reduction by image-adaptive projection. Specifically, we propose to train a network that can separate signal and noise by learn…

Cited by 280PDFcodeScholar
2021

Practical Wide-Angle Portraits Correction With Deep Structured Models

CVPR 2021poster

Wide-angle portraits often enjoy expanded views. However, they contain perspective distortions, especially noticeable when capturing group portrait photos, where the background is skewed and faces are stretched. This paper introduces the first deep learning based approach to remove such artifacts fr…

Cited by 23PDFcodeScholar
2021

UPFlow: Upsampling Pyramid for Unsupervised Optical Flow Learning

CVPR 2021poster

We present an unsupervised learning approach for optical flow estimation by improving the upsampling and learning of pyramid network. We design a self-guided upsample module to tackle the interpolation blur problem caused by bilinear upsampling between pyramid levels. Moreover, we propose a pyramid…

Cited by 111PDFcodeScholar
2020

PVN3D: A Deep Point-Wise 3D Keypoints Voting Network for 6DoF Pose Estimation

CVPR 2020poster

In this work, we present a novel data-driven method for robust 6DoF object pose estimation from a single RGBD image. Unlike previous methods that directly regressing pose parameters, we tackle this challenging task with a keypoint-based approach. Specifically, we propose a deep Hough voting network…

Cited by 613PDFcodeScholar
2017

A Point Set Generation Network for 3D Object Reconstruction From a Single Image

CVPR 2017oral

Generation of 3D data by deep neural network has been attracting increasing attention in the research community. The majority of extant works resort to regular representations such as volumetric grids or collection of images; however, these representations obscure the natural invariance of 3D shapes…

Cited by 2836PDFcodeScholar