← Search

Haibin Huang

30 accepted papers

2026

Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context

CVPR 2026

Scene-consistent video generation aims to create videos that explore 3D scenes based on a camera trajectory. Previous methods rely on video generation models with external memory for consistency, or iterative 3D reconstruction and inpainting, which accumulate errors during inference due to incorrect

Cited by 0SourceScholar
2026

Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation

CVPR 2026

Group Relative Policy Optimization (GRPO) has emerged as an effective and lightweight framework for post-training visual generative models. However, its performance is fundamentally limited by the ambiguity of textual-visual correspondence: a single prompt may validly describe diverse visual outputs

Cited by 0SourceScholar
2026

MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement

ICLR 2026poster

We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual prompts. This task faces persistent challenges, including identity inconsistency, entanglement among multiple reference s…

Cited by 0SourcecodeScholar
2026

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation

CVPR 2026

Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generators with human preferences. However, existing GRPO pipelines rely on a single scalar reward per sample, treating each im

Cited by 0SourceScholar
2026

TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction

CVPR 2026

We present TempoMaster, a novel framework that formulates long video generation as next-frame-rate prediction. Specifically, we first generate a low-frame-rate clip that serves as a coarse blueprint of the entire video sequence, and then progressively increase the frame rate to refine visual details

Cited by 0SourceScholar
2024

DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive Learning

NeurIPS 2024poster

Current techniques for detecting AI-generated text are largely confined to manual feature crafting and supervised binary classification paradigms. These methodologies typically lead to performance bottlenecks and unsatisfactory generalizability. Consequently, these methods are often inapplicable for…

2024

InterFusion: Text-Driven Generation of 3D Human-Object Interaction

ECCV 2024poster

"In this study, we tackle the complex task of generating 3D human-object interactions (HOI) from textual descriptions in a zero-shot text-to-3D manner. We identify and address two key challenges: the unsatisfactory outcomes of direct text-to-3D methods in HOI, largely due to the lack of paired text-…

2024

You Only Look Around: Learning Illumination-Invariant Feature for Low-light Object Detection

NeurIPS 2024poster

In this paper, we introduce YOLA, a novel framework for object detection in low-light scenarios. Unlike previous works, we propose to tackle this challenging problem from the perspective of feature learning. Specifically, we propose to learn illumination-invariant features through the Lambertian ima…

2023

Explicit Motion Disentangling for Efficient Optical Flow Estimation

ICCV 2023poster

In this paper, we propose a novel framework for optical flow estimation that achieves a good balance between performance and efficiency. Our approach involves disentangling global motion learning from local flow estimation, treating global matching and local refinement as separate stages. We offer t…

Cited by 18PDFcodeScholar
2023

HairStep: Transfer Synthetic to Real Using Strand and Depth Maps for Single-View 3D Hair Modeling

CVPR 2023highlight

In this work, we tackle the challenging problem of learning-based single-view 3D hair modeling. Due to the great difficulty of collecting paired real image and 3D hair data, using synthetic data to provide prior knowledge for real domain becomes a leading solution. This unfortunately introduces the…

Cited by 26SourcePDFScholar
2023

Inversion-Based Style Transfer With Diffusion Models

CVPR 2023poster

The artistic style within a painting is the means of expression, which includes not only the painting material, colors, and brushstrokes, but also the high-level attributes, including semantic elements and object shapes. Previous arbitrary example-guided artistic image generation methods often fail…

2023

Self-Supervised Category-Level Articulated Object Pose Estimation with Part-Level SE(3) Equivariance

ICLR 2023poster

Category-level articulated object pose estimation aims to estimate a hierarchy of articulation-aware object poses of an unseen articulated object from a known category. To reduce the heavy annotations needed for supervised learning methods, we present a novel self-supervised strategy that solves thi…

2023

Semi-Weakly Supervised Object Kinematic Motion Prediction

CVPR 2023poster

Given a 3D object, kinematic motion prediction aims to identify the mobile parts as well as the corresponding motion parameters. Due to the large variations in both topological structure and geometric details of 3D objects, this remains a challenging task and the lack of large scale labeled data als…

Cited by 11SourcePDFScholar
2022

D2C-SR: A Divergence to Convergence Approach for Real-World Image Super-Resolution

ECCV 2022poster

"In this paper, we present D2C-SR, a novel framework for the task of real-world image super-resolution. As an ill-posed problem, the key challenge in super-resolution related tasks is there can be multiple predictions for a given low-resolution input. Most classical deep learning based approaches ig…

2022

Deep Constrained Least Squares for Blind Image Super-Resolution

CVPR 2022poster

In this paper, we tackle the problem of blind image super-resolution(SR) with a reformulated degradation model and two novel modules. Following the common practices of blind SR, our method proposes to improve both the kernel estimation as well as the kernel-based high-resolution image restoration. T…

Cited by 131PDFcodeScholar
2021

Arbitrary Video Style Transfer via Multi-Channel Correlation

AAAI 2021technical

Video style transfer is attracting increasing attention from the artificial intelligence community because of its numerous applications, such as augmented reality and animation production. Relative to traditional image style transfer, video style transfer presents new challenges, including how to ef…

2021

FFB6D: A Full Flow Bidirectional Fusion Network for 6D Pose Estimation

CVPR 2021poster

In this work, we present FFB6D, a full flow bidirectional fusion network designed for 6D pose estimation from a single RGBD image. Our key insight is that appearance information in the RGB image and geometry information from the depth image are two complementary data sources, and it still remains un…

Cited by 364PDFcodeScholar
2021

HPNet: Deep Primitive Segmentation Using Hybrid Representations

ICCV 2021poster

This paper introduces HPNet, a novel deep-learning approach for segmenting a 3D shape represented as a point cloud into primitive patches. The key to deep primitive segmentation is learning a feature representation that can separate points of different primitives. Unlike utilizing a single feature r…

Cited by 56PDFcodeScholar
2021

NBNet: Noise Basis Learning for Image Denoising With Subspace Projection

CVPR 2021poster

In this paper, we introduce NBNet, a novel framework for image denoising. Unlike previous works, we propose to tackle this challenging problem from a new perspective: noise reduction by image-adaptive projection. Specifically, we propose to train a network that can separate signal and noise by learn…

Cited by 280PDFcodeScholar
2021

Scene Synthesis via Uncertainty-Driven Attribute Synchronization

ICCV 2021poster

Developing deep neural networks to generate 3D scenes is a fundamental problem in neural synthesis with immediate applications in architectural CAD, computer graphics, as well as in generating virtual robot training environments. This task is challenging because 3D scenes exhibit diverse patterns, r…

Cited by 39PDFcodeScholar
2020

FPConv: Learning Local Flattening for Point Convolution

CVPR 2020poster

We introduce FPConv, a novel surface-style convolution operator designed for 3D point cloud analysis. Unlike previous methods, FPConv doesn't require transforming to intermediate representation like 3D grid or graph and directly works on surface geometry of point cloud. To be more specific, for each…

Cited by 187PDFcodeScholar
2020

PVN3D: A Deep Point-Wise 3D Keypoints Voting Network for 6DoF Pose Estimation

CVPR 2020poster

In this work, we present a novel data-driven method for robust 6DoF object pose estimation from a single RGBD image. Unlike previous methods that directly regressing pose parameters, we tackle this challenging task with a keypoint-based approach. Specifically, we propose a deep Hough voting network…

Cited by 613PDFcodeScholar
2020

Practical Deep Raw Image Denoising on Mobile Devices

ECCV 2020poster

Deep learning-based image denoising approaches have been extensively studied in recent years, prevailing in many public benchmark datasets. However, the stat-of-the-art networks are computationally too expensive to be directly applied on mobile devices. In this work, we propose a light-weight, effic…

2020

SAM: Squeeze-and-Mimic Networks for Conditional Visual Driving Policy Learning

CoRL 2020

We describe a policy learning approach to map visual inputs to driving controls conditioned on turning command that leverages side tasks on semantics and object affordances via a learned representation trained for driving. To learn this representation, we train a squeeze network to drive using annot

2019

GIF2Video: Color Dequantization and Temporal Interpolation of GIF Images

CVPR 2019poster

Graphics Interchange Format (GIF) is a highly portable graphics format that is ubiquitous on the Internet. Despite their small sizes, GIF images often contain undesirable visual artifacts such as flat color regions, false contours, color shift, and dotted patterns. In this paper, we propose GIF2Vide…

Cited by 30PDFScholar
2019

Semi-Supervised Skin Detection by Network With Mutual Guidance

ICCV 2019poster

We present a new data-driven method for robust skin detection from a single human portrait image. Unlike previous methods, we incorporate human body as a weak semantic guidance into this task, considering acquiring large-scale of human labeled skin data is commonly expensive and time-consuming. To b…

Cited by 32PDFScholar
2017

High-Resolution Shape Completion Using Deep Neural Networks for Global Structure and Local Geometry Inference

ICCV 2017spotlight

We propose a data-driven method for recovering missing parts of 3D shapes. Our method is based on a new deep learning architecture consisting of two sub-networks: a global structure inference network and a local geometry refinement network. The global structure inference network incorporates a long…

Cited by 367PDFScholar
2017

Synthesizing 3D Shapes via Modeling Multi-View Depth Maps and Silhouettes With Deep Generative Networks

CVPR 2017poster

We study the problem of learning generative models of 3D shapes. Voxels or 3D parts have been widely used as the underlying representations to build complex 3D shapes; however, voxel-based representations suffer from high memory requirements, and parts-based models require a large collection of cach…

Cited by 256PDFScholar